Evaluations let your team systematically test model quality across a set of prompts — so you can make data-driven decisions about which models to use, rather than going by feel.

Eval modes

Run a set of prompts through one or more models and collect results. You can score results with a judge model that automatically rates each response, or rate them yourself with thumbs up/down. Results are stored per-prompt for analysis and comparison over time.

Running an evaluation

1

Create an eval

Go to Evals in the sidebar and click New Eval. Give it a name and choose the mode — Standard or Pairwise ELO.
2

Select models

Choose the model (or models, for pairwise) you want to evaluate. For pairwise, also select a judge model that will score the responses.
3

Add prompts

Add the prompts you want to test. You can type them directly, import from a conversation, or paste a list. Each prompt becomes one item in the eval set.
4

Run

Click Run Eval. ManyLayers sends each prompt to the selected models, collects responses, and runs judging automatically in the background. You’ll see results populate as they come in.
5

Review results

Open the completed run to see each prompt, the model’s response, and the judge’s verdict or score. For standard evals, you can also add your own rating to any result.

Message feedback as a signal

You don’t have to run a formal eval to influence model rankings. When you give a thumbs up or down on any assistant message in chat, it feeds into the ELO system:
  • Thumbs up — treated as a win against a virtual 1000-rated opponent
  • Thumbs down — treated as a loss
This means everyday feedback from your team accumulates alongside formal eval results, keeping the leaderboard current without extra effort.

Leaderboard

The Leaderboard (accessible from the Evals section in the sidebar) shows ELO ratings for all models that have participated in pairwise evaluations or received message feedback. Higher rated models have performed better on your team’s actual prompts. You can export the leaderboard as a CSV from the leaderboard page — useful for sharing results with stakeholders or tracking changes over time.