Eval modes
- Standard
- Pairwise ELO
Run a set of prompts through one or more models and collect results. You can score results with a judge model that automatically rates each response, or rate them yourself with thumbs up/down. Results are stored per-prompt for analysis and comparison over time.
Running an evaluation
Create an eval
Go to Evals in the sidebar and click New Eval. Give it a name and choose the mode — Standard or Pairwise ELO.
Select models
Choose the model (or models, for pairwise) you want to evaluate. For pairwise, also select a judge model that will score the responses.
Add prompts
Add the prompts you want to test. You can type them directly, import from a conversation, or paste a list. Each prompt becomes one item in the eval set.
Run
Click Run Eval. ManyLayers sends each prompt to the selected models, collects responses, and runs judging automatically in the background. You’ll see results populate as they come in.
Message feedback as a signal
You don’t have to run a formal eval to influence model rankings. When you give a thumbs up or down on any assistant message in chat, it feeds into the ELO system:- Thumbs up — treated as a win against a virtual 1000-rated opponent
- Thumbs down — treated as a loss