This guide shows you how to set up a pairwise evaluation that compares two models on a set of prompts, uses a judge model to pick winners, and builds an ELO leaderboard to track quality over time.

Prerequisites

  • Workspace enabled with at least two chat models configured
  • A third model available to serve as the judge (can be the same as one of the contestants)
1

Create an evaluation

curl -X POST http://localhost:8180/w/evals \
  -H "Cookie: session=..." \
  -H "Content-Type: application/json" \
  -d '{
    "name": "GPT-4o vs Claude Sonnet",
    "mode": "pairwise",
    "models": ["gpt-4o", "claude-sonnet"],
    "judge_model": "gpt-4o"
  }'
Note the returned eval id.
2

Add your test prompts

Add the prompts you want both models to answer:
curl -X POST http://localhost:8180/w/evals/$EVAL_ID/items \
  -H "Cookie: session=..." \
  -H "Content-Type: application/json" \
  -d '{"prompt": "Explain the difference between TCP and UDP."}'

curl -X POST http://localhost:8180/w/evals/$EVAL_ID/items \
  -H "Cookie: session=..." \
  -H "Content-Type: application/json" \
  -d '{"prompt": "Write a Python function to merge two sorted lists."}'

curl -X POST http://localhost:8180/w/evals/$EVAL_ID/items \
  -H "Cookie: session=..." \
  -H "Content-Type: application/json" \
  -d '{"prompt": "What are the main challenges of distributed systems?"}'
3

Run the evaluation

curl -X POST http://localhost:8180/w/evals/$EVAL_ID/run \
  -H "Cookie: session=..."
For each prompt, ManyLayers:
  1. Generates a response from both models
  2. Calls the judge model twice with A/B order swapped (to cancel position bias)
  3. Agreement = clear winner; disagreement = tie
  4. Updates ELO ratings after each match (K=32)
4

Check the leaderboard

curl http://localhost:8180/w/leaderboard \
  -H "Cookie: session=..."
{
  "ratings": [
    {"model": "gpt-4o", "rating": 1032.5, "matches": 3, "votes": 0},
    {"model": "claude-sonnet", "rating": 967.5, "matches": 3, "votes": 0}
  ]
}
5

Export results

Download the leaderboard as CSV for reporting:
curl http://localhost:8180/w/leaderboard/export \
  -H "Cookie: session=..." \
  --output leaderboard.csv

Collect ongoing feedback via message thumbs

Your team members can also influence ELO ratings through message feedback in regular chat:
curl -X POST http://localhost:8180/w/feedback/message \
  -H "Cookie: session=..." \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4o", "vote": 1}'
A thumbs-up is treated as a win against a virtual 1000-rated opponent; a thumbs-down as a loss. This signal accumulates alongside formal eval matches, giving you a continuous measure of model quality in real usage.

View detailed results

# List runs for an eval
curl http://localhost:8180/w/evals/$EVAL_ID/runs \
  -H "Cookie: session=..."

# Get detailed results for a specific run
curl http://localhost:8180/w/eval-runs/$RUN_ID \
  -H "Cookie: session=..."