Prerequisites
- Workspace enabled with at least two chat models configured
- A third model available to serve as the judge (can be the same as one of the contestants)
Run the evaluation
- Generates a response from both models
- Calls the judge model twice with A/B order swapped (to cancel position bias)
- Agreement = clear winner; disagreement = tie
- Updates ELO ratings after each match (K=32)