Prerequisites
- Two providers configured (e.g., OpenAI and Anthropic) — see Providers
- Both target models registered as logical models
- Admin access to create gateway configs
Steps
Create the canary gateway config
A gateway config defines how traffic is split. Here we send 90% to the current model (GPT-4o) and 10% to the new model (Claude 3.5 Sonnet).
Assign the config to a team (or use it per-request)
You can set the canary config as the default for a team, so all their requests use it automatically:Or, opt specific requests into the canary config via header — useful for testing before rolling out to the whole team:
Monitor the split in metrics
Check the Prometheus metrics to see traffic distribution by model:Compare error rates and latency between the two models. If the new model is performing well, increase its weight.
Increase the canary weight gradually
Once you’re satisfied with the 10% canary, update the weights:Continue increasing the weight at each step: 10% → 25% → 50% → 90% → 100%. At each step, validate metrics and audit log quality before proceeding.
What to watch during the canary
- Error rate — any spike in 5xx errors on the new model is a signal to pause and investigate
- Latency — compare p95 latency between models; some models are slower for certain workloads
- Token usage — some models use more tokens to express the same content, affecting cost
- Audit log — sample responses from both models in the audit log to qualitatively assess output quality
Next steps
- Set up evals to quantitatively compare model quality side by side
- Configure failover routing so if the new model fails, traffic falls back automatically
- Review metrics for ongoing monitoring after migration