Outcome: Gradually migrate traffic from one model or provider to another without a big-bang cutover. Validate quality and performance at each step before increasing traffic. This pattern is useful when you’re switching from GPT-4o to Claude, upgrading to a new model version, or shifting from a hosted provider to a self-hosted open-weight model.

Prerequisites

  • Two providers configured (e.g., OpenAI and Anthropic) — see Providers
  • Both target models registered as logical models
  • Admin access to create gateway configs

Steps

1

Create the canary gateway config

A gateway config defines how traffic is split. Here we send 90% to the current model (GPT-4o) and 10% to the new model (Claude 3.5 Sonnet).
curl -X POST http://localhost:8180/admin/gateway-configs \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "gpt4o-to-claude-canary",
    "strategy": "canary",
    "backends": [
      {
        "model": "gpt-4o",
        "provider_id": "openai-main",
        "weight": 90
      },
      {
        "model": "claude-3-5-sonnet",
        "provider_id": "anthropic-main",
        "weight": 10
      }
    ]
  }'
2

Assign the config to a team (or use it per-request)

You can set the canary config as the default for a team, so all their requests use it automatically:
curl -X PATCH http://localhost:8180/admin/teams/$TEAM_ID \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{"default_gateway_config": "gpt4o-to-claude-canary"}'
Or, opt specific requests into the canary config via header — useful for testing before rolling out to the whole team:
curl http://localhost:8180/v1/chat/completions \
  -H "Authorization: Bearer $YOUR_KEY" \
  -H "X-ManyLayers-Config: gpt4o-to-claude-canary" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4o", "messages": [...]}'
3

Monitor the split in metrics

Check the Prometheus metrics to see traffic distribution by model:
# Requests per model over the last 5 minutes
rate(manylayers_requests_total[5m]) by (model)

# Error rate per model
rate(manylayers_requests_total{status=~"5.."}[5m]) by (model)

# Latency comparison
histogram_quantile(0.95, rate(manylayers_request_duration_seconds_bucket[5m])) by (model)
Compare error rates and latency between the two models. If the new model is performing well, increase its weight.
4

Increase the canary weight gradually

Once you’re satisfied with the 10% canary, update the weights:
# Move to 50/50
curl -X PATCH http://localhost:8180/admin/gateway-configs/gpt4o-to-claude-canary \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "backends": [
      {"model": "gpt-4o", "provider_id": "openai-main", "weight": 50},
      {"model": "claude-3-5-sonnet", "provider_id": "anthropic-main", "weight": 50}
    ]
  }'
Continue increasing the weight at each step: 10% → 25% → 50% → 90% → 100%. At each step, validate metrics and audit log quality before proceeding.
5

Complete the migration

When you’re ready to fully migrate, update the config to send 100% traffic to the new model — or simply update the team’s default routing to point directly at the new model and remove the canary config.
curl -X PATCH http://localhost:8180/admin/teams/$TEAM_ID \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{"default_model": "claude-3-5-sonnet", "default_gateway_config": null}'

What to watch during the canary

  • Error rate — any spike in 5xx errors on the new model is a signal to pause and investigate
  • Latency — compare p95 latency between models; some models are slower for certain workloads
  • Token usage — some models use more tokens to express the same content, affecting cost
  • Audit log — sample responses from both models in the audit log to qualitatively assess output quality

Next steps

  • Set up evals to quantitatively compare model quality side by side
  • Configure failover routing so if the new model fails, traffic falls back automatically
  • Review metrics for ongoing monitoring after migration