Routing happens at two levels.
LevelChooses betweenConfigured with
Endpoint routing (the router)the upstream endpoints of one modelthe router: block of gateway.yaml and a model’s upstreams[]
Routing configsdifferent models, for failover, canary, latency and conditional routingnamed configs, created in Gateway → Routing or /admin/gateway/configs
Endpoint routing needs a model with more than one endpoint, which only a gateway.yaml model with upstreams[] has. A model registered through the console has a single endpoint, so it is routed with a routing config instead.

Endpoint routing

The default strategy applies to every request, whatever config it carries. These are the settings with their defaults:
gateway.yaml
router:
  strategy: least_inflight   # least_inflight | round_robin | prefix_affinity | score
  health_interval: 15s       # how often ejected endpoints are probed for recovery
  failure_threshold: 3       # consecutive failures before ejecting an endpoint
  cooldown: 30s              # base ejection window
  session_ttl: 10m           # sticky session lifetime
  prefix_bytes: 1024         # prompt prefix bytes for prefix_affinity
  health_threshold: 0        # eject when the success-rate average falls below this (0 = off, max 1)
  restore_health: 0.5        # health an ejected endpoint returns with (0 = default 0.5)
  max_ejection_duration: 0   # cap on the escalating ejection window (0 = uncapped)
  respect_retry_after: true  # use a rate-limited provider's Retry-After as the ejection window
Operators can also change strategy (except score), health_interval, failure_threshold, cooldown, prefix_bytes and session_ttl in the administration console’s settings (Router section); the other fields are gateway.yaml only.
StrategyDescription
least_inflightRoutes to the endpoint with the fewest active requests; a higher weight breaks ties. Best general-purpose choice.
round_robinRotates across healthy endpoints.
prefix_affinityHashes the first prefix_bytes of the prompt to pick the same endpoint for similar prompts, which helps KV cache locality. Falls back to least in-flight when it has no prompt.
scoreGrades each endpoint on a moving average of success rate and latency, penalised by queue depth: health / (1 + latency × (1 + pending × 0.1)). Two endpoints are drawn at random and the better one wins, so traffic does not herd onto one. It sheds traffic from an endpoint that has gone slow or started erroring before it fails outright.

Sticky sessions

When a request carries an X-Session-Id header (or a user field in the body), the router remembers the endpoint that served that session for session_ttl and prefers it on later requests, under any strategy. If the endpoint is no longer usable, the session is remapped. On a virtual model, the same key also keeps a session on one side of a canary split.

Upstream health

  • Ejection. An endpoint is taken out of rotation after failure_threshold consecutive failures, or when its health average falls below health_threshold. Connection errors, timeouts, 5xx, 429, 408 and 401/402/403 from the provider count as failures; other 4xx responses reject what the caller sent and do not.
  • Window. The ejection lasts cooldown, multiplied by the number of ejections since the last success, up to max_ejection_duration. When the provider states its own wait (Retry-After) and respect_retry_after is on, that wait is used instead, unescalated.
  • Recovery. Every health_interval the gateway probes ejected endpoints with GET /v1/models. One that answers comes back at restore_health, so it earns traffic back instead of taking its old share at once. A real success clears the escalation.
  • Failover. With several endpoints, a request that fails before any byte reaches the client moves to the next endpoint. If every endpoint is ejected, the request fails immediately with provider_unavailable.
Check real-time endpoint health for the models in gateway.yaml:
curl https://app.manylayers.io/admin/gateway/health \
  -H "Authorization: Bearer ml_pat_..."
The response lists each endpoint’s health (healthy, ejected_until, inflight, consecutive_fails, score and latency averages) and the last hour’s request count, error rate and p95 latency per model. It requires the analytics permission in the workspace.

Routing configs

A routing config picks among models. Create one under Gateway → Routing, or with the API; the full schema, retry policy and per-target settings are on Virtual models.
curl -X POST https://app.manylayers.io/admin/gateway/configs \
  -H "Authorization: Bearer ml_pat_..." \
  -H "Content-Type: application/json" \
  -d '{
    "name": "resilient",
    "strategy": "fallback",
    "targets": [
      {"model": "gpt-4o"},
      {"model": "claude-sonnet"}
    ],
    "retry": {
      "attempts": 1,
      "backoff_ms": 200,
      "on_status_codes": [429, 500]
    }
  }'
A personal access token is bound to its workspace; with a console session add ?workspace_id=<id>. Configs belong to one workspace and serve requests made in it.

Routing strategies

Try targets in order. A target is retried per retry, then the request moves to the next target on a status in fallback_status_codes (by default 400 401 403 404 408 413 422 429 500 502 503 504, plus the retry codes).
{
  "strategy": "fallback",
  "targets": [
    {"model": "gpt-4o"},
    {"model": "claude-sonnet"}
  ],
  "retry": {"attempts": 1, "backoff_ms": 200, "on_status_codes": [429, 500]}
}

Using a routing config

A config applies to a request in one of three ways, in this order: the request’s model is the name of a config that has model_types (a virtual model), the X-ManyLayers-Config header names a config by id or name, or the team’s default_config_id applies. An unknown header value returns 400 unknown_config.
curl https://app.manylayers.io/v1/chat/completions \
  -H "Authorization: Bearer $MANYLAYERS_API_KEY" \
  -H "X-ManyLayers-Config: resilient" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4o", "messages": [{"role": "user", "content": "Hi"}]}'
The response names what was decided in X-ManyLayers-Config, X-ManyLayers-Routing-Strategy, X-ManyLayers-Routing-Model-Order and X-ManyLayers-Target.

Config management API

All routes are on https://app.manylayers.io and need the routing permission in the workspace (gateway.routing.read to read, gateway.routing.manage to change).
MethodPathDescription
GET/admin/gateway/configsList the workspace’s routing configs
POST/admin/gateway/configsCreate a config
GET/admin/gateway/configs/{id}Get a config
PUT/admin/gateway/configs/{id}Update a config
DELETE/admin/gateway/configs/{id}Delete a config
GET/admin/gateway/configs/usageLast 30 days of requests per virtual model and target