A virtual model is a routing config with model_types set: clients send its name as model, and the gateway serves the request from the config’s targets. You change providers, add a fallback or shift traffic without touching client code. A config belongs to one workspace and only serves requests made in that workspace. This page covers routing across models. To choose between endpoints of one model, see Routing.

How it works

A config is applied to a request in one of three ways, in this order of precedence:
  1. Called by name. The request’s model is the name of an enabled config that has model_types.
  2. X-ManyLayers-Config header. The value is a config’s id or name. An unknown one returns 400 unknown_config.
  3. Team default. The team’s default_config_id. A stale or disabled default is ignored, and so is one whose model_types do not cover the endpoint.
Real model names always win. A config is only considered when the requested name resolves to no model, so creating a config can never change what an existing model name serves. The API refuses a virtual model whose name is already a model. A virtual model is called by name but is not listed by GET /v1/models.

Configuration structure

Configs are created in the console under Gateway → Routing (the Virtual Models page, Add Virtual Model) or through /admin/gateway/configs on https://app.manylayers.io (GET, POST, and GET/PUT/DELETE on /{id}). The console labels the strategies by what they do: Weight is loadbalance (or canary with sticky routing on), Priority is fallback, plus Latency and Auto Routing; single and conditional are also available.
{
  "name": "support-chat",              // what clients send as "model"
  "strategy": "fallback",              // single | fallback | loadbalance | canary | conditional | latency | complexity
  "model_types": ["chat"],             // chat | completion | responses | embedding; omit for header/team-default only
  "enabled": true,                     // default true on create
  "request_timeout_ms": 30000,         // deadline over the whole target loop, 0-600000; 0 = none
  "strict_openai_compliance": false,   // true drops X-ManyLayers-* headers (except Cache) and extension fields
  "retry": {                           // config-wide retry policy, per target
    "attempts": 1,                     // retries after the first call, 0-10
    "backoff_ms": 200,                 // grows linearly: backoff_ms x attempt
    "max_backoff_ms": 2000,            // cap on one wait, 0 = none
    "jitter_pct": 20,                  // +/- spread; 0 = default 20, -1 = off
    "ignore_retry_after": false,       // true ignores the provider's Retry-After
    "on_status_codes": [429, 503]      // default 429, 500, 502, 503, 504
  },
  "targets": [
    {
      "model": "gpt-4o",               // a workspace or gateway.yaml model, or an "account/model" address; "" = the requested model (not for virtual models)
      "weight": 80,                    // loadbalance / canary only; default 1, 0 = off
      "when": {                        // all conditions must match
        "header": "X-Tier", "equals": "gold",
        "metadata": {"region": "us"}   // matched against X-ManyLayers-Metadata
      },
      "override_params": {"temperature": 0.2},
      "override_headers": {"OpenAI-Organization": "org-123"},
      "retry": {"attempts": 2},        // replaces the config retry for this target
      "fallback_status_codes": [429, 500, 503],
      "fallback_candidate": true       // false = only serves when picked first
    }
  ]
}

Key fields

strategy
string
required
  • single — the first target only.
  • fallback — targets in declared order.
  • loadbalance — a weighted draw for every position, so a failed target’s share spreads across the rest in proportion.
  • canary — one weighted pick, sticky per X-Session-Id (or body user); no cross-target fallback.
  • conditional — the first target whose when matches, else the first target with no when; no cross-target fallback.
  • latency — the fastest recently measured available target first, then the next fastest, then unmeasured targets; 5% of requests try another target to keep measurements fresh.
  • complexity — tiered Auto Routing.
model_types
string[]
Makes the config a virtual model, callable by name on the endpoints those types name: chat (/v1/chat/completions, /v1/messages), completion (/v1/completions), responses (/v1/responses), embedding (/v1/embeddings). Calling it on any other endpoint returns 400. Every target of a virtual model must name a model.
targets[].when
object
Under conditional, selects the target. Under every other strategy, removes targets that do not match before the plan is made. Up to 32 metadata pairs; keys up to 128 bytes, values up to 512.
targets[].fallback_status_codes
int[]
Upstream statuses (400–599) that move the request to the next target once retries are spent. Default: 400 401 403 404 408 413 422 429 500 502 503 504, plus the target’s retry codes.
targets[].override_params
object
Merged over the request body for this target only. Nothing one target overrides reaches the next.
targets[].override_headers
object
Set on the upstream request for this target. At most 32. Credential headers (Authorization, Proxy-Authorization, Api-Key, X-Api-Key, Anthropic-Api-Key, X-Goog-Api-Key, AWS signing headers), framing headers (Content-Type, Content-Length, Transfer-Encoding), Cookie, Host and any X-ManyLayers-* header are refused.
tier, prompt_version and classification_strategy are valid only with complexity; see Auto Routing. prompt_version is refused on this gateway because no prompt registry is configured. A config that fails validation answers 400 bad_request with the sentence to act on; a name already used by another config in the workspace answers 409 conflict. Each target’s model must be a model that resolves in the workspace (by its gateway name or account/model address) or in gateway.yaml.

Common configurations

curl -X POST https://app.manylayers.io/admin/gateway/configs \
  -H "Authorization: Bearer ml_pat_..." -H "Content-Type: application/json" \
  -d '{
    "name": "support-chat", "strategy": "fallback", "model_types": ["chat"],
    "targets": [{"model": "gpt-4o"}, {"model": "claude-sonnet"}, {"model": "gpt-4o-mini"}],
    "retry": {"attempts": 1, "backoff_ms": 200}
  }'
A personal access token is bound to the workspace it was issued for. With a console session, add ?workspace_id=<id> to name the workspace.
{"name": "bulk-chat", "strategy": "loadbalance", "model_types": ["chat"],
 "targets": [{"model": "gpt-4o-mini", "weight": 70}, {"model": "llama-fast", "weight": 30}]}
The second position is drawn from the remaining targets, so a failing target hands its traffic to the others.
{"name": "assistant", "strategy": "canary", "model_types": ["chat", "responses"],
 "targets": [{"model": "gpt-4o", "weight": 95}, {"model": "gpt-4.1", "weight": 5}]}
Send X-Session-Id so each conversation stays on one side of the split.
{"name": "regional-chat", "strategy": "conditional", "model_types": ["chat"],
 "targets": [
   {"model": "gpt-4o-eu", "when": {"metadata": {"region": "eu"}}},
   {"model": "gpt-4o"}
 ]}
Clients send X-ManyLayers-Metadata: {"region":"eu"}. The target with no when is the default.
{"name": "fast-chat", "strategy": "latency", "model_types": ["chat"],
 "targets": [{"model": "gpt-4o-mini"}, {"model": "gemini-flash"}, {"model": "llama-fast"}]}
Targets whose endpoints are all ejected are skipped while any other is available.
{"name": "embed", "strategy": "fallback", "model_types": ["embedding"],
 "targets": [{"model": "text-embedding-3-small"}, {"model": "embed-azure"}]}

Calling a virtual model

curl https://app.manylayers.io/v1/chat/completions \
  -H "Authorization: Bearer ml_vat_..." -H "Content-Type: application/json" \
  -d '{"model": "support-chat", "messages": [{"role": "user", "content": "Hi"}]}'
The response carries X-ManyLayers-Config (the config name), X-ManyLayers-Routing-Strategy, X-ManyLayers-Routing-Model-Order (the order targets would be tried), X-ManyLayers-Target (the model that served), X-ManyLayers-Retries, and X-ManyLayers-Skipped-Targets when targets were passed over, as model=reason pairs (for example gpt-4o=model_restriction; reasons are metadata_mismatch, not_allowed, model_restriction, unresolvable, ambiguous, budget_exceeded and rate_limited). With strict_openai_compliance on, the X-ManyLayers-* headers are not sent.

Behaviour notes

A provider is chosen before the first byte reaches the client and is never switched mid-stream. Statuses are held for retry or fallback only while another attempt could still use them; the last attempt relays directly.
Model restrictions and RBAC apply target by target: a target the caller may not use is skipped. A team with an explicit model list must also include the virtual model’s name, and a virtual account token must have it in its allowed models.
A retry honours the provider’s Retry-After (capped by max_backoff_ms). If the wait would outlast request_timeout_ms, the gateway moves to the next target instead of sleeping. When the deadline passes, the request fails with 504.

Next steps

Auto Routing

Route each request to a model sized for it.

Routing

Endpoint-level load balancing and health checks.

Providers

The models your targets can point at.

Rate limiting

Limits that apply to the requested name.