model_types set: clients send its name as model, and the gateway serves the request from the config’s targets. You change providers, add a fallback or shift traffic without touching client code.
A config belongs to one workspace and only serves requests made in that workspace. This page covers routing across models. To choose between endpoints of one model, see Routing.
How it works
A config is applied to a request in one of three ways, in this order of precedence:- Called by name. The request’s
modelis the name of an enabled config that hasmodel_types. X-ManyLayers-Configheader. The value is a config’s id or name. An unknown one returns400 unknown_config.- Team default. The team’s
default_config_id. A stale or disabled default is ignored, and so is one whosemodel_typesdo not cover the endpoint.
Real model names always win. A config is only considered when the requested name resolves to no model, so creating a config can never change what an existing model name serves. The API refuses a virtual model whose name is already a model. A virtual model is called by name but is not listed by
GET /v1/models.Configuration structure
Configs are created in the console under Gateway → Routing (the Virtual Models page, Add Virtual Model) or through/admin/gateway/configs on https://app.manylayers.io (GET, POST, and GET/PUT/DELETE on /{id}). The console labels the strategies by what they do: Weight is loadbalance (or canary with sticky routing on), Priority is fallback, plus Latency and Auto Routing; single and conditional are also available.
Key fields
single— the first target only.fallback— targets in declared order.loadbalance— a weighted draw for every position, so a failed target’s share spreads across the rest in proportion.canary— one weighted pick, sticky perX-Session-Id(or bodyuser); no cross-target fallback.conditional— the first target whosewhenmatches, else the first target with nowhen; no cross-target fallback.latency— the fastest recently measured available target first, then the next fastest, then unmeasured targets; 5% of requests try another target to keep measurements fresh.complexity— tiered Auto Routing.
Makes the config a virtual model, callable by name on the endpoints those types name:
chat (/v1/chat/completions, /v1/messages), completion (/v1/completions), responses (/v1/responses), embedding (/v1/embeddings). Calling it on any other endpoint returns 400. Every target of a virtual model must name a model.Under
conditional, selects the target. Under every other strategy, removes targets that do not match before the plan is made. Up to 32 metadata pairs; keys up to 128 bytes, values up to 512.Upstream statuses (400–599) that move the request to the next target once retries are spent. Default:
400 401 403 404 408 413 422 429 500 502 503 504, plus the target’s retry codes.Merged over the request body for this target only. Nothing one target overrides reaches the next.
Set on the upstream request for this target. At most 32. Credential headers (
Authorization, Proxy-Authorization, Api-Key, X-Api-Key, Anthropic-Api-Key, X-Goog-Api-Key, AWS signing headers), framing headers (Content-Type, Content-Length, Transfer-Encoding), Cookie, Host and any X-ManyLayers-* header are refused.tier, prompt_version and classification_strategy are valid only with complexity; see Auto Routing. prompt_version is refused on this gateway because no prompt registry is configured.
A config that fails validation answers 400 bad_request with the sentence to act on; a name already used by another config in the workspace answers 409 conflict. Each target’s model must be a model that resolves in the workspace (by its gateway name or account/model address) or in gateway.yaml.
Common configurations
Fallback chain across providers
Fallback chain across providers
?workspace_id=<id> to name the workspace.Weighted split
Weighted split
Canary rollout
Canary rollout
X-Session-Id so each conversation stays on one side of the split.Conditional on metadata
Conditional on metadata
X-ManyLayers-Metadata: {"region":"eu"}. The target with no when is the default.Lowest latency
Lowest latency
Embeddings with failover
Embeddings with failover
Calling a virtual model
X-ManyLayers-Config (the config name), X-ManyLayers-Routing-Strategy, X-ManyLayers-Routing-Model-Order (the order targets would be tried), X-ManyLayers-Target (the model that served), X-ManyLayers-Retries, and X-ManyLayers-Skipped-Targets when targets were passed over, as model=reason pairs (for example gpt-4o=model_restriction; reasons are metadata_mismatch, not_allowed, model_restriction, unresolvable, ambiguous, budget_exceeded and rate_limited). With strict_openai_compliance on, the X-ManyLayers-* headers are not sent.
Behaviour notes
A provider is chosen before the first byte reaches the client and is never switched mid-stream. Statuses are held for retry or fallback only while another attempt could still use them; the last attempt relays directly.
Model restrictions and RBAC apply target by target: a target the caller may not use is skipped. A team with an explicit model list must also include the virtual model’s name, and a virtual account token must have it in its allowed models.
Next steps
Auto Routing
Route each request to a model sized for it.
Routing
Endpoint-level load balancing and health checks.
Providers
The models your targets can point at.
Rate limiting
Limits that apply to the requested name.