Auto Routing sends easy requests to a cheap model and hard ones to a capable model, behind one name. Your client calls "model": "smart-chat"; the gateway classifies the turn, picks the tier, and walks a chain of targets that falls back and escalates on failure. It is the complexity strategy of a virtual model, so it shares the same admin API, retry policy, per-target overrides and streaming guarantee. In the console, create it under Gateway → Routing → Add Virtual Model and choose Auto Routing. The config belongs to one workspace and serves requests made in it.

How it works

Configuration structure

{
  "name": "smart-chat",
  "strategy": "complexity",
  "model_types": ["chat", "responses"],      // optional; default chat, completion, responses
  "classification_strategy": {
    "type": "llm-classifier",                // heuristic (default) | llm-classifier
    "model": "gpt-4o-mini",                  // a concrete chat model, not a routing config
    "timeout_ms": 2000,                      // default 2000, max 30000
    "fallback_strategy": {"type": "static", "default_tier": "medium"}   // or {"type": "heuristic"}
  },
  "targets": [
    {"model": "gpt-4o-mini",   "tier": "simple"},
    {"model": "gpt-4o",        "tier": "medium"},
    {"model": "claude-sonnet", "tier": "complex"},
    {"model": "claude-bedrock","tier": "complex", "fallback_candidate": true}
  ],
  "retry": {"attempts": 1, "backoff_ms": 200, "on_status_codes": [429, 503]},
  "request_timeout_ms": 60000
}

Key fields

targets[].tier
string
required
simple, medium or complex. A model (including through an alias) can sit in only one tier. Array order is the order within a tier; any number of targets per tier is allowed and tiers may be empty. A tier that has no target sends its requests to the next tier up that has one.
targets[].model
string
required
A concrete chat, completion or responses model, by its gateway name or account/model address. Another routing config cannot be a target, and weight is refused.
classification_strategy
object
Omit for the heuristic classifier. llm-classifier requires model and a fallback_strategy; a static fallback requires default_tier.
Every other target field — when, override_params, override_headers, retry, fallback_status_codes, fallback_candidate — behaves as on Virtual Models.

Classifiers

Deterministic, in-process, no network. It reads only the last user message (first 8,000 characters) and the last system/developer message (first 2,000). Each signal adds or subtracts points:
SignalEffect
Two or more distinct reasoning asks (compare, trade-offs, step by step, why, design…)complex, overriding all else
One reasoning ask+2
Code indicators+3; +4 for a fenced block or 3+ groups
Technical vocabulary+1 per group, up to +3
Prompt over 1,500 / 4,000 characters / truncated+2 / +3 / +4
Multi-step structure (first/then/finally, lists, many requirements)up to +3
More than three question marks+2
Long or technical system context+1
Short greeting / short lookup (“what is”) / light task (“translate”) / very short−3 / −2 / −1 / −1
A score of 1 or less is simple, 2–4 medium, 5 or more complex. Empty input is simple.

Conversation pinning

A conversation can move up a tier and never back down while active.
  • Identity: X-Session-Id when sent; otherwise a fingerprint of the first system and first user message, read only on multi-turn requests (ones with an assistant message).
  • Scope: organization, routing config and caller. The pin holds no prompt text.
  • Rule: effective tier is max(classified, pinned). A conversation pinned at complex skips classification. A pinned target leads its tier when still eligible.
  • Write: only after a successful, uninterrupted response; failed turns leave the pin alone.
  • Lifetime: 10 minutes of inactivity. With Redis the pin is shared across replicas; without it, it lives in process memory.

Target chain

For effective tier T the chain is: T’s targets in order, then higher tiers ascending, then lower tiers descending.
Configuredsimplemediumcomplex
S1 · M1 · C1S1 M1 C1M1 C1 S1C1 M1 S1
S1 · (empty) · C1S1 C1C1 S1C1 S1
Retry repeats the same target on retry.on_status_codes, up to attempts. Fallback moves to the next entry on fallback_status_codes (default 400 401 403 404 408 413 422 429 500 502 503 504 plus retry codes). A target with fallback_candidate: false serves only when its own tier selects it first.

Response headers

X-ManyLayers-Routing-Strategy
string
complexity
X-ManyLayers-Complexity-Tier
string
The effective tier routing started from.
X-ManyLayers-Complexity-Cause
string
heuristic, llm-classifier, pinned, heuristic-fallback or static-fallback.
X-ManyLayers-Routing-Model-Order
string
The full target chain, comma-separated.
X-ManyLayers-Served-Tier
string
The tier of the target that answered.
X-ManyLayers-Target
string
The concrete model that served.
All are suppressed when the config sets strict_openai_compliance.

Traces and cost

Each request trace in Gateway → Request Traces has an Auto Routing section: the classification (tier, method, matched signal names, latency, and classifier usage), the pin as read and written, the chain, every attempt, and filtered targets with a reason (metadata_mismatch, not_allowed, model_restriction, unresolvable, ambiguous, budget_exceeded, rate_limited). Cost is priced from the served model as for any request, plus:
FieldMeaning
cost.actual_usdWhat the served target cost
cost.classifier_usdThe LLM classifier call (0 for heuristic)
cost.top_tier_estimate_usdThe same usage priced at the first eligible target of the highest tier
cost.estimated_savings_usd / _pctTop-tier estimate minus actual
cost.net_savings_usdSavings minus the classifier cost
Savings are omitted when either model has no price.

Failure behaviour

FailureResult
LLM classifier fails or times outfallback_strategy decides; the request continues
Pin store unavailableTurn classified on its own; response unaffected
Every target failsThe gateway’s normalized error; pin untouched
request_timeout_ms passes504; pin untouched
No eligible target404 model_not_found; nothing sent upstream
Embeddings request to a complexity config400

Examples

{"name": "smart-chat", "strategy": "complexity",
 "targets": [{"model": "gpt-4o-mini", "tier": "simple"}, {"model": "claude-sonnet", "tier": "complex"}]}
{"name": "smart-chat", "strategy": "complexity",
 "targets": [
   {"model": "gpt-4o-mini", "tier": "simple", "override_params": {"max_tokens": 512}},
   {"model": "gpt-4o", "tier": "medium", "when": {"metadata": {"region": "us"}}},
   {"model": "gpt-4o-eu", "tier": "medium", "when": {"metadata": {"region": "eu"}}},
   {"model": "claude-sonnet", "tier": "complex"}]}
curl https://app.manylayers.io/v1/chat/completions \
  -H "Authorization: Bearer ml_vat_..." -H "Content-Type: application/json" \
  -H "X-Session-Id: conv-42" -H 'X-ManyLayers-Metadata: {"region":"eu"}' \
  -d '{"model": "smart-chat", "messages": [{"role": "user", "content": "Design a distributed rate limiter"}]}'
Rate limits and budgets are checked once against the requested name; spend is metered against the model that served. A team with an explicit model list must include smart-chat, and each target is checked separately too.

Next steps

Virtual Models

The other routing strategies and the full config schema.

Budgets & cost tracking

How the served model’s cost is computed.