"model": "smart-chat"; the gateway classifies the turn, picks the tier, and walks a chain of targets that falls back and escalates on failure.
It is the complexity strategy of a virtual model, so it shares the same admin API, retry policy, per-target overrides and streaming guarantee. In the console, create it under Gateway → Routing → Add Virtual Model and choose Auto Routing. The config belongs to one workspace and serves requests made in it.
How it works
Configuration structure
Key fields
simple, medium or complex. A model (including through an alias) can sit in only one tier. Array order is the order within a tier; any number of targets per tier is allowed and tiers may be empty. A tier that has no target sends its requests to the next tier up that has one.A concrete chat, completion or responses model, by its gateway name or
account/model address. Another routing config cannot be a target, and weight is refused.Omit for the heuristic classifier.
llm-classifier requires model and a fallback_strategy; a static fallback requires default_tier.when, override_params, override_headers, retry, fallback_status_codes, fallback_candidate — behaves as on Virtual Models.
Classifiers
- Heuristic (default)
- LLM classifier
Deterministic, in-process, no network. It reads only the last
A score of 1 or less is simple, 2–4 medium, 5 or more complex. Empty input is simple.
user message (first 8,000 characters) and the last system/developer message (first 2,000). Each signal adds or subtracts points:| Signal | Effect |
|---|---|
| Two or more distinct reasoning asks (compare, trade-offs, step by step, why, design…) | complex, overriding all else |
| One reasoning ask | +2 |
| Code indicators | +3; +4 for a fenced block or 3+ groups |
| Technical vocabulary | +1 per group, up to +3 |
| Prompt over 1,500 / 4,000 characters / truncated | +2 / +3 / +4 |
| Multi-step structure (first/then/finally, lists, many requirements) | up to +3 |
| More than three question marks | +2 |
| Long or technical system context | +1 |
| Short greeting / short lookup (“what is”) / light task (“translate”) / very short | −3 / −2 / −1 / −1 |
Conversation pinning
A conversation can move up a tier and never back down while active.- Identity:
X-Session-Idwhen sent; otherwise a fingerprint of the first system and first user message, read only on multi-turn requests (ones with an assistant message). - Scope: organization, routing config and caller. The pin holds no prompt text.
- Rule: effective tier is
max(classified, pinned). A conversation pinned atcomplexskips classification. A pinned target leads its tier when still eligible. - Write: only after a successful, uninterrupted response; failed turns leave the pin alone.
- Lifetime: 10 minutes of inactivity. With Redis the pin is shared across replicas; without it, it lives in process memory.
Target chain
For effective tier T the chain is: T’s targets in order, then higher tiers ascending, then lower tiers descending.| Configured | simple | medium | complex |
|---|---|---|---|
| S1 · M1 · C1 | S1 M1 C1 | M1 C1 S1 | C1 M1 S1 |
| S1 · (empty) · C1 | S1 C1 | C1 S1 | C1 S1 |
retry.on_status_codes, up to attempts. Fallback moves to the next entry on fallback_status_codes (default 400 401 403 404 408 413 422 429 500 502 503 504 plus retry codes). A target with fallback_candidate: false serves only when its own tier selects it first.
Response headers
complexityThe effective tier routing started from.
heuristic, llm-classifier, pinned, heuristic-fallback or static-fallback.The full target chain, comma-separated.
The tier of the target that answered.
The concrete model that served.
strict_openai_compliance.
Traces and cost
Each request trace in Gateway → Request Traces has an Auto Routing section: the classification (tier, method, matched signal names, latency, and classifier usage), the pin as read and written, the chain, every attempt, and filtered targets with a reason (metadata_mismatch, not_allowed, model_restriction, unresolvable, ambiguous, budget_exceeded, rate_limited).
Cost is priced from the served model as for any request, plus:
| Field | Meaning |
|---|---|
cost.actual_usd | What the served target cost |
cost.classifier_usd | The LLM classifier call (0 for heuristic) |
cost.top_tier_estimate_usd | The same usage priced at the first eligible target of the highest tier |
cost.estimated_savings_usd / _pct | Top-tier estimate minus actual |
cost.net_savings_usd | Savings minus the classifier cost |
Failure behaviour
| Failure | Result |
|---|---|
| LLM classifier fails or times out | fallback_strategy decides; the request continues |
| Pin store unavailable | Turn classified on its own; response unaffected |
| Every target fails | The gateway’s normalized error; pin untouched |
request_timeout_ms passes | 504; pin untouched |
| No eligible target | 404 model_not_found; nothing sent upstream |
| Embeddings request to a complexity config | 400 |
Examples
Two tiers: medium escalates to complex
Two tiers: medium escalates to complex
Regional targets in one tier
Regional targets in one tier
Calling it
Calling it
Rate limits and budgets are checked once against the requested name; spend is metered against the model that served. A team with an explicit model list must include
smart-chat, and each target is checked separately too.Next steps
Virtual Models
The other routing strategies and the full config schema.
Budgets & cost tracking
How the served model’s cost is computed.