How it works
The gateway reads its own headers, acts on them, and removes the control headers (X-ManyLayers-Guardrails, X-ManyLayers-Store-Logs, X-ManyLayers-Cache-Control, X-ManyLayers-Credential) before forwarding the request, so providers never see them. An unrecognized value in a control header is refused with 400 rather than ignored.
Request headers
Bearer <token>. An API key (ml-...), personal access token (ml_pat_...), virtual account token (ml_vat_...) or an OIDC JWT. x-api-key is not read. See Authentication.Id or name of a routing config to apply to this request. Overrides the team’s default config. Unknown names return
400 unknown_config; a config that doesn’t serve the endpoint you called is refused. Ignored when the model is itself a virtual model, which brings its own config.Flat key/value pairs, for example
{"tenant":"acme","env":"prod"}. Recorded on the request trace and matched by conditional routing targets, Auto Routing conditions and guardrail rules. At most 8 KB and 64 keys; values over 512 bytes, objects and arrays are dropped. A malformed header counts as no metadata.Adds guardrails for this request:
{"input_guardrails": ["pii-strict"], "output_guardrails": ["toxicity"]}. It can only add to what your organization’s rules apply, never remove. Up to 32 names and 8 KB. Requires the gateway.guardrails.read permission; unknown names return 400 unknown_guardrail.no-cache skips the cache lookup; no-store skips writing the answer to the cache. Combine with a comma. Cannot make a request cacheable that your team’s settings exclude.false keeps this request’s prompt and response bodies out of the audit log. true is accepted and changes nothing: it can’t turn capture on where a logging rule turns it off.Sticky-session key. Sticky-session routing and canary/loadbalance configs send requests with the same key to the same target; if the header is absent they use the body’s
user field. Auto Routing reads only this header to pin a conversation’s tier.Model for
/v1/files and /v1/fine_tuning/jobs calls whose body names no model. You can use ?model= instead. POST /v1/fine_tuning/jobs takes the model in its JSON body, which wins over this header. Missing returns 400 missing_model.Console sessions only. Id of one of your own personal access tokens; the request is attributed to that token. Sending it with an API key returns
400 credential_header_not_allowed.Free-form label for the calling surface (for example
playground). Lowercased, limited to letters, digits, _ and -, 32 characters. Used in logs and traces only; it grants nothing.Correlation id. If you send one, the gateway uses it as is in logs and audit records and echoes it back; otherwise it generates a UUID.
W3C trace context (
tracestate is read too). When tracing is configured, the gateway continues your trace instead of starting a new one.Response headers
The request’s correlation id, on every response.
Id of the audit record for this request. Open it in the console’s trace view.
W3C trace id of the distributed trace, present when a tracing destination is configured. Use it to search your own tracing backend.
hit, semantic, miss, bypass (you sent no-cache) or disabled (the request isn’t cache-eligible).Name of the routing config that handled the request.
The model that served the response after routing and failover.
Number of retries spent on the serving target.
Strategy of the config that routed the request, for example
fallback or complexity.Comma-separated order in which targets would be tried.
Targets passed over without being called, as
model=reason pairs. Reasons: metadata_mismatch, not_allowed, model_restriction, unresolvable, ambiguous, budget_exceeded, rate_limited.Auto Routing only. The tier the request was classified as:
simple, medium or complex.Auto Routing only. How the tier was chosen, for example
heuristic or pinned. A -fallback suffix means the LLM classifier failed and the heuristic decided.Auto Routing only. The tier of the target that actually answered, which differs from the classified tier after a failover.
buffered-streaming when a streamed response was held until output guardrails finished.On a policy refusal, the level whose limit was hit:
organization, workspace, team, service_account, user or api_key. Sent with Retry-After (seconds) when waiting will help. The gateway sends no X-RateLimit-* headers.Estimated cost in USD for the image and audio endpoints when the model has a price configured.
Common configurations
Debug why a fallback target answered
Debug why a fallback target answered
Read
X-ManyLayers-Routing-Model-Order, X-ManyLayers-Skipped-Targets, X-ManyLayers-Target and X-ManyLayers-Retries together. Then open the trace with X-ManyLayers-Trace-Id.Send a sensitive request
Send a sensitive request
Read headers from the OpenAI SDK
Read headers from the OpenAI SDK
A routing config with
strict_openai_compliance: true stops the gateway from adding the routing headers (X-ManyLayers-Config, -Target, -Retries, -Routing-Strategy, -Routing-Model-Order, -Skipped-Targets, -Complexity-*, -Served-Tier) and strips gateway extension fields from non-streaming JSON bodies. X-ManyLayers-Cache is always kept.Next steps
Making Requests
SDK setup, streaming and error handling.
Routing
Configs, strategies and sticky sessions.
Guardrails
Rules that per-request guardrails add to.
Caching
What makes a request cache-eligible.