Headers let you steer one request without changing your organization’s configuration: pick a routing config, add guardrails, skip the cache, or keep bodies out of the logs. Response headers tell you how the gateway handled the request.

How it works

The gateway reads its own headers, acts on them, and removes the control headers (X-ManyLayers-Guardrails, X-ManyLayers-Store-Logs, X-ManyLayers-Cache-Control, X-ManyLayers-Credential) before forwarding the request, so providers never see them. An unrecognized value in a control header is refused with 400 rather than ignored.
curl https://app.manylayers.io/v1/chat/completions \
  -H "Authorization: Bearer $ML_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-ManyLayers-Config: prod-fallback" \
  -H 'X-ManyLayers-Metadata: {"tenant":"acme"}' \
  -H "X-ManyLayers-Cache-Control: no-cache" \
  -H "X-Session-Id: conv-8f2a" \
  -d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Hi"}]}' -i

Request headers

Authorization
string
required
Bearer <token>. An API key (ml-...), personal access token (ml_pat_...), virtual account token (ml_vat_...) or an OIDC JWT. x-api-key is not read. See Authentication.
X-ManyLayers-Config
string
Id or name of a routing config to apply to this request. Overrides the team’s default config. Unknown names return 400 unknown_config; a config that doesn’t serve the endpoint you called is refused. Ignored when the model is itself a virtual model, which brings its own config.
X-ManyLayers-Metadata
JSON object
Flat key/value pairs, for example {"tenant":"acme","env":"prod"}. Recorded on the request trace and matched by conditional routing targets, Auto Routing conditions and guardrail rules. At most 8 KB and 64 keys; values over 512 bytes, objects and arrays are dropped. A malformed header counts as no metadata.
X-ManyLayers-Guardrails
JSON object
Adds guardrails for this request: {"input_guardrails": ["pii-strict"], "output_guardrails": ["toxicity"]}. It can only add to what your organization’s rules apply, never remove. Up to 32 names and 8 KB. Requires the gateway.guardrails.read permission; unknown names return 400 unknown_guardrail.
X-ManyLayers-Cache-Control
string
no-cache skips the cache lookup; no-store skips writing the answer to the cache. Combine with a comma. Cannot make a request cacheable that your team’s settings exclude.
X-ManyLayers-Store-Logs
string
false keeps this request’s prompt and response bodies out of the audit log. true is accepted and changes nothing: it can’t turn capture on where a logging rule turns it off.
X-Session-Id
string
Sticky-session key. Sticky-session routing and canary/loadbalance configs send requests with the same key to the same target; if the header is absent they use the body’s user field. Auto Routing reads only this header to pin a conversation’s tier.
X-ManyLayers-Model
string
Model for /v1/files and /v1/fine_tuning/jobs calls whose body names no model. You can use ?model= instead. POST /v1/fine_tuning/jobs takes the model in its JSON body, which wins over this header. Missing returns 400 missing_model.
X-ManyLayers-Credential
string
Console sessions only. Id of one of your own personal access tokens; the request is attributed to that token. Sending it with an API key returns 400 credential_header_not_allowed.
X-ManyLayers-Source
string
Free-form label for the calling surface (for example playground). Lowercased, limited to letters, digits, _ and -, 32 characters. Used in logs and traces only; it grants nothing.
X-Request-Id
string
Correlation id. If you send one, the gateway uses it as is in logs and audit records and echoes it back; otherwise it generates a UUID.
traceparent
string
W3C trace context (tracestate is read too). When tracing is configured, the gateway continues your trace instead of starting a new one.

Response headers

X-Request-Id
string
The request’s correlation id, on every response.
X-ManyLayers-Trace-Id
string
Id of the audit record for this request. Open it in the console’s trace view.
X-ManyLayers-Otel-Trace-Id
string
W3C trace id of the distributed trace, present when a tracing destination is configured. Use it to search your own tracing backend.
X-ManyLayers-Cache
string
hit, semantic, miss, bypass (you sent no-cache) or disabled (the request isn’t cache-eligible).
X-ManyLayers-Config
string
Name of the routing config that handled the request.
X-ManyLayers-Target
string
The model that served the response after routing and failover.
X-ManyLayers-Retries
integer
Number of retries spent on the serving target.
X-ManyLayers-Routing-Strategy
string
Strategy of the config that routed the request, for example fallback or complexity.
X-ManyLayers-Routing-Model-Order
string
Comma-separated order in which targets would be tried.
X-ManyLayers-Skipped-Targets
string
Targets passed over without being called, as model=reason pairs. Reasons: metadata_mismatch, not_allowed, model_restriction, unresolvable, ambiguous, budget_exceeded, rate_limited.
X-ManyLayers-Complexity-Tier
string
Auto Routing only. The tier the request was classified as: simple, medium or complex.
X-ManyLayers-Complexity-Cause
string
Auto Routing only. How the tier was chosen, for example heuristic or pinned. A -fallback suffix means the LLM classifier failed and the heuristic decided.
X-ManyLayers-Served-Tier
string
Auto Routing only. The tier of the target that actually answered, which differs from the classified tier after a failover.
X-ManyLayers-Guardrails-Output
string
buffered-streaming when a streamed response was held until output guardrails finished.
X-ManyLayers-Limit
string
On a policy refusal, the level whose limit was hit: organization, workspace, team, service_account, user or api_key. Sent with Retry-After (seconds) when waiting will help. The gateway sends no X-RateLimit-* headers.
X-ManyLayers-Cost
number
Estimated cost in USD for the image and audio endpoints when the model has a price configured.

Common configurations

Read X-ManyLayers-Routing-Model-Order, X-ManyLayers-Skipped-Targets, X-ManyLayers-Target and X-ManyLayers-Retries together. Then open the trace with X-ManyLayers-Trace-Id.
-H "X-ManyLayers-Store-Logs: false" -H "X-ManyLayers-Cache-Control: no-cache, no-store"
The audit row is still written (who, which model, tokens, cost), but without bodies, and the answer is neither read from nor written to the cache.
raw = client.chat.completions.with_raw_response.create(model="gpt-4o-mini", messages=[...])
print(raw.headers.get("X-ManyLayers-Target"), raw.headers.get("X-ManyLayers-Cache"))
completion = raw.parse()
A routing config with strict_openai_compliance: true stops the gateway from adding the routing headers (X-ManyLayers-Config, -Target, -Retries, -Routing-Strategy, -Routing-Model-Order, -Skipped-Targets, -Complexity-*, -Served-Tier) and strips gateway extension fields from non-streaming JSON bodies. X-ManyLayers-Cache is always kept.
Cross-origin browser access is off by default. A deployment enables it only for the console origin it names in PUBLIC_WORKSPACE_URL; for that origin the CORS policy exposes X-Request-Id and the X-ManyLayers-* response headers above, except X-ManyLayers-Otel-Trace-Id and X-ManyLayers-Limit. Browser apps on other origins should call the gateway from their backend.

Next steps

Making Requests

SDK setup, streaming and error handling.

Routing

Configs, strategies and sticky sessions.

Guardrails

Rules that per-request guardrails add to.

Caching

What makes a request cache-eligible.