How it works
Every request goes tohttps://app.manylayers.io/v1 and carries a credential in the Authorization header. The gateway authenticates the caller, applies your organization’s policies, picks a provider for the requested model, and returns the answer in the shape your client expects. Console and admin routes live on https://app.manylayers.io, not on the /v1 host.
Connection settings
Key fields
- Base URL: The gateway serves the OpenAI surface under
/v1. The Anthropic SDK adds/v1itself, so give it the URL without the suffix. - Credential: Send an API key (
ml-...), a personal access token (ml_pat_...) or a virtual account token (ml_vat_...) asAuthorization: Bearer <token>. The gateway also accepts OIDC JWTs from your identity provider. It does not readx-api-key. See Authentication and Credentials. - Model: Any name returned by
GET /v1/models. A name your team isn’t allowed to use is refused.
Send a request
If your Anthropic SDK version has no
auth_token option, pass default_headers={"Authorization": "Bearer ml-..."} instead. See Anthropic Messages API for the full translation rules.Stream a response
Set"stream": true to receive Server-Sent Events in the usual OpenAI chunk format.
When an output guardrail applies to a streamed request, the gateway holds the stream until the guardrail finishes and sets
X-ManyLayers-Guardrails-Output: buffered-streaming. You still receive SSE, but the first token arrives later.Common configurations
Call a virtual model or routing config
Call a virtual model or routing config
A routing config with To apply a config to a request for a regular model, name it (by id or name) in a header. An unknown config returns See Routing for strategies and virtual models.
model_types can be called by name as the model. Real model names always win over a config with the same name.400 unknown_config.Attach metadata
Attach metadata
Send a flat JSON object in Limits: 8 KB, 64 keys, values up to 512 bytes. Nested objects and arrays are ignored. A malformed header is treated as no metadata.
X-ManyLayers-Metadata. The gateway stores it on the request trace, and conditional routing and guardrail rules can match on it.Keep a conversation on one upstream
Keep a conversation on one upstream
Send
X-Session-Id (or the OpenAI user field) to make sticky-session and canary routing send a conversation to the same place.List the models you can call
List the models you can call
GET /v1/models returns the models your credential may use. Each entry has the standard OpenAI fields plus provider and source.GET /v1/models/{model} returns one entry, or 404 model_not_found if the model doesn’t exist or you can’t use it.Errors
Every/v1 error uses the OpenAI error envelope, so SDKs raise their usual exception types:
param is included when a specific request field caused the error. /v1/messages returns the Anthropic shape ({"type": "error", "error": {...}}) for errors raised after authentication; see Anthropic Messages API.
| Status | code | Meaning |
|---|---|---|
| 400 | unknown_config | X-ManyLayers-Config names no config |
| 400 | unknown_guardrail | X-ManyLayers-Guardrails names an unknown guardrail |
| 400 | unsupported_parameter | The target provider can’t honour a request parameter; param names it |
| 400 | max_input_tokens_exceeded, max_output_tokens_exceeded, max_total_tokens_exceeded | A policy’s per-request token ceiling was exceeded |
| 401 | invalid_api_key | Missing, unknown or disabled credential |
| 401 | key_expired | The credential’s expiry has passed |
| 402 | budget_exceeded | A spend budget is used up (insufficient_quota type) |
| 403 | model_not_allowed | Your team, virtual account or policy doesn’t allow this model |
| 403 | product_not_entitled | The organization doesn’t hold the AI Gateway product |
| 404 | model_not_found | The model doesn’t exist or you can’t use it |
| 404 | unsupported_endpoint | The gateway doesn’t implement this path |
| 405 | method_not_allowed | Known path, wrong method |
| 429 | rpm_exceeded, tpm_exceeded | A request or token rate limit was hit |
| 429 | budget_exhausted | A token budget is used up (insufficient_quota type) |
| 429 | provider_rate_limited, provider_quota_exceeded | The upstream provider rate limited the request or its account is out of quota |
| 400 | model_ambiguous | More than one provider account serves the model name |
| 502 | provider_unavailable, provider_auth_failed, upstream_malformed_response | The upstream provider failed, rejected the gateway’s credential or returned an unreadable answer |
| 504 | provider_timeout | The upstream provider did not respond in time |
X-ManyLayers-Limit (the level whose limit was hit) and, when waiting helps, Retry-After in seconds. The gateway sends no X-RateLimit-* headers.
For which request parameters each provider honours, see OpenAI Compatibility.
Next steps
Request & Response Headers
Every header the gateway reads and returns.
Endpoints
The full list of supported
/v1 paths.Anthropic Messages API
Use the Anthropic SDK or Claude Code against any provider.
Routing
Fallback, load balancing and virtual models.