The gateway speaks the OpenAI API, so most clients work after you change two settings: the base URL and the API key. This page shows how to configure the common SDKs, use virtual models and routing configs, attach metadata, and read errors.

How it works

Every request goes to https://app.manylayers.io/v1 and carries a credential in the Authorization header. The gateway authenticates the caller, applies your organization’s policies, picks a provider for the requested model, and returns the answer in the shape your client expects. Console and admin routes live on https://app.manylayers.io, not on the /v1 host.

Connection settings

base_url: https://app.manylayers.io/v1         # the hosted gateway's data plane
authorization: Bearer ml-...               # API key, ml_pat_... or ml_vat_...
model: gpt-4o-mini                         # a model name returned by GET /v1/models

Key fields

  • Base URL: The gateway serves the OpenAI surface under /v1. The Anthropic SDK adds /v1 itself, so give it the URL without the suffix.
  • Credential: Send an API key (ml-...), a personal access token (ml_pat_...) or a virtual account token (ml_vat_...) as Authorization: Bearer <token>. The gateway also accepts OIDC JWTs from your identity provider. It does not read x-api-key. See Authentication and Credentials.
  • Model: Any name returned by GET /v1/models. A name your team isn’t allowed to use is refused.

Send a request

curl https://app.manylayers.io/v1/chat/completions \
  -H "Authorization: Bearer $ML_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Say hello in French."}]
  }'
If your Anthropic SDK version has no auth_token option, pass default_headers={"Authorization": "Bearer ml-..."} instead. See Anthropic Messages API for the full translation rules.

Stream a response

Set "stream": true to receive Server-Sent Events in the usual OpenAI chunk format.
curl -N https://app.manylayers.io/v1/chat/completions \
  -H "Authorization: Bearer $ML_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4o-mini", "stream": true,
       "messages": [{"role": "user", "content": "Count to five."}]}'
When an output guardrail applies to a streamed request, the gateway holds the stream until the guardrail finishes and sets X-ManyLayers-Guardrails-Output: buffered-streaming. You still receive SSE, but the first token arrives later.

Common configurations

A routing config with model_types can be called by name as the model. Real model names always win over a config with the same name.
client.chat.completions.create(model="support-router", messages=[...])
To apply a config to a request for a regular model, name it (by id or name) in a header. An unknown config returns 400 unknown_config.
client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[...],
    extra_headers={"X-ManyLayers-Config": "prod-fallback"},
)
See Routing for strategies and virtual models.
Send a flat JSON object in X-ManyLayers-Metadata. The gateway stores it on the request trace, and conditional routing and guardrail rules can match on it.
import json
client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[...],
    extra_headers={"X-ManyLayers-Metadata": json.dumps({"tenant": "acme", "env": "prod"})},
)
Limits: 8 KB, 64 keys, values up to 512 bytes. Nested objects and arrays are ignored. A malformed header is treated as no metadata.
Send X-Session-Id (or the OpenAI user field) to make sticky-session and canary routing send a conversation to the same place.
-H "X-Session-Id: conv-8f2a"
GET /v1/models returns the models your credential may use. Each entry has the standard OpenAI fields plus provider and source.
curl https://app.manylayers.io/v1/models -H "Authorization: Bearer $ML_API_KEY"
GET /v1/models/{model} returns one entry, or 404 model_not_found if the model doesn’t exist or you can’t use it.

Errors

Every /v1 error uses the OpenAI error envelope, so SDKs raise their usual exception types:
{
  "error": {
    "message": "team \"web\" is not allowed to use model \"gpt-4o\"",
    "type": "permission_error",
    "code": "model_not_allowed"
  }
}
param is included when a specific request field caused the error. /v1/messages returns the Anthropic shape ({"type": "error", "error": {...}}) for errors raised after authentication; see Anthropic Messages API.
StatuscodeMeaning
400unknown_configX-ManyLayers-Config names no config
400unknown_guardrailX-ManyLayers-Guardrails names an unknown guardrail
400unsupported_parameterThe target provider can’t honour a request parameter; param names it
400max_input_tokens_exceeded, max_output_tokens_exceeded, max_total_tokens_exceededA policy’s per-request token ceiling was exceeded
401invalid_api_keyMissing, unknown or disabled credential
401key_expiredThe credential’s expiry has passed
402budget_exceededA spend budget is used up (insufficient_quota type)
403model_not_allowedYour team, virtual account or policy doesn’t allow this model
403product_not_entitledThe organization doesn’t hold the AI Gateway product
404model_not_foundThe model doesn’t exist or you can’t use it
404unsupported_endpointThe gateway doesn’t implement this path
405method_not_allowedKnown path, wrong method
429rpm_exceeded, tpm_exceededA request or token rate limit was hit
429budget_exhaustedA token budget is used up (insufficient_quota type)
429provider_rate_limited, provider_quota_exceededThe upstream provider rate limited the request or its account is out of quota
400model_ambiguousMore than one provider account serves the model name
502provider_unavailable, provider_auth_failed, upstream_malformed_responseThe upstream provider failed, rejected the gateway’s credential or returned an unreadable answer
504provider_timeoutThe upstream provider did not respond in time
On a policy refusal the response carries X-ManyLayers-Limit (the level whose limit was hit) and, when waiting helps, Retry-After in seconds. The gateway sends no X-RateLimit-* headers. For which request parameters each provider honours, see OpenAI Compatibility.

Next steps

Request & Response Headers

Every header the gateway reads and returns.

Endpoints

The full list of supported /v1 paths.

Anthropic Messages API

Use the Anthropic SDK or Claude Code against any provider.

Routing

Fallback, load balancing and virtual models.