The request path
Compatibility surface
Applications written against an OpenAI client can point their base URL athttps://app.manylayers.io/v1. What that covers, exactly:
| Endpoint | Status |
|---|---|
POST /v1/chat/completions, stream: false | Supported, including request fields the gateway does not itself interpret. Models come from the caller’s workspace. |
POST /v1/chat/completions, stream: true | Supported. Chunks are relayed as they arrive, and a client disconnect cancels the provider request. |
POST /v1/responses, POST /v1/messages | Supported. The request is translated to chat and runs through the same pipeline (access, policy, guardrails, routing, metering), then the answer is translated back, streams included. |
GET /v1/models, GET /v1/models/{model} | Supported. Lists what would actually resolve for the caller. |
POST /v1/embeddings, POST /v1/rerank | Resolved in the caller’s workspace first (a model registered with type embedding or rerank), then in gateway.yaml. |
POST /v1/audio/speech, POST /v1/audio/transcriptions | Resolved in the workspace first (models of type audio_speech or audio_transcription), then in gateway.yaml. |
POST /v1/completions | gateway.yaml models, and virtual models that serve completion. |
| Images, moderations, files, fine-tuning, realtime | gateway.yaml models; see Providers. |
embedding on /v1/chat/completions fails as 404 model_not_found, with a message saying it is an embedding model and not a chat model.
Two sources of models
A caller resolves models from exactly one source:- The caller’s workspace. Provider accounts, credentials and models the customer registered themselves in Gateway → Providers. This keeps provider accounts owned by the customer. A token bound to a workspace (a personal access token, a virtual account token, or a workspace-bound API key) resolves chat models here and nowhere else: a name only
gateway.yamldeclares ismodel_not_foundfor it. - The configured catalog. The
models:block ofgateway.yaml, which the operator controls. It serves tokens bound to no workspace (team-scoped keys), and the endpoints a workspace account cannot serve (/v1/completions, images, moderations, files, fine-tuning, realtime). The embeddings, rerank and speech endpoints fall back to it when the workspace has no model of that name.
GET /v1/models lists the same single source a chat request resolves from: what you see listed is what you get. It lists real models only; a virtual model is called by name but is not listed.
Resolution is scoped by workspace: one workspace’s models are never reachable from another’s token, whatever the model is called. Access then applies on top: a virtual account only reaches the models and provider accounts it was given, a team with an explicit model list must include the name, and a provider account can be restricted to named users, teams or virtual accounts.
When the name is not a model
If the name resolves to no model, the gateway checks whether it is a virtual model in the caller’s workspace: a routing config withmodel_types set, or an Auto Routing config. Real models are looked up first and always win, so creating a virtual model can never change what an existing model name serves. The API refuses a virtual model whose name is already a model. If neither matches, the answer is 404 model_not_found.
Separately, the X-ManyLayers-Config header (a config id or name) or a team’s default config can wrap a request for a real model in a routing config. See Virtual models.
Registering a provider
Providers are registered in the console: Gateway → Providers, choose a provider, name the account, attach a credential, then select its models. See Providers and Model selection. The same steps are available through/admin/gateway/providers on https://app.manylayers.io.
For each model the workspace records the gateway name callers use, the provider model name sent upstream, its type (chat, embedding, rerank, audio_speech or audio_transcription), whether it is enabled, and its price. Every provider type is served from the workspace except Azure OpenAI, which needs a deployment path a workspace account does not carry and is configured in gateway.yaml. A model registered with another type, such as image generation, does not resolve.
Resolutions are cached for 30 seconds and dropped immediately when a provider, credential or model changes, so a credential rotation or a newly enabled model takes effect at once. A disabled model, or one never registered, is model_not_found.
Several accounts serving one model
A workspace can register the same model on more than one provider account, such asopenai-main and openai-dev both serving gpt-4o. Every workspace model therefore has two names:
| Name | Means |
|---|---|
gpt-4o | the gateway model name, shared by every account serving it |
openai-dev/gpt-4o | its address: that one account’s model, and nothing else |
- An address always resolves to exactly that account. It is never re-pointed: if the caller may not use the account, the request is refused with
403 provider_access_denied, so typing the name grants nothing. - A bare name resolves among the accounts this caller may use. One account: that one. Several:
400 model_ambiguous, naming the addresses to choose from (only ones the caller could use). The gateway never picks one on the caller’s behalf, because the accounts can differ in billing, region and data handling.
GET /v1/models lists every model the caller may use by its address, and by its bare name too when exactly one of their accounts serves it.
Access and naming:
- Provider account access (users, teams, service accounts; per-account model and API-key scopes) decides which accounts a caller may use. It is managed under Access Control on each account.
- Virtual account model allowlists, team model policies and model restriction policies accept either spelling.
gpt-4ocovers every account’s copy;openai-dev/gpt-4ocovers only that account’s. - Rate limits and budgets accept either spelling too. On
gpt-4othey count every account’s traffic together; onopenai-dev/gpt-4othey count only that account’s. Guardrail model conditions take the bare name only.
/, and a model’s gateway name may not start with an account’s name followed by /. The inference log line carries the serving account as provider_account.
Credentials
Three kinds of credential are involved and none of them is interchangeable:| Credential | Held by | Where it lives |
|---|---|---|
Gateway token (ml_pat_…, ml_vat_…, ml-…) | the client | api_keys.key_hash, a SHA-256 hash, never the token |
| Provider credential | the customer | a ${ENV_VAR} or ${vault:secret/data/path#key} reference in the database, or the pasted key sealed with AES-256-GCM |
| Session token | a browser user | a session cookie issued at sign-in |
Authorization header is dropped before the upstream request is built, and the provider’s own credential is set in its place. The client’s token never reaches a provider, and no provider credential is ever returned by an API or written to a log line. See Credentials.
Error semantics
A provider’s error is never relayed to the client. Each one is classified into a gateway error, which decides both the HTTP status and the wording; the provider’s own status and body are logged as the cause and go no further.| Situation | Client sees | Code |
|---|---|---|
| Missing, invalid or revoked gateway token | 401 | invalid_api_key |
| Expired API key | 401 | key_expired |
| Malformed body, missing model | 400 | bad_json, missing_model |
| Model not allowed for the virtual account or team | 403 | model_not_allowed |
| Provider account the caller may not use | 403 | provider_access_denied |
| Unknown, disabled or unavailable model | 404 | model_not_found |
| Bare model name served by several accounts the caller may use | 400 | model_ambiguous |
| Parameter the provider’s adapter cannot honour | 400 | unsupported_parameter |
| Provider rejected the gateway’s credential | 502 | provider_auth_failed |
| Provider rate limited the request | 429 | provider_rate_limited |
| Provider account out of quota or credit | 429 | provider_quota_exceeded |
| Provider timed out | 504 | provider_timeout |
| Provider unreachable, failing, overloaded, or does not serve the model | 502 | provider_unavailable |
| Provider response could not be translated | 502 | upstream_malformed_response |
| Provider rejected the request itself | 400 | invalid_request |
- A provider rejecting the gateway’s credential surfaces as
502, not401. The caller’s token was fine; telling them otherwise sends them to fix something that is not broken. - A provider’s
404for a model the gateway resolved is502 provider_unavailable: it is a misconfigured account or model name, not a bad client request. - Rate limiting keeps its
429whoever’s limit was hit, because backing off is the caller’s response either way.
Usage and observability
Every request is assigned a correlation id, returned asX-Request-Id and honoured if the caller supplies one. The same id appears in the access log, the inference log line and the usage record. A request’s trace id is returned as X-ManyLayers-Trace-Id.
Each request writes one usage record: requested model, resolved model, serving provider account, token counts, cost, latency and status, so the same figures drive Budgets & cost tracking and Analytics. Prompts and completions are not stored there. Request and response bodies appear in Request Traces only when the deployment’s audit.log_bodies or the workspace’s logging rules keep them, and a caller can opt a single request out with X-ManyLayers-Store-Logs: false.
The gateway also emits one structured log line per inference request. It never logs the Authorization header, tokens, provider credentials, prompts or completions.