ManyLayers is a gateway, not an inference provider. You bring your own provider accounts and credentials, and pay those providers directly. The gateway sits in front of them and owns everything between the client and the provider call: authentication, validation, access checks, model resolution, provider translation, error normalization, usage recording and observability.

The request path

Client
  ↓  POST https://app.manylayers.io/v1/chat/completions
  ↓  Authorization: Bearer <gateway token>
Authentication
  ↓
Validation                       400 on a malformed body or parameter
  ↓
Access checks                    403 model_not_allowed
  ↓  requested model → provider account → provider model
Model resolution                 404 model_not_found
  ↓
Provider account access          403 provider_access_denied
  ↓
Policy: rate limits, budgets
  ↓
Safety: firewall, PII, guardrails, cache
  ↓
Routing and provider call        the provider's own credential, never the client's
  ↓
Response normalization
  ↓
Usage recording
  ↓
Client

Compatibility surface

Applications written against an OpenAI client can point their base URL at https://app.manylayers.io/v1. What that covers, exactly:
EndpointStatus
POST /v1/chat/completions, stream: falseSupported, including request fields the gateway does not itself interpret. Models come from the caller’s workspace.
POST /v1/chat/completions, stream: trueSupported. Chunks are relayed as they arrive, and a client disconnect cancels the provider request.
POST /v1/responses, POST /v1/messagesSupported. The request is translated to chat and runs through the same pipeline (access, policy, guardrails, routing, metering), then the answer is translated back, streams included.
GET /v1/models, GET /v1/models/{model}Supported. Lists what would actually resolve for the caller.
POST /v1/embeddings, POST /v1/rerankResolved in the caller’s workspace first (a model registered with type embedding or rerank), then in gateway.yaml.
POST /v1/audio/speech, POST /v1/audio/transcriptionsResolved in the workspace first (models of type audio_speech or audio_transcription), then in gateway.yaml.
POST /v1/completionsgateway.yaml models, and virtual models that serve completion.
Images, moderations, files, fine-tuning, realtimegateway.yaml models; see Providers.
Response bodies for supported endpoints are OpenAI-shaped whichever provider served them. Error bodies use the OpenAI error envelope but are the gateway’s errors, not the provider’s; see Error semantics. A workspace model is only served on the endpoints its type names. Calling a model registered as embedding on /v1/chat/completions fails as 404 model_not_found, with a message saying it is an embedding model and not a chat model.

Two sources of models

A caller resolves models from exactly one source:
  1. The caller’s workspace. Provider accounts, credentials and models the customer registered themselves in Gateway → Providers. This keeps provider accounts owned by the customer. A token bound to a workspace (a personal access token, a virtual account token, or a workspace-bound API key) resolves chat models here and nowhere else: a name only gateway.yaml declares is model_not_found for it.
  2. The configured catalog. The models: block of gateway.yaml, which the operator controls. It serves tokens bound to no workspace (team-scoped keys), and the endpoints a workspace account cannot serve (/v1/completions, images, moderations, files, fine-tuning, realtime). The embeddings, rerank and speech endpoints fall back to it when the workspace has no model of that name.
GET /v1/models lists the same single source a chat request resolves from: what you see listed is what you get. It lists real models only; a virtual model is called by name but is not listed. Resolution is scoped by workspace: one workspace’s models are never reachable from another’s token, whatever the model is called. Access then applies on top: a virtual account only reaches the models and provider accounts it was given, a team with an explicit model list must include the name, and a provider account can be restricted to named users, teams or virtual accounts.

When the name is not a model

If the name resolves to no model, the gateway checks whether it is a virtual model in the caller’s workspace: a routing config with model_types set, or an Auto Routing config. Real models are looked up first and always win, so creating a virtual model can never change what an existing model name serves. The API refuses a virtual model whose name is already a model. If neither matches, the answer is 404 model_not_found. Separately, the X-ManyLayers-Config header (a config id or name) or a team’s default config can wrap a request for a real model in a routing config. See Virtual models.

Registering a provider

Providers are registered in the console: Gateway → Providers, choose a provider, name the account, attach a credential, then select its models. See Providers and Model selection. The same steps are available through /admin/gateway/providers on https://app.manylayers.io. For each model the workspace records the gateway name callers use, the provider model name sent upstream, its type (chat, embedding, rerank, audio_speech or audio_transcription), whether it is enabled, and its price. Every provider type is served from the workspace except Azure OpenAI, which needs a deployment path a workspace account does not carry and is configured in gateway.yaml. A model registered with another type, such as image generation, does not resolve. Resolutions are cached for 30 seconds and dropped immediately when a provider, credential or model changes, so a credential rotation or a newly enabled model takes effect at once. A disabled model, or one never registered, is model_not_found.

Several accounts serving one model

A workspace can register the same model on more than one provider account, such as openai-main and openai-dev both serving gpt-4o. Every workspace model therefore has two names:
NameMeans
gpt-4othe gateway model name, shared by every account serving it
openai-dev/gpt-4oits address: that one account’s model, and nothing else
A request may use either:
  • An address always resolves to exactly that account. It is never re-pointed: if the caller may not use the account, the request is refused with 403 provider_access_denied, so typing the name grants nothing.
  • A bare name resolves among the accounts this caller may use. One account: that one. Several: 400 model_ambiguous, naming the addresses to choose from (only ones the caller could use). The gateway never picks one on the caller’s behalf, because the accounts can differ in billing, region and data handling.
GET /v1/models lists every model the caller may use by its address, and by its bare name too when exactly one of their accounts serves it. Access and naming:
  • Provider account access (users, teams, service accounts; per-account model and API-key scopes) decides which accounts a caller may use. It is managed under Access Control on each account.
  • Virtual account model allowlists, team model policies and model restriction policies accept either spelling. gpt-4o covers every account’s copy; openai-dev/gpt-4o covers only that account’s.
  • Rate limits and budgets accept either spelling too. On gpt-4o they count every account’s traffic together; on openai-dev/gpt-4o they count only that account’s. Guardrail model conditions take the bare name only.
Account names may not contain /, and a model’s gateway name may not start with an account’s name followed by /. The inference log line carries the serving account as provider_account.

Credentials

Three kinds of credential are involved and none of them is interchangeable:
CredentialHeld byWhere it lives
Gateway token (ml_pat_…, ml_vat_…, ml-…)the clientapi_keys.key_hash, a SHA-256 hash, never the token
Provider credentialthe customera ${ENV_VAR} or ${vault:secret/data/path#key} reference in the database, or the pasted key sealed with AES-256-GCM
Session tokena browser usera session cookie issued at sign-in
The client’s Authorization header is dropped before the upstream request is built, and the provider’s own credential is set in its place. The client’s token never reaches a provider, and no provider credential is ever returned by an API or written to a log line. See Credentials.

Error semantics

A provider’s error is never relayed to the client. Each one is classified into a gateway error, which decides both the HTTP status and the wording; the provider’s own status and body are logged as the cause and go no further.
SituationClient seesCode
Missing, invalid or revoked gateway token401invalid_api_key
Expired API key401key_expired
Malformed body, missing model400bad_json, missing_model
Model not allowed for the virtual account or team403model_not_allowed
Provider account the caller may not use403provider_access_denied
Unknown, disabled or unavailable model404model_not_found
Bare model name served by several accounts the caller may use400model_ambiguous
Parameter the provider’s adapter cannot honour400unsupported_parameter
Provider rejected the gateway’s credential502provider_auth_failed
Provider rate limited the request429provider_rate_limited
Provider account out of quota or credit429provider_quota_exceeded
Provider timed out504provider_timeout
Provider unreachable, failing, overloaded, or does not serve the model502provider_unavailable
Provider response could not be translated502upstream_malformed_response
Provider rejected the request itself400invalid_request
Three of these deserve their reasoning stated:
  • A provider rejecting the gateway’s credential surfaces as 502, not 401. The caller’s token was fine; telling them otherwise sends them to fix something that is not broken.
  • A provider’s 404 for a model the gateway resolved is 502 provider_unavailable: it is a misconfigured account or model name, not a bad client request.
  • Rate limiting keeps its 429 whoever’s limit was hit, because backing off is the caller’s response either way.

Usage and observability

Every request is assigned a correlation id, returned as X-Request-Id and honoured if the caller supplies one. The same id appears in the access log, the inference log line and the usage record. A request’s trace id is returned as X-ManyLayers-Trace-Id. Each request writes one usage record: requested model, resolved model, serving provider account, token counts, cost, latency and status, so the same figures drive Budgets & cost tracking and Analytics. Prompts and completions are not stored there. Request and response bodies appear in Request Traces only when the deployment’s audit.log_bodies or the workspace’s logging rules keep them, and a caller can opt a single request out with X-ManyLayers-Store-Logs: false. The gateway also emits one structured log line per inference request. It never logs the Authorization header, tokens, provider credentials, prompts or completions.