A provider account is inert until it has models. Model resolution matches the model field of an incoming request against a registered model row, so an account with a perfect credential and no models answers model_not_found for everything.
You can register models by name, and for some providers you must. Where the provider’s API can enumerate them, ManyLayers asks it instead: the set of models differs per account, names are long and versioned, and a typo is otherwise not discovered until a request fails at the gateway.
In the console this is the second step of Gateway → Providers → Add account (Models Selection), and Add Model on an existing account. The endpoints below are what the console calls.
These are control-plane endpoints on https://app.manylayers.io, not on the /v1 data plane. Authenticate with a session or a personal access token (Authorization: Bearer ml_pat_…); a personal access token is bound to the workspace it was issued for, and a session caller names the workspace with workspace_id when it has more than one. Discovery makes one call upstream per request, and saving a selection clears the resolver’s cache for the workspace. Neither is on the request path.
The flow
POST /admin/gateway/providers create the provider account
POST /admin/gateway/providers/{id}/credentials store its credential
GET /admin/gateway/providers/{id}/models/available ask what it can reach
POST /admin/gateway/providers/{id}/models/test check they answer
PUT /admin/gateway/providers/{id}/models choose what to serve
GET /admin/gateway/providers/{id}/models see what is configured
curl -X POST https://app.manylayers.io/admin/gateway/providers \
-H "Authorization: Bearer ml_pat_..." -H "Content-Type: application/json" \
-d '{"name": "openai-main", "provider_type": "openai"}'
curl -X POST https://app.manylayers.io/admin/gateway/providers/$PROVIDER_ID/credentials \
-H "Authorization: Bearer ml_pat_..." -H "Content-Type: application/json" \
-d '{"secret_reference": "${OPENAI_API_KEY}"}'
A credential takes either api_key (a pasted key, sealed at rest) or secret_reference, never both. See Credentials. Single models can also be added with POST /admin/gateway/providers/{id}/models.
What this gateway can talk to
GET /admin/gateway/provider-types
Lists the provider types this build can register, whether each can be asked for its models, its default endpoint, and the models the console offers before a credential exists. The console’s provider and model choosers both read it, so they cannot drift from what the API will accept. Types that are addressed differently also carry requires_region (AWS accounts), requires_project (Vertex), requires_upstream_workspace (Claude Platform on AWS) and api_versions (Cartesia). A type with no default_base_url must be given one.
{
"provider_types": [
{
"provider_type": "openai",
"discovery": true,
"default_base_url": "https://api.openai.com",
"models": [
{ "model_name": "gpt-5", "display_name": "GPT-5",
"input_price_per_1k": 0.00125, "output_price_per_1k": 0.01 }
]
}
]
}
The models here are suggestions, not entitlements. Only the provider knows what a given account may call; discovery is what asks it. See Providers for the types.
Discovering models
GET /admin/gateway/providers/{id}/models/available
Resolves the provider’s stored credential and makes one authenticated call to its model listing, through the same credential injector the request path uses, so a listing that succeeds proves the credential real traffic would present is accepted. It needs gateway.models.read and times out after 15 seconds.
{
"provider_id": "…",
"provider_type": "openai",
"models": [
{ "model_name": "gpt-5", "model_type": "chat", "configured": true,
"enabled": true, "gateway_model_name": "gpt-5",
"input_price_per_1k": 0.00125, "output_price_per_1k": 0.01 },
{ "model_name": "gpt-5-mini", "model_type": "chat", "configured": false,
"enabled": false, "gateway_model_name": "gpt-5-mini" }
]
}
configured and enabled describe what this account already does, so a console renders its checkboxes in the state the gateway is actually in. A price is filled in from the provider’s own listing when it gives one, otherwise from the published list price when the gateway knows the model.
Which providers can be asked
| Provider | Listing |
|---|
openai and the OpenAI-wire types (groq, deepseek, together, perplexity, xai, fireworks, openrouter, cerebras, sambanova, mistral, ollama, baseten, self_hosted, …) | GET /v1/models |
deepinfra | GET /v1/openai/models |
anthropic, aws_claude_platform | GET /v1/models, paginated with after_id |
gemini | GET /v1beta/models, paginated with pageToken |
cohere | GET /v1/models, paginated with page_token |
databricks | GET /api/2.0/serving-endpoints |
bedrock | AWS control-plane model listing (SigV4) |
sagemaker | the account’s SageMaker endpoints |
azure, vertex, typesafe, ai21, snowflake_cortex, wafer | not supported: 501 discovery_unsupported |
Azure, Vertex and the others above have no listing the gateway can call with the account’s credential: an Azure deployment name is chosen by whoever created it, Vertex enumerates per project and location, and the rest document registering models by ID. For these the console offers its suggested models, and you register anything else by name. A self_hosted account with no URL of its own has no one server to ask (400 missing_base_url); give each model its server’s URL.
Failure modes
| Condition | Status | Code |
|---|
| Provider has no active credential | 400 | missing_credential |
| Credential reference does not resolve, or ciphertext will not open | 400 | invalid_credential |
| Provider rejected the credential (401/403 upstream) | 400 | invalid_credential |
| Provider type cannot enumerate models | 501 | discovery_unsupported |
| Provider unreachable, or answered an error | 502 | discovery_failed |
| Provider is in another workspace, or does not exist | 404 | not_found |
The resolved secret never appears in a response, an error or a log line.
Testing a selection
POST /admin/gateway/providers/{id}/models/test
Discovery says what an account lists. This says what it serves, by sending one real, minimal request per model through the same adapter, the same credential injector and the same base URL live traffic uses. It needs gateway.models.manage, since it spends the account’s credential. In the console it is the Test Connection button.
The two are not the same fact, and the gap between them is what nothing else catches: a listing carries models the organization is not entitled to, models retired behind the listing, and models whose name is right for a different account. Each of those is a 404 at request time, hours after someone saved the selection.
{
"models": [
{ "model_name": "gpt-5", "model_type": "chat" },
{ "model_name": "text-embedding-3-small", "model_type": "embedding" },
{ "model_name": "gpt-4o-transcribe", "model_type": "audio_transcription" }
]
}
{
"provider_id": "…",
"provider_type": "openai",
"results": [
{ "model_name": "gpt-5", "status": "passed", "stage": "ok",
"message": "api.openai.com answered HTTP 200 in 412ms", "latency_ms": 412 },
{ "model_name": "text-embedding-3-small", "status": "failed", "stage": "model",
"message": "api.openai.com does not serve \"text-embedding-3-small\" on this account (HTTP 404)" },
{ "model_name": "gpt-4o-transcribe", "status": "skipped", "stage": "unsupported",
"message": "audio transcription models are not called on a path the gateway can probe" }
],
"passed": 1, "failed": 1, "skipped": 1
}
It writes nothing. A model that failed may be one you are about to fix by correcting its name, and unregistering it under you is not this endpoint’s call. Save the selection separately, with PUT.
The three outcomes
skipped is a state of its own rather than a failure, because “we did not ask” and “we asked and it said no” have opposite fixes. A model is skipped when:
- its type is not one the gateway calls on a modelled path: speech and image models on an account that is not a speech-only provider;
- the provider cannot serve that endpoint (for example an embedding model on a chat-only provider);
- the account is Azure, which calls a deployment rather than a model name.
Chat, completion, embedding and rerank models are probed on /v1/chat/completions, /v1/completions, /v1/embeddings and /v1/rerank respectively. A model with no model_type is treated as chat, the same reading the model rows use. Speech-only providers (ElevenLabs, Deepgram, Cartesia, Smallest AI) are checked against a free listing or voice lookup rather than by synthesizing or transcribing anything. Bedrock and Vertex accounts are probed like the rest: the account’s region and project are part of the address.
Stages
stage is the coarse reason behind a failed, so identical failures group into one thing to fix rather than N rows to read:
| Stage | What it means |
|---|
authentication | the provider refused the credential (401/403) |
model | the credential works; this account cannot call that model |
quota | accepted and refused (429): a quota or rate limit |
reachability | the endpoint could not be reached at all |
upstream | the provider answered 5xx |
request | the provider rejected the request itself |
endpoint | the model’s own URL (Cloudera, a self-hosted server) is unusable |
Each entry is a billed request against your own account, so a batch is capped at 50 models (too_many_models), at most five run at once, each is bounded at 20 seconds, and the whole batch at 90. The console probes a longer selection in successive batches.
A reasoning model that rejects max_tokens by name is retried once with max_completion_tokens, so a working model is not reported as broken over the probe’s own request shape.
Failure modes
| Condition | Status | Code |
|---|
models empty or absent | 400 | missing_models |
| More than 50 models in one body | 400 | too_many_models |
| Empty name, whitespace, control characters, or over 256 bytes | 400 | invalid_model_name |
| Provider has no active credential | 400 | missing_credential |
| Credential reference does not resolve | 400 | invalid_credential |
| Provider is in another workspace, or does not exist | 404 | not_found |
The resolved secret never appears in a response, an error or a log line, and a provider that quotes the key back in its own error message is not passed through verbatim.
Choosing models
PUT /admin/gateway/providers/{id}/models
Applies a whole selection in one transaction: a named model that is not registered is created, and one that is has its enabled flag, display name and (for self-hosted and Cloudera models) URL and server brought into line. It needs gateway.models.manage.
curl -X PUT https://app.manylayers.io/admin/gateway/providers/$PROVIDER_ID/models \
-H "Authorization: Bearer ml_pat_..." -H "Content-Type: application/json" \
-d '{"models": [
{"model_name": "gpt-5", "enabled": true},
{"model_name": "gpt-5-mini", "enabled": true},
{"model_name": "gpt-4.1", "enabled": false}
]}'
provider_model_name is accepted as a synonym for model_name. gateway_model_name, display_name, model_type, context_window, input_price_per_1k and output_price_per_1k are optional. model_type is chat (the default), embedding or rerank, plus audio_speech and audio_transcription on OpenAI, Azure and the speech-only providers. enabled defaults to true: selecting a model means serving it.
{
"models": [ /* the provider's models after the change */ ],
"created": 2,
"updated": 1,
"unchanged": 0
}
It adds and toggles; it never removes. A model the request does not mention keeps whatever state it has. A model listing is a page of what a provider offers, so treating an unmentioned model as “delete it” would let a filtered or paginated view silently unregister models that are serving traffic, and would break routing configs that target them. Turning a model off is what "enabled": false is for. Removing one is DELETE /admin/gateway/providers/{id}/models/{modelID}, which answers 204, or 409 model_in_use while a routing config still targets the model.
Pricing
A model registered with no price is costed at zero for ever, which reads downstream as free rather than as unpriced. An omitted price (both 0) therefore falls back to the model’s published list price when the gateway knows it, written into the row so the number is visible and editable. Pass an explicit price only if your rate differs, and 0 only if the model really is free. To change prices later, PATCH /admin/gateway/providers/{id}/models/{modelID} takes input_price_per_1k and output_price_per_1k together.
Validation
| Condition | Status | Code |
|---|
models empty or absent | 400 | missing_models |
| More than 500 models in one body | 400 | too_many_models |
| Empty name, whitespace, control characters, or over 256 bytes | 400 | invalid_model_name |
| Negative price | 400 | invalid_price |
A gateway name that starts with a provider account’s name and a / (the gateway adds that prefix itself) | 400 | reserved_model_name |
| Same model named twice in one body | 409 | model_exists |
| A gateway name that already points at a different upstream model | 409 | model_exists |
| An attempt to rename a registered model’s gateway name | 409 | model_exists |
A rejected selection writes nothing. A gateway name is a name callers have hardcoded, so this endpoint will not move one: register the new name and retire the old one, or use PATCH /admin/gateway/providers/{id}/models/{modelID} with provider_model_name to repoint the existing name at a different upstream model.
After saving
The resolver caches “what does this name resolve to, in this workspace”, credential and all, for thirty seconds. Saving a selection drops that workspace’s cache, so a model you have just enabled serves immediately and one you have just disabled stops immediately, rather than somewhere in the next half minute.
A request for a model that is disabled, or was never registered, fails at the gateway with model_not_found and is never sent upstream. The two are deliberately indistinguishable to the caller.
Isolation and permissions
Every route here is scoped to the caller’s organization and workspace, in the SQL rather than by a comparison afterwards. A provider in another workspace is 404, not 403: a distinguishable “forbidden” would confirm that a guessed id names something real.
Permissions: gateway.models.read to list and discover, gateway.models.manage to select, test and delete; gateway.providers.manage to create accounts and credentials. An account can also be restricted to named users, teams or virtual accounts under Access Control; see Model resolution.