Two ways to add a provider
| Where | Provider types | What it serves | Best for |
|---|---|---|---|
Console: Gateway → Providers (https://app.manylayers.io/ui/gateway/providers) | 34 | /v1/chat/completions, plus /v1/responses and /v1/messages, which are translated to chat. Embedding, rerank and speech models registered on an account are served on /v1/embeddings, /v1/rerank, /v1/audio/speech and /v1/audio/transcriptions. | Workspace teams managing their own accounts and keys |
gateway.yaml models[] | 20 | Every endpoint the provider supports, including /v1/completions, images, moderation, files, fine-tuning and realtime | Operators, and anything the console does not serve |
models[] in gateway.yaml serves keys bound to no workspace and the endpoints above that a console account cannot serve. See Model resolution.
Azure OpenAI can be selected in the console, but a workspace provider row has no place for a deployment name, so requests to it fail with
provider_unavailable. Configure Azure in gateway.yaml. Every other console type, including Bedrock and Vertex, is served from the console.Provider types
provider_type is the id used by the API and by gateway.yaml. The default base URL is what the console fills in when you leave the field empty.
Provider (provider_type) | Default base URL | Credential | In gateway.yaml |
|---|---|---|---|
OpenAI (openai) | https://api.openai.com | API key | yes |
Anthropic (anthropic) | https://api.anthropic.com | API key | yes |
Azure OpenAI (azure) | none, your resource URL | API key | yes (only here) |
Google Gemini (gemini) | https://generativelanguage.googleapis.com | API key | yes |
Google Vertex (vertex) | https://aiplatform.googleapis.com (express mode), or built from project and region | API key, service account, external account, or the gateway’s own Google identity | yes |
AWS Bedrock (bedrock) | built from the account’s region | access key pair, assumed role, or Bedrock API key | yes |
AWS Bedrock Mantle (bedrock_mantle) | https://bedrock-mantle.{region}.api.aws | same AWS methods | console only |
AWS Claude Platform (aws_claude_platform) | https://aws-external-anthropic.{region}.api.aws | same AWS methods, plus a Claude workspace id (wrkspc_…) | console only |
AWS SageMaker (sagemaker) | built from the account’s region | access key pair or assumed role | console only |
Databricks (databricks) | none, your workspace URL | personal access token or service principal | console only |
Cohere (cohere) | https://api.cohere.com | API key | yes |
TypeSafe AI (typesafe) | https://api.typesafe.ai | API key | yes |
Mistral AI (mistral) | https://api.mistral.ai | API key | yes |
Groq (groq) | https://api.groq.com/openai | API key | yes |
DeepSeek (deepseek) | https://api.deepseek.com | API key | yes |
Together AI (together) | https://api.together.xyz | API key | yes |
Perplexity AI (perplexity) | https://api.perplexity.ai | API key | yes |
xAI (xai) | https://api.x.ai | API key | yes |
Fireworks AI (fireworks) | https://api.fireworks.ai/inference | API key | yes |
OpenRouter (openrouter) | https://openrouter.ai/api | API key | yes |
Cerebras (cerebras) | https://api.cerebras.ai | API key | yes |
AI21 Labs (ai21) | https://api.ai21.com/studio | API key | yes |
SambaNova (sambanova) | https://api.sambanova.ai | API key | yes |
DeepInfra (deepinfra) | https://api.deepinfra.com | API key | console only |
Baseten (baseten) | https://inference.baseten.co | API key | console only |
Wafer (wafer) | https://pass.wafer.ai | API key | console only |
Ollama (ollama) | console: https://ollama.com; gateway.yaml: http://localhost:11434 | optional | yes |
Self-Hosted Model (self_hosted) | none, always your server’s URL | optional | no |
Snowflake Cortex (snowflake_cortex) | none; the console builds it from the Snowflake account identifier | programmatic access token | console only |
Cloudera (cloudera) | none; each model has its own endpoint | workload token or CDP access key | console only |
ElevenLabs (elevenlabs) | https://api.elevenlabs.io | API key | console only |
Deepgram (deepgram) | https://api.deepgram.com | API key | console only |
Cartesia (cartesia) | https://api.cartesia.ai | API key | console only |
Smallest AI (smallest) | https://waves-api.smallest.ai | API key | console only |
self_hosted and ollama may be saved without a credential; every other type needs one, except that a Vertex account can use the gateway’s own Google identity and an AWS assumed-role account can be assumed with the gateway’s own AWS identity. A pasted key is checked against the provider’s key format when the account uses the default endpoint (it is skipped when you set your own base URL): OpenAI keys start with sk-, Anthropic sk-ant- (Admin keys are refused), xAI xai-, Cerebras csk-, and Google Gemini accounts accept both the older AIza… keys and the newer AQ.… AI Studio keys.
How requests reach each provider
| Family | Provider types | Behaviour |
|---|---|---|
| OpenAI wire | openai, mistral, groq, deepseek, together, fireworks, openrouter, cerebras, xai, ai21, sambanova, ollama, baseten, wafer, self_hosted, snowflake_cortex, cloudera, bedrock_mantle | The body is forwarded unchanged. Whether the call succeeds depends on whether that upstream implements the endpoint or parameter. deepinfra is the same under /v1/openai. |
| Translated | anthropic, aws_claude_platform, gemini, vertex, bedrock, cohere, perplexity, sagemaker, databricks, typesafe | The gateway rewrites the request and response. See the endpoint table below. |
| Azure | azure | The path becomes /openai/deployments/{upstream_model}/… with an api-version query (default 2024-06-01). |
| Speech only | elevenlabs, deepgram, cartesia, smallest | Serve /v1/audio/speech and /v1/audio/transcriptions only, translated to the provider’s own API. They serve no chat models. |
| Endpoint | Served by |
|---|---|
Chat (/v1/chat/completions) | every type except the speech-only ones |
| Embeddings | OpenAI-wire types, azure, gemini, vertex, bedrock, cohere, databricks |
| Rerank | cohere (translated to Cohere’s v2 rerank); openai, azure and OpenAI-wire types are forwarded to {base}/v1/rerank |
| Text-to-speech and transcription | openai, azure, and the four speech-only types |
| Image generation, edits and variations | gateway.yaml openai and azure models only; the console cannot register image models |
| Files, fine-tuning, realtime | gateway.yaml openai and azure models only |
- A translated provider that cannot serve an endpoint at all (for example
/v1/completionsto Anthropic) is refused by the gateway without an upstream call, and the error names the endpoints it does serve. - Moderation is forwarded to a gateway.yaml
openaiorazuremodel. For any other model, or no model, the gateway answers itself from its PII and prompt-firewall scoring. - Bedrock embeddings take one input per upstream call, so the gateway fans a batch out and reassembles one OpenAI response.
- TypeSafe answers typed questions: every request must set
response_formatto ajson_schema, andstream: trueis refused. - Perplexity is translated to its Agent API; Sonar model names map to Agent API presets.
Parameter support on translated providers
Native-format adapters honour only what they can translate. A parameter that would change the result and cannot be honoured is refused with a400 (code unsupported_parameter) naming it in error.param, never silently dropped. Inert values (n: 1, tool_choice: "auto", parallel_tool_calls: true, a zero penalty) always pass.
| Parameter | anthropic, aws_claude_platform | gemini, vertex | cohere | bedrock | perplexity | typesafe |
|---|---|---|---|---|---|---|
tools, forced tool_choice | yes | yes | yes | yes | — | — |
parallel_tool_calls: false | yes | — | — | — | — | — |
response_format json_object | — | yes | yes | yes | — | — |
response_format json_schema | — | yes | yes | yes | yes | required |
n > 1 | — | yes | — | — | — | — |
seed | — | yes | yes | — | — | — |
logprobs: true | — | yes | yes | — | — | — |
presence_penalty, frequency_penalty | — | yes | yes | — | — | — |
logit_bias | — | — | — | — | — | — |
reasoning_effort | yes | yes | yes | yes | yes | — |
azure, databricks and sagemaker are not checked: the body goes upstream as sent.
Authentication upstream
| Provider | How the gateway authenticates |
|---|---|
openai, cohere, typesafe, databricks, snowflake_cortex, cloudera, smallest and OpenAI-wire types | Authorization: Bearer <key> (omitted when there is no credential) |
azure | api-key: <key> |
anthropic | x-api-key: <key> plus anthropic-version |
gemini | x-goog-api-key: <key> |
vertex | express mode: x-goog-api-key; project mode: an OAuth access token minted from the service account or external account, or from the gateway’s own Google identity when the account has no credential |
bedrock, bedrock_mantle, sagemaker, aws_claude_platform | AWS SigV4 with an access key pair ACCESS_KEY:SECRET_KEY[:SESSION_TOKEN], or the role’s temporary credentials for an assumed role; a Bedrock API key (ABSK… or bedrock-api-key-…) is sent as a bearer token. SageMaker has no API keys. |
elevenlabs | xi-api-key |
deepgram | Authorization: Token <key> |
cartesia | X-API-Key plus a date-based Cartesia-Version the account may pin |
Authorization header is never forwarded.
Adding a provider in the console
Open Gateway → Providers and choose a provider from the gallery (Add other providers once one is connected). The setup has two steps:- Configure Account. Give the account a name (unique in the workspace; a second account of the same provider serves its models as
account/model), a base URL where the type needs one, and a credential: paste the key, or give a reference such as${OPENAI_API_KEY}or${vault:secret/data/llm#openai}(see Credentials). Bedrock-family accounts ask for an AWS region, Vertex for a Google Cloud project and region (leave the project empty for express mode with an API key), and Claude Platform on AWS for the Claude workspace id. Test Credentials checks the account before you save. - Models Selection. Pick the models the account serves from the provider’s own listing, or add one manually with Add Model Manually. See Model selection.
gateway.yaml examples
OpenAI
OpenAI
gateway.yaml
Azure OpenAI
Azure OpenAI
The gateway rewrites the path to
/openai/deployments/{upstream_model}/... and adds api-version (default 2024-06-01, set per upstream with api_version).gateway.yaml
Anthropic
Anthropic
In the console: type Anthropic, base URL left empty, paste the key.
gateway.yaml
Gemini and Vertex AI
Gemini and Vertex AI
In gateway.yaml, Vertex uses express mode: an API key in
x-goog-api-key, calls to {upstream_url}/v1/publishers/google/models/{model}. Project-scoped Vertex accounts (service account, external account or the gateway’s Google identity) are set up in the console.gateway.yaml
AWS Bedrock
AWS Bedrock
Requests use the Converse API and are signed with SigV4.
aws_region is required. The key is ACCESS_KEY:SECRET_KEY[:SESSION_TOKEN], or a Bedrock API key.gateway.yaml
Cohere
Cohere
gateway.yaml
OpenAI-wire presets (Groq, Mistral, DeepSeek, ...)
OpenAI-wire presets (Groq, Mistral, DeepSeek, ...)
Omit
upstream_url and the preset default is used. In the console, choose the type and paste the key.gateway.yaml
Ollama
Ollama
Neither a credential nor, in gateway.yaml, a URL is required for a local Ollama. The console’s Ollama type defaults to
https://ollama.com instead; set the base URL to point it at your own server.gateway.yaml
Self-hosted (vLLM, SGLang, TGI, LM Studio, llama.cpp)
Self-hosted (vLLM, SGLang, TGI, LM Studio, llama.cpp)
In the console,
self_hosted has no default address: the account or each model names its server’s URL, with the server recorded as vLLM, Ollama, SGLang or TGI. A URL ending in /v1 is accepted. In gateway.yaml, serve a self-hosted server as a plain OpenAI-wire model with an explicit URL.gateway.yaml
TypeSafe
TypeSafe
Every request must carry
response_format: {"type": "json_schema", ...}; each schema property is a question. Answers come back as JSON in choices[0].message.content, with probabilities in a typesafe field beside choices.gateway.yaml
Model entry fields (gateway.yaml)
| Field | Description |
|---|---|
logical_name | The name clients send as model and see in /v1/models. Required and unique. |
aliases | Extra names that resolve to this model. |
provider | One of the 20 gateway.yaml types. Default openai. |
upstream_url, upstream_model, upstream_api_key | Single-endpoint shorthand. Keys support ${ENV}, ${ENV:-default} and ${vault:path#key}. Use either this or upstreams[], not both. |
upstreams[] | Several endpoints for one model: url, model, api_key, weight (default 1), provider, api_version (Azure), region (Bedrock). See Routing. |
aws_region | Bedrock region, copied to upstreams that do not set region. |
encoding | Tokenizer for counting. Default cl100k_base. |
input_price_per_1k, output_price_per_1k, cached_input_price_per_1k, cache_write_price_per_1k, media_prices | Pricing; see Budgets & cost tracking. |
pii_output_mode | passthrough (default) or buffer. |
cache | Cache responses even when temperature is not 0. |
Next steps
Model resolution
How a requested name becomes a provider call.
Virtual models
Put fallback, weighting and canaries behind one model name.
Credentials
Gateway keys for callers and how provider credentials are stored.
Budgets & cost tracking
How prices become spend and limits.