OpenAI
Anthropic
The gateway translates OpenAI chat completions to Anthropic’s/v1/messages format automatically — your clients send standard OpenAI requests.
Azure OpenAI
AWS Bedrock
From the console
In Models → Add account → AWS Bedrock, each account brings its own AWS credentials:-
AWS region: where the models are called, such as
us-east-1. Requests are signed for it, and the endpoint isbedrock-runtime.<region>.amazonaws.com. Set a Base URL under Show advanced fields only for a VPC endpoint. -
AWS Account Auth Data: choose one method.
-
Access Key Based: an IAM user’s access key ID and secret. A session token field appears for temporary (
ASIA…) keys. -
Assumed Role Based: the ARN of a role in your account, plus an optional external ID. The gateway assumes the role with its own AWS identity: its environment variables, EKS service-account role, ECS task role or EC2 instance role, found in that order. The role’s trust policy must allow that identity:
Use an external ID. Without one, anyone who can add accounts on the same gateway and knows your role ARN could assume it.
-
API Key Based: a key from the Bedrock console’s API keys page (
ABSK…), sent as a bearer token.
${ENV_VAR}or${vault:path#key}). -
Access Key Based: an IAM user’s access key ID and secret. A session token field appears for temporary (
List* actions are optional; they let the console list your models for you. Without them, add models by ID.
us.anthropic.claude-sonnet-4-5-20250929-v1:0, global.anthropic.…) rather than the bare model ID. A cross-region profile routes to other regions, so the policy must allow them too, which the wildcards above do. An application inference profile is used by its full ARN.
In gateway.yaml
Pass credentials asACCESS_KEY:SECRET_KEY (or append :SESSION_TOKEN for STS temporary credentials), or pass a Bedrock API key as is:
AI21 Labs (Jamba)
In Models → Add account → AI21 Labs, enter an AI21 Studio API key. Jamba speaks the OpenAI format athttps://api.ai21.com/studio/v1/chat/completions, so requests pass through unchanged: streaming, tools and response_format: {"type": "json_object"}. AI21 has no model listing, so the console offers jamba-large and jamba-mini; add any other Jamba ID by name.
Cohere
In Models → Add account → Cohere, enter a Cohere API key. The models list comes from the key, and each model is typed by what it serves:- Chat (
command-a-03-2025, …) on/v1/chat/completions, with tools and structured output.command-a-vision-07-2025accepts images;command-a-reasoning-08-2025takesreasoning_effortand returnsreasoning_content. - Embeddings (
embed-v4.0) on/v1/embeddings. Passinput_type:search_queryfor queries,search_document(the default) for what they search.dimensionsandtruncateare passed through. - Rerank (
rerank-v3.5) on/v1/rerank:{"model", "query", "documents", "top_n"}in,{"results": [{"index", "relevance_score"}]}out.
gateway.yaml is used, as before.
Google Gemini (AI Studio)
In Models → Add account → Google Gemini, enter an AI Studio API key (AIza… or the newer AQ.…). The models list comes from the key itself.
Chat supports:
- Media: images, audio, video and PDFs, as base64 data URIs or
https://URLs.gs://objects need a Vertex account. - Reasoning:
reasoning_effortmaps to Gemini’s thinking budget, and the reasoning comes back asreasoning_content. - Built-in tools: pass
{"type": "google_search"},{"type": "url_context"}or{"type": "code_execution"}intools. Search sources come back asurl_citationannotations, and executed code and its output appear in the answer as code blocks. - Embeddings:
task_type(RETRIEVAL_QUERY,RETRIEVAL_DOCUMENT, …),titleanddimensions, for example withgemini-embedding-001.
Google Vertex AI
In Models → Add account → Google Vertex:- Project ID and Region: the account is called at
https://{region}-aiplatform.googleapis.com/v1/projects/{project}/locations/{region}. Thegloballocation usesaiplatform.googleapis.com. Set a Base URL under Show advanced fields only for Private Service Connect. - Google Cloud Auth Data: choose one method.
- Service Account Key: paste the JSON key of a service account with Vertex AI User (
roles/aiplatform.user). Afterwards the console shows only the service account’s email. - Workload Identity Federation: paste the
credential-config.jsonfromgcloud iam workload-identity-pools create-cred-config. Its token file must be a mounted Kubernetes token under/var/run/secrets/; setGCP_WIF_FILE_PREFIXESon the gateway to allow another directory. The token endpoints must be Google’s own. - API Key: a Vertex express-mode key. It reaches Google’s Gemini models only, and is sent without the project.
- Switched off: the gateway’s own Google identity (GKE Workload Identity, a GCE service account, or
GOOGLE_APPLICATION_CREDENTIALS).
- Service Account Key: paste the JSON key of a service account with Vertex AI User (
| Publisher | Model ID | Sent as |
|---|---|---|
google/gemini-2.5-pro | generateContent | |
| Anthropic | anthropic/claude-sonnet-4-5@20250929 | rawPredict (Anthropic Messages) |
| Mistral AI | mistralai/mistral-medium-3@001 | rawPredict |
| Others (Meta, Qwen, …) | meta/llama-3.3-70b-instruct-maas | Vertex’s OpenAI-compatible endpoint |
TypeSafe (Jev)
In Models → Add account → TypeSafe AI, enter a TypeSafe API key and addjev-latest, or a pinned version such as jev-1.13.0. TypeSafe bills input tokens only, so set the output price to 0. Call Jev either way:
- Native:
POST /v1/systemonewith{"model", "state", "questions"}. It is forwarded as is and answers{"model", "answers", "usage"}, with every question type and its calibrated probabilities. - OpenAI SDK:
POST /v1/chat/completionswith aresponse_formatJSON schema, whose properties become the questions.
Cerebras and xAI
Both use the OpenAI format and list their models live. Enter acsk-… key for Cerebras or an xai-… key for xAI; a key of the wrong shape is refused when you save it.
OpenAI-compatible providers
These providers use the OpenAI wire format natively. You can omitupstream_url — it defaults to the provider’s public API:
- Groq
- Together AI
- Ollama (local)
- DeepSeek
Add the model to a team
After adding a model, make sure it appears in the team’s model allow-list:["*"] to allow all configured models.