This guide shows how to add real AI providers to your gateway configuration. Once configured, your clients can use any provider’s models through the standard OpenAI API — no changes needed in your application code.

OpenAI

models:
  - logical_name: gpt-4o
    provider: openai
    upstream_url: https://api.openai.com
    upstream_model: gpt-4o
    upstream_api_key: ${OPENAI_API_KEY}
    input_price_per_1k: 0.005
    output_price_per_1k: 0.015
Set the environment variable and reload the config:
export OPENAI_API_KEY=sk-...
docker compose kill -s HUP gateway

Anthropic

The gateway translates OpenAI chat completions to Anthropic’s /v1/messages format automatically — your clients send standard OpenAI requests.
models:
  - logical_name: claude-sonnet
    provider: anthropic
    upstream_url: https://api.anthropic.com
    upstream_model: claude-sonnet-4-6
    upstream_api_key: ${ANTHROPIC_API_KEY}
    input_price_per_1k: 0.003
    output_price_per_1k: 0.015

Azure OpenAI

models:
  - logical_name: gpt-4o-azure
    provider: azure
    upstream_url: https://myresource.openai.azure.com
    upstream_model: my-gpt4o-deployment
    upstream_api_key: ${AZURE_OPENAI_API_KEY}
    # api_version: 2024-06-01
    input_price_per_1k: 0.005
    output_price_per_1k: 0.015

AWS Bedrock

From the console

In Models → Add account → AWS Bedrock, each account brings its own AWS credentials:
  • AWS region: where the models are called, such as us-east-1. Requests are signed for it, and the endpoint is bedrock-runtime.<region>.amazonaws.com. Set a Base URL under Show advanced fields only for a VPC endpoint.
  • AWS Account Auth Data: choose one method.
    • Access Key Based: an IAM user’s access key ID and secret. A session token field appears for temporary (ASIA…) keys.
    • Assumed Role Based: the ARN of a role in your account, plus an optional external ID. The gateway assumes the role with its own AWS identity: its environment variables, EKS service-account role, ECS task role or EC2 instance role, found in that order. The role’s trust policy must allow that identity:
      {
        "Effect": "Allow",
        "Principal": { "AWS": "arn:aws:iam::<gateway-account-id>:role/<gateway-role>" },
        "Action": "sts:AssumeRole",
        "Condition": { "StringEquals": { "sts:ExternalId": "<your-external-id>" } }
      }
      
      Use an external ID. Without one, anyone who can add accounts on the same gateway and knows your role ARN could assume it.
    • API Key Based: a key from the Bedrock console’s API keys page (ABSK…), sent as a bearer token.
    Each secret field’s ••• menu switches between entering the value and naming a secret reference (${ENV_VAR} or ${vault:path#key}).
The identity that calls Bedrock (the IAM user, the assumed role, or the API key’s user) needs this policy. The two List* actions are optional; they let the console list your models for you. Without them, add models by ID.
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"],
      "Resource": [
        "arn:aws:bedrock:*::foundation-model/*",
        "arn:aws:bedrock:*:*:inference-profile/*",
        "arn:aws:bedrock:*:*:application-inference-profile/*"
      ]
    },
    {
      "Effect": "Allow",
      "Action": ["bedrock:ListFoundationModels", "bedrock:ListInferenceProfiles", "bedrock:GetInferenceProfile"],
      "Resource": "*"
    }
  ]
}
Claude 4 and later models are served only through inference profiles. Use the profile ID (us.anthropic.claude-sonnet-4-5-20250929-v1:0, global.anthropic.…) rather than the bare model ID. A cross-region profile routes to other regions, so the policy must allow them too, which the wildcards above do. An application inference profile is used by its full ARN.

In gateway.yaml

Pass credentials as ACCESS_KEY:SECRET_KEY (or append :SESSION_TOKEN for STS temporary credentials), or pass a Bedrock API key as is:
models:
  - logical_name: claude-bedrock
    provider: bedrock
    upstream_url: https://bedrock-runtime.us-east-1.amazonaws.com
    upstream_model: us.anthropic.claude-sonnet-4-5-20250929-v1:0
    upstream_api_key: ${AWS_ACCESS_KEY_ID}:${AWS_SECRET_ACCESS_KEY}   # or ${AWS_BEARER_TOKEN_BEDROCK}
    aws_region: us-east-1
    input_price_per_1k: 0.003
    output_price_per_1k: 0.015

AI21 Labs (Jamba)

In Models → Add account → AI21 Labs, enter an AI21 Studio API key. Jamba speaks the OpenAI format at https://api.ai21.com/studio/v1/chat/completions, so requests pass through unchanged: streaming, tools and response_format: {"type": "json_object"}. AI21 has no model listing, so the console offers jamba-large and jamba-mini; add any other Jamba ID by name.

Cohere

In Models → Add account → Cohere, enter a Cohere API key. The models list comes from the key, and each model is typed by what it serves:
  • Chat (command-a-03-2025, …) on /v1/chat/completions, with tools and structured output. command-a-vision-07-2025 accepts images; command-a-reasoning-08-2025 takes reasoning_effort and returns reasoning_content.
  • Embeddings (embed-v4.0) on /v1/embeddings. Pass input_type: search_query for queries, search_document (the default) for what they search. dimensions and truncate are passed through.
  • Rerank (rerank-v3.5) on /v1/rerank: {"model", "query", "documents", "top_n"} in, {"results": [{"index", "relevance_score"}]} out.
A model is used only on its own endpoint. Embedding and rerank models registered on a console account are found first; otherwise gateway.yaml is used, as before.

Google Gemini (AI Studio)

In Models → Add account → Google Gemini, enter an AI Studio API key (AIza… or the newer AQ.…). The models list comes from the key itself. Chat supports:
  • Media: images, audio, video and PDFs, as base64 data URIs or https:// URLs. gs:// objects need a Vertex account.
  • Reasoning: reasoning_effort maps to Gemini’s thinking budget, and the reasoning comes back as reasoning_content.
  • Built-in tools: pass {"type": "google_search"}, {"type": "url_context"} or {"type": "code_execution"} in tools. Search sources come back as url_citation annotations, and executed code and its output appear in the answer as code blocks.
  • Embeddings: task_type (RETRIEVAL_QUERY, RETRIEVAL_DOCUMENT, …), title and dimensions, for example with gemini-embedding-001.

Google Vertex AI

In Models → Add account → Google Vertex:
  • Project ID and Region: the account is called at https://{region}-aiplatform.googleapis.com/v1/projects/{project}/locations/{region}. The global location uses aiplatform.googleapis.com. Set a Base URL under Show advanced fields only for Private Service Connect.
  • Google Cloud Auth Data: choose one method.
    • Service Account Key: paste the JSON key of a service account with Vertex AI User (roles/aiplatform.user). Afterwards the console shows only the service account’s email.
    • Workload Identity Federation: paste the credential-config.json from gcloud iam workload-identity-pools create-cred-config. Its token file must be a mounted Kubernetes token under /var/run/secrets/; set GCP_WIF_FILE_PREFIXES on the gateway to allow another directory. The token endpoints must be Google’s own.
    • API Key: a Vertex express-mode key. It reaches Google’s Gemini models only, and is sent without the project.
    • Switched off: the gateway’s own Google identity (GKE Workload Identity, a GCE service account, or GOOGLE_APPLICATION_CREDENTIALS).
Name models by publisher:
PublisherModel IDSent as
Googlegoogle/gemini-2.5-progenerateContent
Anthropicanthropic/claude-sonnet-4-5@20250929rawPredict (Anthropic Messages)
Mistral AImistralai/mistral-medium-3@001rawPredict
Others (Meta, Qwen, …)meta/llama-3.3-70b-instruct-maasVertex’s OpenAI-compatible endpoint

TypeSafe (Jev)

In Models → Add account → TypeSafe AI, enter a TypeSafe API key and add jev-latest, or a pinned version such as jev-1.13.0. TypeSafe bills input tokens only, so set the output price to 0. Call Jev either way:
  • Native: POST /v1/systemone with {"model", "state", "questions"}. It is forwarded as is and answers {"model", "answers", "usage"}, with every question type and its calibrated probabilities.
  • OpenAI SDK: POST /v1/chat/completions with a response_format JSON schema, whose properties become the questions.

Cerebras and xAI

Both use the OpenAI format and list their models live. Enter a csk-… key for Cerebras or an xai-… key for xAI; a key of the wrong shape is refused when you save it.

OpenAI-compatible providers

These providers use the OpenAI wire format natively. You can omit upstream_url — it defaults to the provider’s public API:
- logical_name: groq-llama
  provider: groq
  upstream_model: llama-3.3-70b-versatile
  upstream_api_key: ${GROQ_API_KEY}

Add the model to a team

After adding a model, make sure it appears in the team’s model allow-list:
teams:
  - name: engineering
    models: ["gpt-4o", "claude-sonnet", "groq-llama"]
Use ["*"] to allow all configured models.

Verify

# List models visible to your key
curl http://localhost:8180/v1/models \
  -H "Authorization: Bearer $KEY"

# Test the new model
curl http://localhost:8180/v1/chat/completions \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "claude-sonnet", "messages": [{"role": "user", "content": "Hi"}]}'