Key features
Making requests
Use the OpenAI SDK unchanged: chat, completions, embeddings, responses, streaming.
Anthropic Messages API
Call
/v1/messages in Anthropic format and serve it from any provider.Providers
Connect 20 provider types, from OpenAI and Bedrock to vLLM and Ollama.
Virtual models
Expose a routing config under a model name your clients can call directly.
Routing and fallback
Fallback, load balancing, canary, conditional and latency-based routing.
Auto routing
Classify each turn’s complexity and send it to the right tier of model.
Caching
Exact and semantic response caching, keyed on the redacted request.
Async requests
Queue completions with
/v1/async/*, poll for the result or get a webhook.Access control
Roles, team model allow-lists and provider account access.
Rate limits
Request and token limits per minute, hour or day at every level.
Budgets
Spending caps in USD or tokens; the most restrictive one wins.
Guardrails
Validate and mutate inputs and outputs with built-in or custom checks.
PII redaction
Redact personal data before it reaches a provider, with an optional reversible vault.
Prompt injection
Score prompts for injection attempts and audit or block them.
MCP gateway
Proxy Model Context Protocol servers through
/v1/mcp/{server}.Tracing
Send OpenTelemetry traces to your own collector per workspace.
Metrics
Prometheus metrics for requests, tokens, cache, policy and upstream health.
Self-hosting
Run the gateway, workspace and deployer services on your own infrastructure.
Supported providers
Each provider type has an adapter. Providers marked as passthrough speak the OpenAI wire format, so the gateway forwards the request body unchanged and the provider’s own API decides what it accepts.| Provider | provider_type | Chat | Embeddings | Images | Audio | Rerank |
|---|---|---|---|---|---|---|
| OpenAI | openai | Yes | Yes | Yes | Yes | Passthrough |
| Azure OpenAI | azure | Yes | Yes | Yes | Yes | Passthrough |
| Anthropic | anthropic | Yes | No | No | No | No |
| Google Gemini | gemini | Yes | Yes | No | No | No |
| Google Vertex AI | vertex | Yes | Yes | No | No | No |
| AWS Bedrock | bedrock | Yes | Yes | No | No | No |
| Cohere | cohere | Yes | Yes | No | No | Yes (translated) |
| TypeSafe | typesafe | Yes (structured only) | No | No | No | No |
| Self-hosted (vLLM, TGI, LM Studio, …) | self_hosted | Yes | Passthrough | No | No | Passthrough |
| Ollama | ollama | Yes | Passthrough | No | No | Passthrough |
| Mistral | mistral | Yes | Passthrough | No | No | Passthrough |
| Groq | groq | Yes | Passthrough | No | No | Passthrough |
| DeepSeek | deepseek | Yes | Passthrough | No | No | Passthrough |
| Together AI | together | Yes | Passthrough | No | No | Passthrough |
| Perplexity | perplexity | Yes | Passthrough | No | No | Passthrough |
| xAI | xai | Yes | Passthrough | No | No | Passthrough |
| Fireworks AI | fireworks | Yes | Passthrough | No | No | Passthrough |
| OpenRouter | openrouter | Yes | Passthrough | No | No | Passthrough |
| Cerebras | cerebras | Yes | Passthrough | No | No | Passthrough |
| AI21 Labs | ai21 | Yes | Passthrough | No | No | Passthrough |
| SambaNova | sambanova | Yes | Passthrough | No | No | Passthrough |
For adapters that translate to a native API (Anthropic, Gemini, Vertex, Bedrock, Cohere, TypeSafe), the gateway checks each request parameter it cannot translate faithfully, such as
n, seed or response_format, and returns 400 unsupported_parameter naming that parameter. It never drops the parameter silently. TypeSafe requires response_format with a JSON schema.Supported APIs
All endpoints live underhttps://gateway.example.com/v1 and accept Authorization: Bearer <key>.
| Endpoint | Method | Notes |
|---|---|---|
/v1/chat/completions | POST | OpenAI Chat Completions, streaming supported |
/v1/completions | POST | Legacy text completions |
/v1/embeddings | POST | Embeddings |
/v1/responses | POST | OpenAI Responses API, stateless; translated to chat so it works with every provider |
/v1/messages | POST | Anthropic Messages format, including SSE, served by any provider |
/v1/models, /v1/models/{model} | GET | Models the caller can use |
/v1/async/chat/completions, /v1/async/completions | POST | Enqueue a job |
/v1/async/jobs/{id} | GET | Poll a queued job |
/v1/realtime | GET (WebSocket) | Realtime API, openai and azure only; pass ?model= |
/v1/images/generations, /edits, /variations | POST | openai and azure only |
/v1/audio/speech, /transcriptions, /translations | POST | openai and azure only |
/v1/moderations | POST | Passthrough to openai/azure, built-in fallback otherwise |
/v1/rerank | POST | Translated for Cohere, passthrough for OpenAI-wire providers |
/v1/files, /v1/fine_tuning/jobs | GET, POST, DELETE | openai and azure only; pick the model with X-ManyLayers-Model or ?model= |
/v1/mcp/{server} | POST | MCP server proxy |
/v1/kb, /v1/docsets | GET, POST, PUT, DELETE | Knowledge bases and document sets |
/v1 returns 404 with code unsupported_endpoint in the standard OpenAI error envelope. The gateway also serves /health, /healthz, /ready, /readyz, /version and /metrics.
Next steps
Quick Start
Add a provider, create a key and send your first request.
Gateway Architecture
The three services, the data stores and the request pipeline.
Authentication
API keys, personal access tokens, virtual accounts and OIDC.
Providers
Configure each provider type in detail.