The ManyLayers AI Gateway is a single OpenAI-compatible endpoint that sits between your applications and the model providers you use. You point any OpenAI SDK at it, and it authenticates the caller, applies your access rules, limits, guardrails and PII redaction, routes the request to the right provider, and records cost, metrics and a tamper-evident audit trail.

Key features

Making requests

Use the OpenAI SDK unchanged: chat, completions, embeddings, responses, streaming.

Anthropic Messages API

Call /v1/messages in Anthropic format and serve it from any provider.

Providers

Connect 20 provider types, from OpenAI and Bedrock to vLLM and Ollama.

Virtual models

Expose a routing config under a model name your clients can call directly.

Routing and fallback

Fallback, load balancing, canary, conditional and latency-based routing.

Auto routing

Classify each turn’s complexity and send it to the right tier of model.

Caching

Exact and semantic response caching, keyed on the redacted request.

Async requests

Queue completions with /v1/async/*, poll for the result or get a webhook.

Access control

Roles, team model allow-lists and provider account access.

Rate limits

Request and token limits per minute, hour or day at every level.

Budgets

Spending caps in USD or tokens; the most restrictive one wins.

Guardrails

Validate and mutate inputs and outputs with built-in or custom checks.

PII redaction

Redact personal data before it reaches a provider, with an optional reversible vault.

Prompt injection

Score prompts for injection attempts and audit or block them.

MCP gateway

Proxy Model Context Protocol servers through /v1/mcp/{server}.

Tracing

Send OpenTelemetry traces to your own collector per workspace.

Metrics

Prometheus metrics for requests, tokens, cache, policy and upstream health.

Self-hosting

Run the gateway, workspace and deployer services on your own infrastructure.

Supported providers

Each provider type has an adapter. Providers marked as passthrough speak the OpenAI wire format, so the gateway forwards the request body unchanged and the provider’s own API decides what it accepts.
Providerprovider_typeChatEmbeddingsImagesAudioRerank
OpenAIopenaiYesYesYesYesPassthrough
Azure OpenAIazureYesYesYesYesPassthrough
AnthropicanthropicYesNoNoNoNo
Google GeminigeminiYesYesNoNoNo
Google Vertex AIvertexYesYesNoNoNo
AWS BedrockbedrockYesYesNoNoNo
CoherecohereYesYesNoNoYes (translated)
TypeSafetypesafeYes (structured only)NoNoNoNo
Self-hosted (vLLM, TGI, LM Studio, …)self_hostedYesPassthroughNoNoPassthrough
OllamaollamaYesPassthroughNoNoPassthrough
MistralmistralYesPassthroughNoNoPassthrough
GroqgroqYesPassthroughNoNoPassthrough
DeepSeekdeepseekYesPassthroughNoNoPassthrough
Together AItogetherYesPassthroughNoNoPassthrough
PerplexityperplexityYesPassthroughNoNoPassthrough
xAIxaiYesPassthroughNoNoPassthrough
Fireworks AIfireworksYesPassthroughNoNoPassthrough
OpenRouteropenrouterYesPassthroughNoNoPassthrough
CerebrascerebrasYesPassthroughNoNoPassthrough
AI21 Labsai21YesPassthroughNoNoPassthrough
SambaNovasambanovaYesPassthroughNoNoPassthrough
For adapters that translate to a native API (Anthropic, Gemini, Vertex, Bedrock, Cohere, TypeSafe), the gateway checks each request parameter it cannot translate faithfully, such as n, seed or response_format, and returns 400 unsupported_parameter naming that parameter. It never drops the parameter silently. TypeSafe requires response_format with a JSON schema.
Providers you register in the console serve /v1/chat/completions, /v1/responses and /v1/messages. Every other endpoint (completions, embeddings, images, audio, rerank, moderations, files, fine-tuning, realtime) resolves only from models declared under models: in gateway.yaml. Azure, Bedrock and Vertex need per-account addressing (and, for Bedrock, a region and SigV4 key pair), so configure those three in gateway.yaml.

Supported APIs

All endpoints live under https://gateway.example.com/v1 and accept Authorization: Bearer <key>.
EndpointMethodNotes
/v1/chat/completionsPOSTOpenAI Chat Completions, streaming supported
/v1/completionsPOSTLegacy text completions
/v1/embeddingsPOSTEmbeddings
/v1/responsesPOSTOpenAI Responses API, stateless; translated to chat so it works with every provider
/v1/messagesPOSTAnthropic Messages format, including SSE, served by any provider
/v1/models, /v1/models/{model}GETModels the caller can use
/v1/async/chat/completions, /v1/async/completionsPOSTEnqueue a job
/v1/async/jobs/{id}GETPoll a queued job
/v1/realtimeGET (WebSocket)Realtime API, openai and azure only; pass ?model=
/v1/images/generations, /edits, /variationsPOSTopenai and azure only
/v1/audio/speech, /transcriptions, /translationsPOSTopenai and azure only
/v1/moderationsPOSTPassthrough to openai/azure, built-in fallback otherwise
/v1/rerankPOSTTranslated for Cohere, passthrough for OpenAI-wire providers
/v1/files, /v1/fine_tuning/jobsGET, POST, DELETEopenai and azure only; pick the model with X-ManyLayers-Model or ?model=
/v1/mcp/{server}POSTMCP server proxy
/v1/kb, /v1/docsetsGET, POST, PUT, DELETEKnowledge bases and document sets
Any other OpenAI path under /v1 returns 404 with code unsupported_endpoint in the standard OpenAI error envelope. The gateway also serves /health, /healthz, /ready, /readyz, /version and /metrics.

Next steps

Quick Start

Add a provider, create a key and send your first request.

Gateway Architecture

The three services, the data stores and the request pipeline.

Authentication

API keys, personal access tokens, virtual accounts and OIDC.

Providers

Configure each provider type in detail.