ManyLayers speaks the OpenAI API. Point an existing application’s base_url at your
gateway, swap the key for a ManyLayers key, and the code keeps working — across every
provider the gateway routes to, not only OpenAI.
from openai import OpenAI
client = OpenAI(api_key="ml-your-key", base_url="https://your-gateway/v1")
client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Hello"}],
)
This page is the precise version of that claim. Each item is marked:
| Mark | Meaning |
|---|
| Supported | Works the same way it does against OpenAI. |
| Partially supported | Works, with a stated limit. |
| Provider dependent | Works where the provider’s adapter can express it; a clear 400 where it cannot. See the parameter matrix. |
| Not supported | Rejected with an error saying so. Never silently ignored. |
Endpoints
| Endpoint | Method | Status | Notes |
|---|
/v1/chat/completions | POST | Supported | Streaming and non-streaming, tools, structured outputs. |
/v1/responses | POST | Partially supported | Translated onto chat completions, so it works for every provider. Stateless only — see Responses API. |
/v1/completions | POST | Supported | Legacy text completions. |
/v1/embeddings | POST | Supported | Translated for Gemini, Vertex, Cohere and Bedrock. |
/v1/models | GET | Supported | Lists what resolves for your key, including aliases. |
/v1/models/{model} | GET | Supported | A model you may not use reports 404, as OpenAI does for an unknown model. |
/v1/images/generations | POST | Provider dependent | Passthrough to the resolved provider. |
/v1/images/edits | POST | Provider dependent | Passthrough. |
/v1/images/variations | POST | Provider dependent | Passthrough. |
/v1/audio/speech | POST | Provider dependent | Passthrough. |
/v1/audio/transcriptions | POST | Provider dependent | Passthrough. |
/v1/audio/translations | POST | Provider dependent | Passthrough. |
/v1/moderations | POST | Supported | Built-in classifier, or a configured upstream. |
/v1/files, /v1/fine_tuning/jobs | — | Provider dependent | Passthrough to the resolved provider. |
/v1/messages | POST | Supported | Anthropic-native ingress, for the Anthropic SDKs. |
/v1/assistants, /v1/threads, /v1/vector_stores | any | Not supported | Stateful, OpenAI-hosted surfaces. Every method and sub-path answers 404 unsupported_endpoint. Use Knowledge and Agents instead. |
/v1/realtime | GET | Partially supported | WebSocket tunnel; no policy translation inside the socket. |
ManyLayers also serves endpoints OpenAI does not: /v1/rerank, /v1/kb/*,
/v1/docsets/*, /v1/mcp/{server} and /v1/async/*. See Endpoints.
Chat completions parameters
Supported — accepted, validated, and either forwarded or translated for the
resolved provider:
model · messages · stream · stream_options · temperature · top_p ·
max_tokens · max_completion_tokens · stop · user · tools · tool_choice ·
response_format
Provider dependent — honored where the provider’s adapter can express them,
otherwise refused with 400 and the parameter named:
n · seed · presence_penalty · frequency_penalty · logprobs ·
top_logprobs · logit_bias · parallel_tool_calls · reasoning_effort ·
response_format
Passed through untouched — unknown parameters are forwarded to
OpenAI-wire-format providers rather than stripped, so a new OpenAI parameter works
before ManyLayers has heard of it.
Validation
Requests are validated against the OpenAI schema before any provider is contacted.
An invalid request comes back as a 400 naming the field, which is what the SDKs
surface as error.param:
{
"error": {
"message": "temperature must be between 0 and 2, got 5",
"type": "invalid_request_error",
"param": "temperature",
"code": "invalid_value"
}
}
Validated bounds: temperature 0–2, top_p 0–1, presence_penalty and
frequency_penalty −2–2, n 1–128, top_logprobs 0–20, stop at most four
sequences, max_tokens and max_completion_tokens above zero, reasoning_effort
one of minimal/low/medium/high, and the structure of messages, tools,
tool_choice and response_format.
Provider parameter matrix
The gateway’s public API is uniform; the providers beneath it are not. This is what
each adapter can translate. A ✗ is not a silent drop — the request is refused
with 400 unsupported_parameter and error.param set, so you learn it from the
response rather than from a wrong answer.
| Parameter | OpenAI · Azure · Groq · Mistral · Together · xAI · DeepSeek · Fireworks · OpenRouter · Perplexity · Cerebras · SambaNova · Ollama | Anthropic | Gemini · Vertex | Cohere | Bedrock |
|---|
tools, tool_choice | ✓ | ✓ | ✓ | ✓ | ✓ |
parallel_tool_calls: false | ✓ | ✓ | ✗ | ✗ | ✗ |
response_format: json_object | ✓ | ✗ | ✓ | ✓ | ✗ |
response_format: json_schema | ✓ | ✗ | ✓ | ✓ | ✗ |
n > 1 | ✓ | ✗ | ✓ | ✗ | ✗ |
seed | ✓ | ✗ | ✓ | ✓ | ✗ |
presence_penalty, frequency_penalty | ✓ | ✗ | ✓ | ✓ | ✗ |
logprobs | ✓ | ✗ | ✓ | ✓ | ✗ |
logit_bias | ✓ | ✗ | ✗ | ✗ | ✗ |
reasoning_effort | ✓ | ✓ | ✗ | ✗ | ✗ |
A parameter set to its default is never refused: n: 1, presence_penalty: 0,
parallel_tool_calls: true, tool_choice: "auto", logprobs: false and an empty
tools array ask for the behavior you would get anyway, so they pass on every
provider. SDK wrappers that set them unconditionally keep working.
Provider-specific translations
- Anthropic —
reasoning_effort becomes an extended-thinking budget
(minimal 1k, low 2k, medium 8k, high 16k tokens); temperature and
top_p are dropped while thinking is on, because Anthropic rejects them.
parallel_tool_calls: false becomes tool_choice.disable_parallel_tool_use.
user becomes metadata.user_id. Thinking traces stream as
delta.reasoning_content.
- Gemini / Vertex —
response_format becomes responseMimeType plus
responseSchema. JSON Schema keywords Gemini’s OpenAPI-subset schema rejects
(additionalProperties, $defs, allOf, oneOf, const, …) are stripped, so
a schema OpenAI accepts does not turn into a provider 400.
- Cohere —
response_format: json_schema maps to Cohere v2’s
response_format: json_object with the schema attached. A named
tool_choice degrades to REQUIRED: Cohere can force a tool call but not a
specific one.
- Bedrock —
tool_choice: "none" has no Converse equivalent, so the tool
config is withheld for that turn.
Streaming
Supported. Content-Type: text/event-stream, standard data: framing, a role
chunk first, incremental content deltas, tool-call deltas correlated by index,
a finish_reason chunk, and data: [DONE].
The gateway relays each SSE line as it arrives; a provider chunk is never held
waiting for the next one. Streams are only buffered when a policy requires the
whole completion before any byte may leave — output PII redaction in buffer mode,
or an enforcing after-phase guardrail.
Provider streams in native formats (Anthropic, Gemini, Vertex, Bedrock, Cohere) are
translated to OpenAI chunk framing incrementally, event by event.
stream = client.chat.completions.create(
model="claude-model",
messages=[{"role": "user", "content": "Hello"}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
Providers translated from a native format always emit a final usage chunk, whether
or not stream_options.include_usage was set. It is an extra chunk with an empty
choices array, which the OpenAI SDKs ignore; OpenAI-wire-format providers follow
include_usage exactly.
Cancellation
Closing the connection cancels the request end to end: the client disconnect
cancels the gateway’s request context, which cancels the upstream HTTP request. No
generation is left running for a client that has gone away.
Supported on every provider in the matrix, streaming and non-streaming:
tools, tool_choice (auto / none / required / a named function)
- single and parallel tool calls
- stable
tool_calls[].id, correlated by index across streamed fragments
arguments delivered as fragments that concatenate into valid JSON
- assistant tool-call turns and
role: "tool" result messages on the way back in
finish_reason: "tool_calls"
Provider-native tool formats — Anthropic tool_use/tool_result, Gemini
functionCall/functionResponse, Bedrock toolUse/toolResult, Cohere v2 tool
calls — are translated in both directions.
Structured outputs
Provider dependent. response_format of type text, json_object and
json_schema is validated by the gateway and then mapped to the provider’s own
mechanism. Where a provider has no mechanism, the request is refused:
{
"error": {
"message": "response_format is not supported for this model's provider (anthropic): this provider does not support schema-constrained output",
"type": "invalid_request_error",
"param": "response_format",
"code": "unsupported_parameter"
}
}
The gateway does not emulate structured output with prompt instructions on
providers that lack it. A schema that is not enforced is worse than one that is
refused.
Usage
Supported. Every provider’s usage is normalized into OpenAI’s shape:
{
"usage": {
"prompt_tokens": 100,
"completion_tokens": 25,
"total_tokens": 125,
"prompt_tokens_details": { "cached_tokens": 90 },
"completion_tokens_details": { "reasoning_tokens": 12 }
}
}
prompt_tokens means the same number on every provider. Anthropic reports cache
reads outside input_tokens, so they are folded in; Gemini counts thinking tokens
outside candidatesTokenCount, so they are folded into completion_tokens. The
detail blocks are present only when the provider reported them.
When a provider reports no usage, the gateway estimates it with the model’s
tokenizer and records usage_source: "estimated" in the audit
log — the response body still carries whatever the
provider sent.
Models
Supported. GET /v1/models returns what would actually resolve for the calling
key: your workspace’s own models for a workspace-bound key, or the configured
catalog for a team-scoped key, filtered by the team’s model policy.
{
"object": "list",
"data": [
{ "id": "gpt-4o", "object": "model", "created": 0, "owned_by": "manylayers",
"provider": "openai", "source": "config" }
]
}
provider and source are ManyLayers extensions; they name the upstream without
exposing endpoints or credentials. created is always 0: a catalog entry has no
training date to report, and the field is typed as required by the SDKs.
Resolution works for direct model ids, aliases, workspace-registered models,
routed models and fallback targets. See Model resolution.
Aliases
A model may declare additional names. Both resolve, both are listed, and a team
policy naming either one grants both:
models:
- logical_name: house-fast
aliases: ["gpt-4o", "gpt-4o-mini"]
provider: groq
upstream_model: llama-3.3-70b-versatile
Authentication
Supported. Standard bearer authentication is all an OpenAI SDK needs:
Authorization: Bearer ml-your-api-key
Clients never send provider credentials. The gateway resolves the provider key
from the workspace’s stored credential or the operator’s configuration, and never
forwards the client’s Authorization header upstream.
| Condition | Status | error.code |
|---|
| Missing key | 401 | invalid_api_key |
| Invalid or revoked key | 401 | invalid_api_key |
| Expired key | 401 | key_expired |
| Model outside the team’s policy | 403 | model_not_allowed |
| Team request rate exceeded | 429 | rpm_exceeded |
| Team token rate exceeded | 429 | tpm_exceeded |
| Per-key rate exceeded | 429 | rpm_exceeded / tpm_exceeded |
| Monthly token budget exhausted | 429 | budget_exhausted |
| Monthly key budget exceeded | 402 | key_budget_exceeded |
Workspaces are isolated: a key resolves models from its own workspace and its
team’s policy only. A model another team can use is reported as forbidden on the
inference path and as absent on GET /v1/models/{model}, because telling an
unauthorized caller that a model exists is itself a disclosure.
Errors
Supported. Every failure uses the OpenAI envelope, with param set whenever a
specific field is at fault:
{
"error": {
"message": "...",
"type": "invalid_request_error",
"param": "temperature",
"code": "invalid_value"
}
}
Provider errors are translated, never relayed. A provider’s status and body would
leak upstream detail and make behavior depend on which provider happened to serve
the model, so both are logged server-side and answered in the gateway’s own terms:
| Failure | Status | error.type |
|---|
| Request invalid (gateway validation) | 400 | invalid_request_error |
| Parameter unsupported by the provider | 400 | invalid_request_error |
| Provider rejected the request (4xx) | 400 | invalid_request_error |
| Authentication | 401 | invalid_request_error |
| Authorization | 403 | permission_error |
| Model not found | 404 | invalid_request_error |
| Endpoint not implemented by the gateway | 404 | invalid_request_error |
| Payload too large | 413 | invalid_request_error |
| Rate limited (gateway or provider) | 429 | rate_limit_error |
| Monthly token budget exhausted | 429 | insufficient_quota |
| Monthly key budget exceeded | 402 | insufficient_quota |
| Provider credential rejected | 502 | server_error |
| Provider unavailable or 5xx | 502 | server_error |
| Provider response unreadable | 502 | server_error |
| Provider timeout | 504 | server_error |
A provider 4xx becomes a 400 because the request is the caller’s to fix; a
provider 401 becomes a 502 because the credential that was rejected is the
gateway’s, not the caller’s.
Context-length errors arrive as the provider’s own 4xx and surface as 400
invalid_request_error. The provider’s wording stays in the gateway log rather
than in the response body.
Not supported
| Feature | Why |
|---|
| Assistants, Threads, Vector Stores | Stateful OpenAI-hosted surfaces. Use Knowledge and Agents. Refused with 404 unsupported_endpoint. |
previous_response_id and stored Responses conversations | The data plane is stateless by design. Send the whole conversation in input. Refused with 400 unsupported_parameter. |
OpenAI-hosted tools (web_search, file_search, computer_use) | ManyLayers routes to model providers and does not run hosted tools. Refused with 400, error.param naming the tool. Use web search. |
Deprecated functions / function_call | Superseded by tools / tool_choice. Refused with 400 unsupported_parameter on every endpoint and every provider — an OpenAI-wire provider would honor them while every native adapter drops them, so the same request would mean two different things. A null value, and the per-message function_call of a stored assistant turn, are still accepted. |