POST /v1/messages accepts Anthropic Messages API requests and answers in Anthropic format. Use it when your code or tools already speak Anthropic: the gateway translates the request, so the model behind it can be any provider — OpenAI, Bedrock, Gemini, a self-hosted model, or Anthropic itself.
How it works
- The gateway converts the Anthropic body into an OpenAI chat completion.
- The request runs through the same pipeline as
/v1/chat/completions: RBAC, policies, guardrails, PII redaction, cache, routing, metering and audit. - The response — or the SSE stream — is converted back into Anthropic framing.
Request structure
Key fields
| Anthropic field | Becomes | Notes |
|---|---|---|
system | A leading system message | Text blocks are joined with blank lines. |
messages[].content | Message content | String or blocks. text, tool_use and tool_result blocks are translated. |
tool_use (assistant) | tool_calls | input becomes the JSON arguments. |
tool_result (user) | A tool message | content must be a string. |
tools[].input_schema | function.parameters | Missing schema defaults to {"type":"object"}. |
tool_choice | auto, required (any), none, or a named function (tool) | Applied only when tools is set. |
stop_sequences | stop | |
temperature, top_p, max_tokens, stream | Same name |
top_k, metadata, thinking, cache_control, and image or document blocks — are ignored rather than rejected. Requests without model or with max_tokens missing or 0 return 400 invalid_request_error.
Responses
A non-streaming response is an Anthropicmessage:
stop_reason maps from the upstream finish reason: stop → end_turn, length → max_tokens, tool_calls → tool_use. Tool calls come back as tool_use content blocks.
Streaming
With"stream": true the gateway emits Anthropic SSE events in this order: message_start, then content_block_start / content_block_delta / content_block_stop for each text block (text_delta) or tool call (input_json_delta), then message_delta with the stop_reason, and finally message_stop. No ping events are sent.
Errors
Errors raised by the translation or the inference pipeline use the Anthropic envelope with the gateway’s HTTP status:error.type is the gateway’s error type, such as invalid_request_error, permission_error, rate_limit_error or server_error; the gateway’s code field is not included. Failures that happen before the request reaches /v1/messages (a missing or invalid credential, product_not_entitled) and unrouted paths such as /v1/messages/count_tokens still use the OpenAI envelope.
Authentication
The gateway reads onlyAuthorization: Bearer <token>. The Anthropic SDKs send x-api-key by default, which the gateway ignores, so configure them to send a bearer token.
Common configurations
Claude Code
Claude Code
Point Claude Code at the gateway with environment variables. Make sure the model names Claude Code requests exist in your workspace, as models, aliases or virtual models.
ANTHROPIC_AUTH_TOKEN is sent as Authorization: Bearer, which is what the gateway needs; don’t use ANTHROPIC_API_KEY, which is sent as x-api-key.Route Anthropic-format traffic to another provider
Route Anthropic-format traffic to another provider
Send any model name from
GET /v1/models, or a virtual model, as model. The translation is the same whatever the provider.Tool use loop
Tool use loop
Return each tool’s output as a
tool_result block whose content is a plain string, with tool_use_id set to the id of the matching tool_use block.X-ManyLayers-* request headers (config, metadata, guardrails, cache control) work on /v1/messages, and response headers such as X-ManyLayers-Trace-Id are passed through.Next steps
Making Requests
OpenAI SDK, LangChain and LlamaIndex setup.
Request & Response Headers
Steer and inspect requests with headers.
Routing
Fallback and virtual models behind one name.
Endpoints
Everything else the gateway serves.