POST /v1/messages accepts Anthropic Messages API requests and answers in Anthropic format. Use it when your code or tools already speak Anthropic: the gateway translates the request, so the model behind it can be any provider — OpenAI, Bedrock, Gemini, a self-hosted model, or Anthropic itself.

How it works

  1. The gateway converts the Anthropic body into an OpenAI chat completion.
  2. The request runs through the same pipeline as /v1/chat/completions: RBAC, policies, guardrails, PII redaction, cache, routing, metering and audit.
  3. The response — or the SSE stream — is converted back into Anthropic framing.
Because the request runs as a chat completion, models from workspace providers and routing configs (including virtual models) work here too.

Request structure

{
  "model": "gpt-4o-mini",          // required: any model GET /v1/models lists
  "max_tokens": 1024,              // required, > 0
  "system": "You are concise.",    // string or array of text blocks
  "messages": [
    {"role": "user", "content": "What's the weather in Paris?"}
  ],
  "tools": [{
    "name": "get_weather",
    "description": "Current weather for a city",
    "input_schema": {"type": "object", "properties": {"city": {"type": "string"}}}
  }],
  "tool_choice": {"type": "auto"}, // auto | any | none | tool (+ name)
  "temperature": 0.2,
  "top_p": 0.9,
  "stop_sequences": ["END"],
  "stream": false
}

Key fields

Anthropic fieldBecomesNotes
systemA leading system messageText blocks are joined with blank lines.
messages[].contentMessage contentString or blocks. text, tool_use and tool_result blocks are translated.
tool_use (assistant)tool_callsinput becomes the JSON arguments.
tool_result (user)A tool messagecontent must be a string.
tools[].input_schemafunction.parametersMissing schema defaults to {"type":"object"}.
tool_choiceauto, required (any), none, or a named function (tool)Applied only when tools is set.
stop_sequencesstop
temperature, top_p, max_tokens, streamSame name
Fields not listed — for example top_k, metadata, thinking, cache_control, and image or document blocks — are ignored rather than rejected. Requests without model or with max_tokens missing or 0 return 400 invalid_request_error.

Responses

A non-streaming response is an Anthropic message:
{
  "id": "chatcmpl-...",
  "type": "message",
  "role": "assistant",
  "model": "gpt-4o-mini",
  "content": [{"type": "text", "text": "..."}],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {"input_tokens": 21, "output_tokens": 9}
}
stop_reason maps from the upstream finish reason: stop → end_turn, length → max_tokens, tool_calls → tool_use. Tool calls come back as tool_use content blocks.

Streaming

With "stream": true the gateway emits Anthropic SSE events in this order: message_start, then content_block_start / content_block_delta / content_block_stop for each text block (text_delta) or tool call (input_json_delta), then message_delta with the stop_reason, and finally message_stop. No ping events are sent.

Errors

Errors raised by the translation or the inference pipeline use the Anthropic envelope with the gateway’s HTTP status:
{"type": "error", "error": {"type": "permission_error", "message": "team \"web\" is not allowed to use model \"gpt-4o\""}}
The error.type is the gateway’s error type, such as invalid_request_error, permission_error, rate_limit_error or server_error; the gateway’s code field is not included. Failures that happen before the request reaches /v1/messages (a missing or invalid credential, product_not_entitled) and unrouted paths such as /v1/messages/count_tokens still use the OpenAI envelope.

Authentication

The gateway reads only Authorization: Bearer <token>. The Anthropic SDKs send x-api-key by default, which the gateway ignores, so configure them to send a bearer token.
import anthropic

client = anthropic.Anthropic(
    base_url="https://app.manylayers.io",  # the SDK appends /v1/messages
    auth_token="ml-...",                     # sent as Authorization: Bearer
)
msg = client.messages.create(
    model="gpt-4o-mini",
    max_tokens=512,
    messages=[{"role": "user", "content": "Hello"}],
)

Common configurations

Point Claude Code at the gateway with environment variables. ANTHROPIC_AUTH_TOKEN is sent as Authorization: Bearer, which is what the gateway needs; don’t use ANTHROPIC_API_KEY, which is sent as x-api-key.
export ANTHROPIC_BASE_URL="https://app.manylayers.io"
export ANTHROPIC_AUTH_TOKEN="ml_pat_..."
export ANTHROPIC_MODEL="claude-sonnet"   # a model name GET /v1/models lists for your key
claude
Make sure the model names Claude Code requests exist in your workspace, as models, aliases or virtual models.
Send any model name from GET /v1/models, or a virtual model, as model. The translation is the same whatever the provider.
Return each tool’s output as a tool_result block whose content is a plain string, with tool_use_id set to the id of the matching tool_use block.
Translation covers text and tool use. Images, documents, extended thinking and prompt caching blocks are dropped, and a message containing a tool_result whose content is an array of blocks can’t be decoded and is dropped. /v1/messages/count_tokens isn’t implemented and returns 404 unsupported_endpoint. Agents that depend on those features — including some Claude Code workflows — may not behave as they do against Anthropic directly.
X-ManyLayers-* request headers (config, metadata, guardrails, cache control) work on /v1/messages, and response headers such as X-ManyLayers-Trace-Id are passed through.

Next steps

Making Requests

OpenAI SDK, LangChain and LlamaIndex setup.

Request & Response Headers

Steer and inspect requests with headers.

Routing

Fallback and virtual models behind one name.

Endpoints

Everything else the gateway serves.