Create a chat completion

curl https://your-gateway/v1/chat/completions \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "What is the capital of France?"}
    ],
    "temperature": 0.7,
    "max_tokens": 256
  }'

Response

{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "created": 1700000000,
  "model": "gpt-4o",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The capital of France is Paris."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 25,
    "completion_tokens": 8,
    "total_tokens": 33
  }
}

Parameters

The gateway accepts the OpenAI chat completions schema and validates it before contacting a provider. model, messages, stream, stream_options, temperature, top_p, max_tokens, max_completion_tokens, stop, user, tools, tool_choice and response_format work on every provider. n, seed, presence_penalty, frequency_penalty, logprobs, top_logprobs, logit_bias, parallel_tool_calls and reasoning_effort are honored where the resolved provider’s adapter can express them, and refused with a 400 naming the parameter where it cannot — never dropped silently. Unknown parameters are forwarded to OpenAI-wire-format providers untouched. See OpenAI Compatibility for the full matrix.

Streaming

Set "stream": true to receive Server-Sent Events:
curl https://your-gateway/v1/chat/completions \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": true
  }'
Each SSE chunk follows the standard OpenAI format:
data: {"id":"chatcmpl-...","choices":[{"delta":{"content":"Hello"},"index":0}]}

data: [DONE]

Async chat completion

For long-running requests, enqueue as a background job and poll for the result:
curl -X POST https://your-gateway/v1/async/chat/completions \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Write a detailed analysis..."}]
  }'

Response

{"job_id": "job_abc123", "status": "pending"}
Poll status via GET /v1/async/jobs/{id}. When complete, the response contains the full chat completion. Configure async.webhook_signing_secret to receive results via HMAC-signed webhook instead.

ManyLayers-specific headers

HeaderDescription
X-ManyLayers-ConfigNamed routing config to apply (failover, canary, etc.)
X-Session-IdSticky session identifier — pins this request to a specific upstream
X-Request-IdRequest ID (auto-generated if you do not provide one)

Anthropic messages endpoint

For clients using the Anthropic wire format:
POST /v1/messages
Accepts Anthropic-format requests and routes them to the configured Anthropic upstream. Available when a model with provider: anthropic is configured.