/v1/responses is OpenAI’s current inference surface and the one the SDKs reach for by default. ManyLayers implements it by translating requests onto chat completions and translating the result back, so it runs the same pipeline — RBAC, firewall, guardrails, PII, budgets, caching, routing, failover, metering, audit — and works against every provider the gateway supports, including providers whose own API has never implemented Responses.

Create a response

from openai import OpenAI

client = OpenAI(api_key="ml-your-key", base_url="https://your-gateway/v1")

response = client.responses.create(
    model="gpt-4o",
    instructions="You are a terse assistant.",
    input="What is the capital of France?",
)
print(response.output_text)

Response

{
  "id": "resp_abc123",
  "object": "response",
  "created_at": 1700000000,
  "status": "completed",
  "model": "gpt-4o",
  "output": [
    {
      "type": "message",
      "id": "msg_abc123",
      "status": "completed",
      "role": "assistant",
      "content": [
        { "type": "output_text", "text": "Paris.", "annotations": [] }
      ]
    }
  ],
  "usage": {
    "input_tokens": 18,
    "input_tokens_details": { "cached_tokens": 0 },
    "output_tokens": 2,
    "output_tokens_details": { "reasoning_tokens": 0 },
    "total_tokens": 20
  },
  "parallel_tool_calls": true,
  "tool_choice": "auto",
  "tools": [],
  "error": null,
  "incomplete_details": null,
  "metadata": {}
}
status is completed, or incomplete with incomplete_details.reason set to max_output_tokens or content_filter when generation was cut short.

Parameters

ParameterStatusNotes
modelSupportedSame resolution as chat completions, aliases included.
inputSupportedA string, or an array of message / function_call / function_call_output items.
instructionsSupportedBecomes a leading system message.
max_output_tokensSupported
temperature, top_pSupported
streamSupportedSee Streaming.
tools, tool_choiceSupportedFunction tools only.
parallel_tool_callsProvider dependentSee the parameter matrix.
text.formatProvider dependentMaps to response_format; text, json_object and json_schema.
reasoning.effortProvider dependentMaps to reasoning_effort.
user, metadata, storePartially supporteduser is honored for sticky routing and provider metadata. metadata and store are accepted and ignored — the gateway stores no conversations.
previous_response_idNot supportedRejected with 400. See Statelessness.
Hosted tools (web_search, file_search, computer_use)Not supportedRejected with 400.

Input items

input accepts the single-turn shorthand or the full item array:
{
  "model": "gpt-4o",
  "input": [
    { "role": "user", "content": [{ "type": "input_text", "text": "Weather in NYC?" }] },
    { "type": "function_call", "call_id": "call_1", "name": "get_weather",
      "arguments": "{\"city\":\"NYC\"}" },
    { "type": "function_call_output", "call_id": "call_1", "output": "12C" }
  ]
}
input_text and input_image content parts are translated to the chat text and image_url parts. A part type the gateway cannot translate is rejected rather than dropped.

Tool calling

Responses tools are flat where chat tools nest under function; the gateway translates both directions.
response = client.responses.create(
    model="claude-model",
    input="What is the weather in NYC?",
    tools=[{
        "type": "function",
        "name": "get_weather",
        "parameters": {"type": "object", "properties": {"city": {"type": "string"}}},
    }],
)
A tool call comes back as a function_call output item:
{
  "type": "function_call",
  "id": "fc_call_abc",
  "call_id": "call_abc",
  "name": "get_weather",
  "arguments": "{\"city\":\"NYC\"}",
  "status": "completed"
}
Send the result back as a function_call_output item with the same call_id.

Streaming

Set "stream": true for the Responses event stream. Each event carries a sequence_number and is framed with both an SSE event: name and a data: payload.
stream = client.responses.create(
    model="gpt-4o", input="Count to three", stream=True,
)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="")
Emitted events, in order:
EventWhen
response.createdThe first upstream chunk arrives.
response.in_progressImmediately after.
response.output_item.addedA message or function_call item opens.
response.content_part.addedThe text part of a message opens.
response.output_text.deltaEach incremental text fragment.
response.function_call_arguments.deltaEach incremental argument fragment.
response.output_text.doneThe message text is complete.
response.content_part.doneThe text part closes.
response.function_call_arguments.doneA call’s arguments are complete.
response.output_item.doneAn item closes.
response.completed / response.incompleteThe final response object, with usage.
Translation happens as each upstream chunk arrives — nothing is buffered waiting for the stream to end.

Statelessness

previous_response_id is rejected:
{
  "error": {
    "message": "previous_response_id is not supported: the gateway is stateless and does not store conversations; send the full conversation in input instead",
    "type": "invalid_request_error",
    "code": "unsupported_parameter"
  }
}
Storing conversations would make the data plane stateful and put a database read on the request hot path, which is the property this architecture exists to avoid. Send the whole conversation in input — the same thing the SDK does for chat completions. For conversation storage with history, sharing, folders and search, use the Workspace chat API, which is a control-plane surface.

Errors

The error envelope is identical to chat completions, and error.param names the Responses field you sent — not the chat parameter it was translated into:
{
  "error": {
    "message": "max_output_tokens must be greater than 0, got 0",
    "type": "invalid_request_error",
    "param": "max_output_tokens",
    "code": "invalid_value"
  }
}
See OpenAI Compatibility for the full status-code map.