/v1/responses is OpenAI’s current inference surface and the one the SDKs reach
for by default. ManyLayers implements it by translating requests onto chat
completions and translating the result back, so it runs the same pipeline —
RBAC, firewall, guardrails, PII, budgets, caching, routing, failover, metering,
audit — and works against every provider the gateway supports, including
providers whose own API has never implemented Responses.
Create a response
Response
status is completed, or incomplete with incomplete_details.reason set to
max_output_tokens or content_filter when generation was cut short.
Parameters
| Parameter | Status | Notes |
|---|---|---|
model | Supported | Same resolution as chat completions, aliases included. |
input | Supported | A string, or an array of message / function_call / function_call_output items. |
instructions | Supported | Becomes a leading system message. |
max_output_tokens | Supported | |
temperature, top_p | Supported | |
stream | Supported | See Streaming. |
tools, tool_choice | Supported | Function tools only. |
parallel_tool_calls | Provider dependent | See the parameter matrix. |
text.format | Provider dependent | Maps to response_format; text, json_object and json_schema. |
reasoning.effort | Provider dependent | Maps to reasoning_effort. |
user, metadata, store | Partially supported | user is honored for sticky routing and provider metadata. metadata and store are accepted and ignored — the gateway stores no conversations. |
previous_response_id | Not supported | Rejected with 400. See Statelessness. |
Hosted tools (web_search, file_search, computer_use) | Not supported | Rejected with 400. |
Input items
input accepts the single-turn shorthand or the full item array:
input_text and input_image content parts are translated to the chat text and
image_url parts. A part type the gateway cannot translate is rejected rather
than dropped.
Tool calling
Responses tools are flat where chat tools nest underfunction; the gateway
translates both directions.
function_call output item:
function_call_output item with the same call_id.
Streaming
Set"stream": true for the Responses event stream. Each event carries a
sequence_number and is framed with both an SSE event: name and a data:
payload.
| Event | When |
|---|---|
response.created | The first upstream chunk arrives. |
response.in_progress | Immediately after. |
response.output_item.added | A message or function_call item opens. |
response.content_part.added | The text part of a message opens. |
response.output_text.delta | Each incremental text fragment. |
response.function_call_arguments.delta | Each incremental argument fragment. |
response.output_text.done | The message text is complete. |
response.content_part.done | The text part closes. |
response.function_call_arguments.done | A call’s arguments are complete. |
response.output_item.done | An item closes. |
response.completed / response.incomplete | The final response object, with usage. |
Statelessness
previous_response_id is rejected:
input — the same thing the SDK does for chat
completions.
For conversation storage with history, sharing, folders and search, use the
Workspace chat API, which is a control-plane surface.
Errors
The error envelope is identical to chat completions, anderror.param names the
Responses field you sent — not the chat parameter it was translated into: