https://app.manylayers.io/v1, plus an Anthropic-compatible /v1/messages. Use this page to check whether an endpoint exists and which providers can serve it.
On the hosted service, /v1 and the console’s control-plane routes (/ui, /auth/*, /admin/*, /w/*) are both served from https://app.manylayers.io. A self-hosted stack runs them as separate services (Gateway and Workspace).
All /v1 endpoints require Authorization: Bearer <token>: an API key (ml-...), personal access token (ml_pat_...), virtual account token (ml_vat_...) or an OIDC JWT. See Authentication and Making Requests. The caller’s organization must hold the AI Gateway product, otherwise /v1 answers 403 product_not_entitled.
Where models come from
A model is resolved from one of two sources (see Model Resolution):- Workspace providers registered in the console, for callers bound to a workspace.
models:ingateway.yaml, the configured catalog.
Which source an endpoint uses:
- Workspace providers only (for a workspace-bound credential):
/v1/chat/completions,/v1/systemone, and/v1/responsesand/v1/messages, which run as chat completions. - Workspace providers first, then
gateway.yaml:/v1/embeddings,/v1/rerank,/v1/audio/speechand/v1/audio/transcriptions. A workspace model of the wrong type (a chat model on/v1/embeddings) is refused. gateway.yamlonly:/v1/completions,/v1/images/*,/v1/audio/translations,/v1/moderations, files, fine-tuning, realtime and async.
Chat & completions
| Method | Path | Description |
|---|---|---|
POST | /v1/chat/completions | Chat completions on every provider. Set "stream": true for Server-Sent Events. |
POST | /v1/responses | OpenAI Responses API, stateless. Translated onto chat completions, so it works on every provider. See Responses. |
POST | /v1/messages | Anthropic Messages API. Translated onto chat completions, so it works on any provider, not only Anthropic. See Anthropic Messages API. |
POST | /v1/completions | Legacy text completions. |
GET | /v1/models | Models your credential may use, aliases and virtual models included. |
GET | /v1/models/{model} | One model. A model outside your policy returns 404. |
Embeddings
| Method | Path | Description |
|---|---|---|
POST | /v1/embeddings | Text embeddings. OpenAI-compatible providers are passed through; Azure is re-addressed to the deployment; Gemini, Vertex, Cohere and Bedrock are translated. Bedrock inputs are embedded one call per input and merged into one response. Anthropic has no embeddings. |
Realtime
| Method | Path | Description |
|---|---|---|
GET | /v1/realtime?model=<model> | WebSocket tunnel to OpenAI or Azure OpenAI realtime. See Realtime API. |
Images & audio
| Method | Path | Providers | Description |
|---|---|---|---|
POST | /v1/images/generations | openai, azure | Generate images |
POST | /v1/images/edits | openai, azure | Edit an image with a prompt and optional mask (multipart) |
POST | /v1/images/variations | openai, azure | Variations of an image (multipart) |
POST | /v1/audio/speech | openai, azure, elevenlabs, deepgram, cartesia, smallest | Text-to-speech |
POST | /v1/audio/transcriptions | openai, azure, elevenlabs, deepgram, cartesia, smallest | Speech-to-text (multipart) |
POST | /v1/audio/translations | openai, azure | Translate audio to English text (multipart) |
400 unsupported_provider. Responses carry X-ManyLayers-Cost when the model has a price configured. These endpoints skip token counting, PII redaction, the cache, guardrails and budget checks; access policies, model restrictions and rate limits still apply.
Moderations & rerank
| Method | Path | Description |
|---|---|---|
POST | /v1/moderations | Passed through for openai and azure models. Without a model, or for other providers, a built-in moderation check answers. |
POST | /v1/rerank | Translated to Cohere’s rerank API for cohere; passed through for openai, azure and OpenAI-compatible providers. |
Files & fine-tuning
openai and azure only. Name the model in the X-ManyLayers-Model header or the ?model= query parameter. POST /v1/fine_tuning/jobs reads model from its JSON body instead; without any of these the call returns 400 missing_model.
| Method | Path | Description |
|---|---|---|
POST | /v1/files | Upload a file |
GET | /v1/files | List files |
GET | /v1/files/{id} | File metadata |
GET | /v1/files/{id}/content | Download file content |
DELETE | /v1/files/{id} | Delete a file |
POST | /v1/fine_tuning/jobs | Create a fine-tuning job |
GET | /v1/fine_tuning/jobs | List jobs |
GET | /v1/fine_tuning/jobs/{id} | Job status |
POST | /v1/fine_tuning/jobs/{id}/cancel | Cancel a job |
Async inference
| Method | Path | Description |
|---|---|---|
POST | /v1/async/chat/completions | Queue a chat completion; returns 202 with a job id |
POST | /v1/async/completions | Queue a legacy completion |
GET | /v1/async/jobs/{id} | Job status and result |
System One
| Method | Path | Description |
|---|---|---|
POST | /v1/systemone | TypeSafe System One call: {model, state, questions} in, {model, answers, usage} out. Only TypeSafe decision models; other providers return 400 unsupported_provider. |
Knowledge bases & document sets
| Method | Path | Description |
|---|---|---|
POST | /v1/kb | Create a knowledge base |
GET | /v1/kb | List knowledge bases |
DELETE | /v1/kb/{id} | Delete a knowledge base |
GET | /v1/kb/{id}/documents | List documents |
POST | /v1/kb/{id}/documents | Add a document |
DELETE | /v1/kb/{id}/documents/{docID} | Delete a document |
POST | /v1/kb/{id}/query | Vector search over a knowledge base |
GET | /v1/docsets | List document sets |
POST | /v1/docsets | Create a document set |
GET | /v1/docsets/{id} | Get a document set |
PUT | /v1/docsets/{id} | Update a document set |
DELETE | /v1/docsets/{id} | Delete a document set |
Service endpoints
Every service (Gateway and Workspace) serves these without a credential, outside/v1: GET /health and /healthz (liveness), /ready and /readyz (readiness, which also checks dependencies), /version (build stamp), and /metrics (Prometheus).
Unsupported endpoints
These OpenAI surfaces are refused with404 and code: "unsupported_endpoint", for every method and sub-path. The message names what to use instead.
| Path | Use instead |
|---|---|
/v1/assistants | Workspace Agents, or /v1/chat/completions with tools |
/v1/threads | Workspace Agents, or chat completions with the conversation in messages |
/v1/conversations | /v1/responses with the conversation in input |
/v1/vector_stores | /v1/kb and /v1/docsets |
/v1/containers | Workspace Agents |
/v1/uploads | /v1/files (single multipart upload) |
/v1/evals | The evaluation suite under /admin/evals on the Workspace service |
/v1/organization | The Workspace admin API under /admin |
/v1 path — including /v1/responses/{id} — also returns 404 unsupported_endpoint. A known path called with the wrong method returns 405 method_not_allowed. Both use the standard OpenAI error envelope.
Next steps
Making Requests
Configure your SDK and send a first request.
Request & Response Headers
Headers that steer and describe each request.
Model Resolution
How a model name becomes a provider deployment.
OpenAI Compatibility
Parameter support per provider.