Async requests let you submit a completion, get a job id back immediately, and collect the result later. Use them for long generations, bulk offline work, or callers that can’t hold a connection open.

How it works

  1. Submit. The gateway checks that model exists in the configured catalog and that your team may use it, then stores the job and returns 202.
  2. Run. A worker claims the job and replays it through the normal inference pipeline as the submitting key’s team — policies, rate limits, budgets, firewall, PII redaction, guardrails, cache, metering and audit all apply.
  3. Collect. Poll the job, or let the gateway POST the result to your webhook_url.

Configuration structure

Self-hosted installs configure the worker pool and queue in gateway.yaml:
async:
  workers: 2                                       # worker goroutines per gateway (default 2)
  webhook_signing_secret: ${ASYNC_WEBHOOK_SECRET}   # HMAC key; empty = unsigned webhooks

queue:
  backend: db          # db (default, claim-based polling) | kafka
  kafka:
    brokers: ["kafka:9092"]
    prefix: "prod."    # optional topic prefix

Key fields

  • async.workers: Number of concurrent job runners on each gateway replica. Unset or 0 means 2.
  • async.webhook_signing_secret: When set, every webhook carries an X-ManyLayers-Signature header.
  • queue.backend: How workers learn about new jobs. db needs no extra infrastructure; kafka suits multi-replica deployments.

Submit a job

Endpoints: POST /v1/async/chat/completions and POST /v1/async/completions. The body is a normal chat or text completion, plus an optional webhook_url.
curl https://app.manylayers.io/v1/async/chat/completions \
  -H "Authorization: Bearer $ML_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Summarize this report..."}],
    "webhook_url": "https://app.example.com/hooks/manylayers"
  }'
Response (202 Accepted):
{"id": "6f1c2a9e-0b7d-4d1e-9c4a-2f8e5b1d7a33", "status": "queued"}
webhook_url must be an absolute http or https URL (otherwise 400 bad_webhook_url); it is removed before the request reaches the provider. stream is forced to false. The body is limited to the server’s maximum request size (413 body_too_large).

Poll for the result

GET /v1/async/jobs/{id} returns the job. You can only read jobs that belong to your team; anything else is 404 job_not_found.
{
  "id": "6f1c2a9e-0b7d-4d1e-9c4a-2f8e5b1d7a33",
  "team_id": "...",
  "api_key_id": "...",
  "endpoint": "/v1/chat/completions",
  "model": "gpt-4o-mini",
  "status": "succeeded",
  "response_status": 200,
  "response_body": "{\"id\":\"chatcmpl-...\",\"choices\":[...]}",
  "webhook_url": "https://app.example.com/hooks/manylayers",
  "webhook_delivered": true,
  "created_at": "2026-10-06T09:12:03.104000000Z",
  "started_at": "2026-10-06T09:12:03.211000000Z",
  "finished_at": "2026-10-06T09:12:09.874000000Z"
}
status moves from queued to running to succeeded or failed. response_body is the gateway’s answer as a JSON string — parse it again to get the completion. A job fails when the replayed request returns a non-2xx status — a policy refusal, for example; response_status and response_body hold the gateway’s error — or when its key was disabled or deleted after submission, in which case error says so.

Webhook delivery

When a job finishes, the gateway POSTs this body to webhook_url:
{
  "id": "6f1c2a9e-0b7d-4d1e-9c4a-2f8e5b1d7a33",
  "status": "succeeded",
  "response_status": 200,
  "response": {"id": "chatcmpl-...", "choices": [...]}
}
response is the completion (or the gateway’s error) as a JSON object; if a failed job’s body isn’t JSON, an error string is sent instead. Any 2xx reply marks the webhook delivered. Otherwise the gateway retries up to 3 attempts in total, waiting 1 s then 2 s, with a 10 s timeout per attempt.

Verify the signature

With async.webhook_signing_secret set, the header is X-ManyLayers-Signature: sha256=<hex>, where <hex> is the HMAC-SHA256 of the raw request body keyed with the secret.
import hmac, hashlib

def verify(secret: str, body: bytes, header: str) -> bool:
    expected = "sha256=" + hmac.new(secret.encode(), body, hashlib.sha256).hexdigest()
    return hmac.compare_digest(expected, header or "")
Compute the HMAC over the bytes you received, before any JSON parsing.

Behaviour notes

Submit jobs with an API key, personal access token or virtual account token. The worker re-authenticates the job by its key when it runs, so a job submitted with a console session or OIDC token has no key to run as and fails.
Only the body is stored. Request headers such as X-ManyLayers-Config, X-ManyLayers-Metadata or X-ManyLayers-Guardrails aren’t replayed; your team’s default routing config and your organization’s guardrail rules still apply.
The model must exist in gateway.yaml. Models from console-registered workspace providers and virtual models are refused at submit time with 404 model_not_found.

Next steps

Endpoints

Every path the gateway serves.

Policies

Limits and budgets applied when jobs run.

Making Requests

The synchronous equivalent.