This guide takes you from a new account to a first chat completion. You connect one provider, issue a credential, and call the gateway with the OpenAI SDK you already use. The hosted service answers on one host, app.manylayers.io. A self-hosted stack runs the two services on separate ports:
WhatHosted addressSelf-hosted default
Gateway (the /v1 data plane)https://app.manylayers.io/v1http://localhost:8180/v1
Console and admin API (/ui/, /auth/*, /admin/*, /w/*)https://app.manylayers.iohttp://localhost:8190
Running your own stack? See Self-hosting and replace the hosted addresses below with yours. For a local run from a source checkout:
make env-files    # seed infra/docker/.env.dev from its template
make dev-migrate  # apply database migrations; no "up" target migrates
make dev-up       # postgres, redis, gateway, workspace, deployer
1

Check the gateway is up

curl https://app.manylayers.io/health   # {"status":"ok"}
curl https://app.manylayers.io/ready    # {"status":"ready","checks":{...}}
/health (also /healthz) is liveness and checks nothing. /ready (also /readyz) checks Postgres, and Redis when configured.
2

Add a provider and models

Sign in at https://app.manylayers.io/ui/, open Providers in the Gateway console, choose Add provider, pick a type (for example openai) and a name, then add a credential (a pasted key, or a reference such as ${OPENAI_API_KEY}) and enable the models you want to serve.
Console providers serve /v1/chat/completions, /v1/responses and /v1/messages for keys bound to that workspace, and also /v1/embeddings, /v1/rerank and speech. See Supported APIs for the rest.
3

Create a credential

Choose the credential that fits. See Authentication for all types.
  • API key (ml-...): in the Gateway console open API Keys and choose New key, or call POST /admin/keys. Pass workspace_id to bind the key to a workspace’s console providers; leave it out for a team-scoped key that resolves models from gateway.yaml.
  • Personal access token (ml_pat_...): in the Gateway console open Access, then Personal Access Tokens. A PAT is issued only from a signed-in console session, acts as you, belongs to one workspace, and every PAT needs an expiry date. Use it to try the gateway yourself.
curl https://app.manylayers.io/admin/keys \
  -H "Authorization: Bearer <admin-credential>" -H "Content-Type: application/json" \
  -d '{"team": "platform", "name": "my-service", "workspace_id": "<workspace-id>", "expires_at": "2027-01-01T00:00:00Z"}'
The response contains the plaintext key exactly once; only its hash is stored. The team must already exist.
4

Send your first request

Point any OpenAI client at https://app.manylayers.io/v1:
curl https://app.manylayers.io/v1/chat/completions \
  -H "Authorization: Bearer ml-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Hello from ManyLayers"}]
  }'
The response carries an X-ManyLayers-Trace-Id header (unless the routing config enables strict_openai_compliance). Use it to find the request under Request Traces in the console. To see the models your credential can reach, call GET /v1/models. You can also try a model without writing code in the console’s Playground.
5

Try routing and caching

  • Routing: create a gateway config (fallback, load balancing, canary and others) under Routing in the console, then select it per request with X-ManyLayers-Config: <name>, or set it as the team default. See Routing.
  • Caching: the cache must be enabled on the deployment (cache.enabled) and for the team. Requests with temperature: 0 are then cached, and the X-ManyLayers-Cache response header reports hit, semantic, miss, bypass or disabled. See Caching.
If a model name resolves to nothing for your credential, the gateway returns 404 model_not_found. A key bound to a workspace resolves chat models only from that workspace’s console providers, never from gateway.yaml. A bad or missing key returns 401 invalid_api_key.

Next steps

Making requests

Streaming, embeddings, the Responses API and gateway headers.

Providers

Configuration for each provider type.

Rate limits

Put limits and budgets on your new key.