Rate limits stop one caller from exhausting a shared provider quota or running up spend in a burst. The gateway counts requests and tokens per identity and model, and refuses with 429 the moment a ceiling is reached — before the provider is called.

Where limits come from

SourceWindowsSet in
Team rpm_limit / tpm_limitminutePOST /admin/teams (organization Owner), or teams[] in gateway.yaml on a self-hosted install
API key rpm_limit / tpm_limitminuteGateway → API Keys, or POST / PATCH /admin/keys
Rate-limit policiesminute, hour, dayGateway → Policies → Rate limits, or /admin/gateway/policies/rate-limits
All three are enforced together. 0 always means unlimited.

How it works

  • Every matching ceiling must allow the request. A request under a 100/min workspace rule, a 50/min team rule and a 10/min user-on-model rule has to satisfy all three. The narrowest scope can never raise a ceiling a broader scope set.
  • Outermost first. When several refuse, you are told about the outermost one — organization before workspace, team, service account, user, API key.
  • Cache hits still count as requests. Request ceilings are checked before the cache lookup; token ceilings are checked after it, so a hit is not charged tokens.
  • Tokens are settled after the response. Admission charges the prompt; the completion is charged once known. A long answer can leave the counter empty, and the overspend is repaid out of the next window. A single prompt larger than the whole token ceiling can never fit and is refused with tpm_exceeded.
  • Windows. With Redis the window slides: the previous interval is weighted by how much of it is still inside the trailing window, so capacity returns continuously instead of resetting on the minute. The in-memory backend uses a token bucket that holds the whole ceiling as burst and refills evenly across the window.

Rate-limit policy structure

{
  "name": "Support team on gpt-4o",
  "scope": "workspace",            // workspace (default) | organization
  "team_id": "team_9",             // at most one of: team_id, service_account_id, user_id, api_key_id
  "model_name": "gpt-4o",          // logical model; "" = every model
  "applies_per": "user",           // "" shared | user | model | "user,model": one counter each
  "requests_per_minute": 60,
  "requests_per_hour": 0,
  "requests_per_day": 10000,
  "tokens_per_minute": 200000,
  "tokens_per_hour": 0,
  "tokens_per_day": 0,
  "enabled": true
}

Key fields

scope
string
workspace files the rule under the workspace the request resolved to. organization governs every workspace and needs gateway.ratelimits.manage across the product; it cannot name an API key or service account, which belong to one workspace.
team_id | service_account_id | user_id | api_key_id
string
The identity the rule governs. Name at most one; leave all empty to govern the whole scope. A user_id must belong to the workspace.
model_name
string
A second dimension beside the identity: user + gpt-4o caps that pair only. It is the logical name of a model the workspace resolves, so re-pointing the model at another provider keeps the limit, and aliases resolve to it. It cannot be changed after the rule is created.
applies_per
string
Whether matching requests share one counter or get one each. A team rule with applies_per: "user" gives every user in the team their own 60/min. Fixed once the rule exists.
requests_per_* | tokens_per_*
integer
Six independent ceilings. Set at least one; a rule with all six at zero is refused. Each window has its own counter, so “60 a minute and 10,000 a day” are both enforced.

Common configurations

curl -X POST https://app.manylayers.io/admin/gateway/policies/rate-limits \
  -H "Authorization: Bearer ml_pat_..." -H "Content-Type: application/json" \
  -d '{"name": "Support burst", "team_id": "team_9",
       "requests_per_minute": 60, "requests_per_day": 10000}'
{"name": "o1 per user", "model_name": "o1", "applies_per": "user",
 "requests_per_minute": 5, "tokens_per_day": 500000}
curl -X PATCH "https://app.manylayers.io/admin/keys/key_123?workspace_id=ws_123" \
  -H "Authorization: Bearer ml_pat_..." -H "Content-Type: application/json" \
  -d '{"rpm_limit": 100, "tpm_limit": 20000, "budget_usd_monthly": 50, "budget_reset_period": "monthly"}'
PATCH /admin/keys/{id} replaces the key’s limits wholesale: send every limit you want to keep, or an omitted one becomes 0 (unlimited).
teams:
  - name: engineering
    rpm_limit: 60
    tpm_limit: 100000
    models: ["gpt-4o", "gpt-4o-mini"]
Manage existing rules with PUT /admin/gateway/policies/rate-limits/{id} (edits the ceilings and name; the model and applies_per cannot change, create a new rule instead), PATCH with {"enabled": false} to pause one, and DELETE to remove it. Pass ?workspace_id= when your credential reaches more than one workspace.

The refusal

HTTP/1.1 429 Too Many Requests
X-ManyLayers-Limit: team
Retry-After: 60

{"error": {"message": "Team rate limit exceeded", "type": "rate_limit_error", "code": "rpm_exceeded"}}
error.code
string
rpm_exceeded for a request ceiling (any window), tpm_exceeded for a token ceiling.
X-ManyLayers-Limit
string
The rung that bound: organization, workspace, team, service_account, user or api_key. The policy id and its values are never disclosed.
Retry-After
integer
Seconds: the length of the window that refused — 60, 3600 or 86400. It is the upper bound on how long room can take to return.
When a virtual model has several targets and every one that could serve is at its limit, the caller gets the same 429 with Retry-After.

Backends

Counters live in each gateway process. Correct for a single replica; with several, each replica enforces the full ceiling on its own, so the effective limit is multiplied by the replica count.
The Redis limiter fails open: if Redis errors, requests are admitted rather than all refused. Monitor Redis if your limits protect a hard provider quota.
Policies are compiled into each replica’s memory, so evaluation adds no database query; only the counters are shared. A change in the console reaches every replica through the policy version bump. See Policies. Refusals are counted in the manylayers_policy_decisions_total Prometheus metric (labels outcome, policy_type, code).

Next steps

Budgets & cost tracking

Cap spend rather than throughput.

Policies

Evaluation order and the other policy kinds.

Access control

Who may edit rate-limit policies.