429 the moment a ceiling is reached — before the provider is called.
Where limits come from
| Source | Windows | Set in |
|---|---|---|
Team rpm_limit / tpm_limit | minute | POST /admin/teams (organization Owner), or teams[] in gateway.yaml on a self-hosted install |
API key rpm_limit / tpm_limit | minute | Gateway → API Keys, or POST / PATCH /admin/keys |
| Rate-limit policies | minute, hour, day | Gateway → Policies → Rate limits, or /admin/gateway/policies/rate-limits |
0 always means unlimited.
How it works
- Every matching ceiling must allow the request. A request under a 100/min workspace rule, a 50/min team rule and a 10/min user-on-model rule has to satisfy all three. The narrowest scope can never raise a ceiling a broader scope set.
- Outermost first. When several refuse, you are told about the outermost one — organization before workspace, team, service account, user, API key.
- Cache hits still count as requests. Request ceilings are checked before the cache lookup; token ceilings are checked after it, so a hit is not charged tokens.
- Tokens are settled after the response. Admission charges the prompt; the completion is charged once known. A long answer can leave the counter empty, and the overspend is repaid out of the next window. A single prompt larger than the whole token ceiling can never fit and is refused with
tpm_exceeded. - Windows. With Redis the window slides: the previous interval is weighted by how much of it is still inside the trailing window, so capacity returns continuously instead of resetting on the minute. The in-memory backend uses a token bucket that holds the whole ceiling as burst and refills evenly across the window.
Rate-limit policy structure
Key fields
workspace files the rule under the workspace the request resolved to. organization governs every workspace and needs gateway.ratelimits.manage across the product; it cannot name an API key or service account, which belong to one workspace.The identity the rule governs. Name at most one; leave all empty to govern the whole scope. A
user_id must belong to the workspace.A second dimension beside the identity:
user + gpt-4o caps that pair only. It is the logical name of a model the workspace resolves, so re-pointing the model at another provider keeps the limit, and aliases resolve to it. It cannot be changed after the rule is created.Whether matching requests share one counter or get one each. A team rule with
applies_per: "user" gives every user in the team their own 60/min. Fixed once the rule exists.Six independent ceilings. Set at least one; a rule with all six at zero is refused. Each window has its own counter, so “60 a minute and 10,000 a day” are both enforced.
Common configurations
Team burst and daily cap
Team burst and daily cap
Protect an expensive model, per user
Protect an expensive model, per user
Per-key limit
Per-key limit
PATCH /admin/keys/{id} replaces the key’s limits wholesale: send every limit you want to keep, or an omitted one becomes 0 (unlimited).Team limits in gateway.yaml (self-hosted)
Team limits in gateway.yaml (self-hosted)
PUT /admin/gateway/policies/rate-limits/{id} (edits the ceilings and name; the model and applies_per cannot change, create a new rule instead), PATCH with {"enabled": false} to pause one, and DELETE to remove it. Pass ?workspace_id= when your credential reaches more than one workspace.
The refusal
rpm_exceeded for a request ceiling (any window), tpm_exceeded for a token ceiling.The rung that bound:
organization, workspace, team, service_account, user or api_key. The policy id and its values are never disclosed.Seconds: the length of the window that refused — 60, 3600 or 86400. It is the upper bound on how long room can take to return.
429 with Retry-After.
Backends
- Memory (default)
- Redis
Counters live in each gateway process. Correct for a single replica; with several, each replica enforces the full ceiling on its own, so the effective limit is multiplied by the replica count.
Policies are compiled into each replica’s memory, so evaluation adds no database query; only the counters are shared. A change in the console reaches every replica through the policy version bump. See Policies. Refusals are counted in the
manylayers_policy_decisions_total Prometheus metric (labels outcome, policy_type, code).Next steps
Budgets & cost tracking
Cap spend rather than throughput.
Policies
Evaluation order and the other policy kinds.
Access control
Who may edit rate-limit policies.