The kinds
| Kind | Question it answers | Where to manage it |
|---|---|---|
| Access policies | May this identity use the gateway at all? | API only |
| Model restrictions | May it request this model? | Policies → Model restrictions |
| Token limits | Is this single request too large? | API only |
| Rate limits | Has it sent too much this minute, hour or day? | Policies → Rate limits |
| Budgets | Has it spent too much this period? | Policies → Budgets |
| Logging config | Are its request and response bodies recorded, and what is redacted from that copy? | Policies → Logging config |
Scope
Every policy names a scope, and the chain runs outermost to innermost:team_id, service_account_id, user_id or api_key_id; naming two is refused). Leaving the target unset governs everything in the scope. scope: "organization" governs every workspace of the product; it needs gateway.ratelimits.manage plus the permission for that kind of policy, held across the product, and it cannot name an API key or service account, which belong to one workspace.
Semantics worth knowing
Model restrictions. No policy means every model the workspace resolves. An allow-list permits only what it names; a deny-list refuses what it names; when both match a model, deny wins. Two allow-lists at different rungs intersect rather than combine. A restriction may name a model on every provider account (gpt-4o) or on one (openai-dev/gpt-4o). A virtual model is judged target by target: an allow-list that does not list it still lets it serve the targets the list does permit, while a deny-list that names it refuses it outright. Team-level models lists and virtual account allowlists are checked separately and also have to pass.
Access policies. A deny bans the identity it names, and must name a team, service account, user or API key. An allow is stronger than it looks: one enabled allow policy turns its scope into an allow-list, and every identity it does not name is refused. A deny always beats an allow. An optional endpoints list (exact request paths such as /v1/embeddings) confines the policy to those paths; empty means all.
Rate limits target a user, a team, a model, or a subject and a model together — user, team, model, user + model, team + model. Arun on gpt-4o-mini is one rule with both dimensions, not two rules: it caps that pair and neither Arun’s other models nor other people’s use of that model. A rule with no model named caps every model.
Each rule is its own ceiling on its own counter, and every rule that matches a request must allow it. A request under a 100/min workspace rule, a 50/min team rule and a 10/min user+model rule has to satisfy all three; the first to refuse, outermost first, is the one you are told about. Two rules on the same dimensions are two independent ceilings rather than one merged one. See Rate limiting.
The model is the logical name — what a request carries in its model field. A limit on enterprise-smart keeps applying when you re-point it from OpenAI to Azure. An alias resolves to the model it points at, so switching to an alias does not escape the limit.
Token limits cap one request — its input, its requested completion, or the total (input plus requested completion). They are not rate limits: tokens_per_minute on a rate limit governs consumption over time, these govern a single request and are checked before the provider is called. Zero means no limit, at least one must be set, and a total cannot be smaller than the input or output ceiling. A request that names no max_tokens (max_completion_tokens and max_output_tokens also count) is not refused by an output ceiling, because the gateway cannot know what a provider will generate.
Budgets are measured from the same usage records your invoice comes from, so a budget and a bill cannot disagree. See Budgets.
Logging config rules decide whether a scope’s request and response bodies are recorded on the audit log (log_requests) and what is redacted from the stored copy (redaction: PII, secrets, custom regex patterns). Redaction changes only what is stored; the model receives and returns the content unchanged.
Disabled is not zero. A disabled policy is the absence of a limit, not a limit of zero. It stays listed so it can be turned back on.
How a change reaches the gateway
500 policy evaluation failed rather than being reported as a limit.
Refusals
| Code | HTTP | type | Meaning |
|---|---|---|---|
access_denied | 403 | permission_error | An access policy refused this identity |
model_not_allowed | 403 | permission_error | A model restriction refused this model |
max_input_tokens_exceeded | 400 | invalid_request_error | The prompt is over a per-request ceiling |
max_output_tokens_exceeded | 400 | invalid_request_error | The requested completion is over a ceiling |
max_total_tokens_exceeded | 400 | invalid_request_error | Prompt plus requested completion is over a ceiling |
rpm_exceeded | 429 | rate_limit_error | A request ceiling (minute, hour or day) is exhausted |
tpm_exceeded | 429 | rate_limit_error | A token ceiling is exhausted, or a single prompt is larger than the ceiling |
budget_exceeded | 402 | insufficient_quota | A spend budget is exhausted |
key_budget_exceeded | 402 | insufficient_quota | The API key’s own budget is exhausted |
budget_exhausted | 429 | insufficient_quota | A team’s monthly token budget is exhausted |
X-ManyLayers-Limit naming the scope that bound (organization, workspace, team, service_account, user or api_key), so you know which ceiling to raise. Rate-limit refusals also carry Retry-After; budget refusals deliberately do not, because waiting does not refill a budget.
A refusal never names the policy that produced it. The scope is enough to act on, and the configuration of other scopes is not the caller’s business.
API
Console API athttps://app.manylayers.io. All routes are workspace-scoped; pass ?workspace_id= when your credential can reach more than one workspace. Every POST takes a name (optional for budgets) and the optional scope / target fields described above; PATCH takes {"enabled": bool}.
| Method | Path |
|---|---|
GET/POST | /admin/gateway/policies/rate-limits |
PUT/PATCH/DELETE | /admin/gateway/policies/rate-limits/{id} |
GET/POST | /admin/gateway/policies/budgets |
GET/PUT/PATCH/DELETE | /admin/gateway/policies/budgets/{id} |
POST | /admin/gateway/policies/budgets/test-alert |
GET/POST | /admin/gateway/policies/model-restrictions |
PUT/PATCH/DELETE | /admin/gateway/policies/model-restrictions/{id} |
GET/POST | /admin/gateway/policies/token-limits |
PUT/PATCH/DELETE | /admin/gateway/policies/token-limits/{id} |
GET/POST | /admin/gateway/policies/access |
PUT/PATCH/DELETE | /admin/gateway/policies/access/{id} |
GET/POST | /admin/gateway/policies/logging-configs |
PUT/PATCH/DELETE | /admin/gateway/policies/logging-configs/{id} |
GET | /admin/gateway/policies/subjects |
| Kind | Body fields (besides name, scope and one target) |
|---|---|
| Rate limit | model_name, applies_per, requests_per_minute/_hour/_day, tokens_per_minute/_hour/_day |
| Budget | model_name, period, limit_amount, currency, applies_per, enforcement, alert_thresholds, alert_emails, alert_slack_webhook |
| Model restriction | mode (allow or deny), models (at least one) |
| Token limit | max_input_tokens, max_output_tokens, max_total_tokens |
| Access policy | effect (allow or deny), endpoints |
| Logging config | log_requests, redaction (pii, pii_categories, secrets, patterns, replacement) |
applies_per are fixed once a policy exists: moving one changes which requests it governs and which counter it uses, which is a new policy rather than an edit. A duplicate name in the same scope returns 409.
Permissions
| Operation | Permission |
|---|---|
| Read rate limits, token limits | gateway.ratelimits.read |
| Change rate limits, token limits, access policies | gateway.ratelimits.manage |
| Read budgets | gateway.budgets.read |
| Change budgets | gateway.budgets.manage |
| Read model restrictions | gateway.models.read |
| Change model restrictions | gateway.models.manage |
| Read access policies | gateway.apikeys.read |
| Read logging configs | gateway.policies.read |
| Change logging configs | gateway.policies.manage |
gateway.ratelimits.manage and its own kind’s permission across the whole product, not just inside one workspace. See Access control.