Policies decide whether an AI request is allowed. They are configured per workspace in Gateway → Policies, stored in PostgreSQL, and enforced by every gateway replica from a compiled snapshot held in local memory. A policy is not RBAC. Your role in the console decides whether you may edit a policy; a policy decides whether a request is served.

The kinds

KindQuestion it answersWhere to manage it
Access policiesMay this identity use the gateway at all?API only
Model restrictionsMay it request this model?Policies → Model restrictions
Token limitsIs this single request too large?API only
Rate limitsHas it sent too much this minute, hour or day?Policies → Rate limits
BudgetsHas it spent too much this period?Policies → Budgets
Logging configAre its request and response bodies recorded, and what is redacted from that copy?Policies → Logging config
The console’s Policies page has four tabs: Rate limits, Budgets, Model restrictions and Logging config. Access policies and token limits have no tab, but they are enforced like the rest and are written through the API. The first five are evaluated in the order shown above, cheapest and most decisive first. A request refused by an access policy never consumes a rate-limit token, and a request refused by any of them never reaches a provider. Two of the checks run after the prompt has been counted, which happens after PII redaction, guardrails and the cache lookup: the input and total token ceilings, and tokens-per-minute limits and token budgets. A cache hit consumes no upstream tokens, so it is not charged against a token rate limit.

Scope

Every policy names a scope, and the chain runs outermost to innermost:
Organization → Workspace → Team → Service account → User → API key
Every level that sets a limit is enforced and the most restrictive one binds. A key allowed 5000 requests a minute inside a workspace capped at 1000 gets 1000 — a narrower scope can never raise a ceiling a broader one set. A policy targets at most one identity (team_id, service_account_id, user_id or api_key_id; naming two is refused). Leaving the target unset governs everything in the scope. scope: "organization" governs every workspace of the product; it needs gateway.ratelimits.manage plus the permission for that kind of policy, held across the product, and it cannot name an API key or service account, which belong to one workspace.

Semantics worth knowing

Model restrictions. No policy means every model the workspace resolves. An allow-list permits only what it names; a deny-list refuses what it names; when both match a model, deny wins. Two allow-lists at different rungs intersect rather than combine. A restriction may name a model on every provider account (gpt-4o) or on one (openai-dev/gpt-4o). A virtual model is judged target by target: an allow-list that does not list it still lets it serve the targets the list does permit, while a deny-list that names it refuses it outright. Team-level models lists and virtual account allowlists are checked separately and also have to pass. Access policies. A deny bans the identity it names, and must name a team, service account, user or API key. An allow is stronger than it looks: one enabled allow policy turns its scope into an allow-list, and every identity it does not name is refused. A deny always beats an allow. An optional endpoints list (exact request paths such as /v1/embeddings) confines the policy to those paths; empty means all. Rate limits target a user, a team, a model, or a subject and a model together — user, team, model, user + model, team + model. Arun on gpt-4o-mini is one rule with both dimensions, not two rules: it caps that pair and neither Arun’s other models nor other people’s use of that model. A rule with no model named caps every model. Each rule is its own ceiling on its own counter, and every rule that matches a request must allow it. A request under a 100/min workspace rule, a 50/min team rule and a 10/min user+model rule has to satisfy all three; the first to refuse, outermost first, is the one you are told about. Two rules on the same dimensions are two independent ceilings rather than one merged one. See Rate limiting. The model is the logical name — what a request carries in its model field. A limit on enterprise-smart keeps applying when you re-point it from OpenAI to Azure. An alias resolves to the model it points at, so switching to an alias does not escape the limit. Token limits cap one request — its input, its requested completion, or the total (input plus requested completion). They are not rate limits: tokens_per_minute on a rate limit governs consumption over time, these govern a single request and are checked before the provider is called. Zero means no limit, at least one must be set, and a total cannot be smaller than the input or output ceiling. A request that names no max_tokens (max_completion_tokens and max_output_tokens also count) is not refused by an output ceiling, because the gateway cannot know what a provider will generate. Budgets are measured from the same usage records your invoice comes from, so a budget and a bill cannot disagree. See Budgets. Logging config rules decide whether a scope’s request and response bodies are recorded on the audit log (log_requests) and what is redacted from the stored copy (redaction: PII, secrets, custom regex patterns). Redaction changes only what is stored; the model receives and returns the content unchanged. Disabled is not zero. A disabled policy is the absence of a limit, not a limit of zero. It stays listed so it can be turned back on.

How a change reaches the gateway

Console → PostgreSQL (source of truth) → version bump
        → compile snapshot → local memory
        → Redis snapshot + update event (ml:policy:updated)
        → every other replica loads it and swaps atomically
Each replica holds its own compiled snapshot in memory, so evaluating a policy costs a map lookup — no database query on the request path. Redis is used for two things only: distributing the snapshot, and the rate-limit and spend counters that have to be shared across replicas to mean anything. Snapshots are versioned and monotonic: a replica ignores a late or duplicate notification and loads the newest version. A replica that misses a notification entirely is repaired by a reconciler that compares versions once a minute. Redis is optional. Without it, the same update events travel on PostgreSQL’s own notification channel and snapshots are read from the database directly. If a replica cannot reload, it keeps enforcing the snapshot it already holds. A replica that has never loaded a snapshot and cannot reach the database falls back to the per-team and per-key limit columns only, and logs a warning. If a rate-limit or spend check itself cannot be evaluated, the request fails with 500 policy evaluation failed rather than being reported as a limit.

Refusals

CodeHTTPtypeMeaning
access_denied403permission_errorAn access policy refused this identity
model_not_allowed403permission_errorA model restriction refused this model
max_input_tokens_exceeded400invalid_request_errorThe prompt is over a per-request ceiling
max_output_tokens_exceeded400invalid_request_errorThe requested completion is over a ceiling
max_total_tokens_exceeded400invalid_request_errorPrompt plus requested completion is over a ceiling
rpm_exceeded429rate_limit_errorA request ceiling (minute, hour or day) is exhausted
tpm_exceeded429rate_limit_errorA token ceiling is exhausted, or a single prompt is larger than the ceiling
budget_exceeded402insufficient_quotaA spend budget is exhausted
key_budget_exceeded402insufficient_quotaThe API key’s own budget is exhausted
budget_exhausted429insufficient_quotaA team’s monthly token budget is exhausted
Every refusal carries X-ManyLayers-Limit naming the scope that bound (organization, workspace, team, service_account, user or api_key), so you know which ceiling to raise. Rate-limit refusals also carry Retry-After; budget refusals deliberately do not, because waiting does not refill a budget. A refusal never names the policy that produced it. The scope is enough to act on, and the configuration of other scopes is not the caller’s business.

API

Console API at https://app.manylayers.io. All routes are workspace-scoped; pass ?workspace_id= when your credential can reach more than one workspace. Every POST takes a name (optional for budgets) and the optional scope / target fields described above; PATCH takes {"enabled": bool}.
MethodPath
GET/POST/admin/gateway/policies/rate-limits
PUT/PATCH/DELETE/admin/gateway/policies/rate-limits/{id}
GET/POST/admin/gateway/policies/budgets
GET/PUT/PATCH/DELETE/admin/gateway/policies/budgets/{id}
POST/admin/gateway/policies/budgets/test-alert
GET/POST/admin/gateway/policies/model-restrictions
PUT/PATCH/DELETE/admin/gateway/policies/model-restrictions/{id}
GET/POST/admin/gateway/policies/token-limits
PUT/PATCH/DELETE/admin/gateway/policies/token-limits/{id}
GET/POST/admin/gateway/policies/access
PUT/PATCH/DELETE/admin/gateway/policies/access/{id}
GET/POST/admin/gateway/policies/logging-configs
PUT/PATCH/DELETE/admin/gateway/policies/logging-configs/{id}
GET/admin/gateway/policies/subjects
KindBody fields (besides name, scope and one target)
Rate limitmodel_name, applies_per, requests_per_minute/_hour/_day, tokens_per_minute/_hour/_day
Budgetmodel_name, period, limit_amount, currency, applies_per, enforcement, alert_thresholds, alert_emails, alert_slack_webhook
Model restrictionmode (allow or deny), models (at least one)
Token limitmax_input_tokens, max_output_tokens, max_total_tokens
Access policyeffect (allow or deny), endpoints
Logging configlog_requests, redaction (pii, pii_categories, secrets, patterns, replacement)
The scope, target and (for rate limits and budgets) the model and applies_per are fixed once a policy exists: moving one changes which requests it governs and which counter it uses, which is a new policy rather than an edit. A duplicate name in the same scope returns 409.

Permissions

OperationPermission
Read rate limits, token limitsgateway.ratelimits.read
Change rate limits, token limits, access policiesgateway.ratelimits.manage
Read budgetsgateway.budgets.read
Change budgetsgateway.budgets.manage
Read model restrictionsgateway.models.read
Change model restrictionsgateway.models.manage
Read access policiesgateway.apikeys.read
Read logging configsgateway.policies.read
Change logging configsgateway.policies.manage
An organization-wide policy needs gateway.ratelimits.manage and its own kind’s permission across the whole product, not just inside one workspace. See Access control.