How cost is computed
base_input_tokens is the prompt minus the tokens the provider reported as read from or written to its prompt cache. Prices come from the model that served the request, so a fallback or an Auto Routing escalation is billed at the target’s rate, not the requested name’s.
Models you register in Gateway → Providers carry input_price_per_1k and output_price_per_1k (set together). When you omit them they are filled from the published list price if the gateway knows one. Edit the row for negotiated rates; the stored price is what bills. See Providers.
Self-hosted: prices in gateway.yaml
Self-hosted: prices in gateway.yaml
gateway.yaml
/v1/images/*, /v1/audio/*) are priced from the model’s media_prices and return the amount in an X-ManyLayers-Cost response header when it is above zero. For token endpoints, read cost from the usage API or the request trace.
Three kinds of budget
| Budget | Measures | Window | Refusal |
|---|---|---|---|
Team monthly_token_budget | Tokens | Calendar month (UTC) | 429, insufficient_quota, budget_exhausted |
API key budget_usd_monthly | USD | budget_reset_period | 402, insufficient_quota, key_budget_exceeded |
| Budget policy | USD | period | 402, insufficient_quota, budget_exceeded |
X-ManyLayers-Limit naming the rung that bound. None carries Retry-After: waiting does not refill a budget. The team token budget is checked against the prompt as well: a request whose prompt would take the team past its allowance is refused.
Budget policy structure
Budget policies live in Gateway → Policies → Budgets.409.
Key fields
Windows are calendar-aligned in UTC: a day from midnight, a week from Monday, a month from the 1st, a quarter from 1 January/April/July/October.
never is a lifetime total that nothing resets.Compared with spend in USD.
currency is stored with the budget but does not convert anything, so leave it USD.enforce refuses once the allowance is spent. soft refuses only when no other budget matching the same request still has room. audit never refuses; it records and alerts, which is how you introduce a budget against existing traffic.Each threshold is notified once per period, the first time spend crosses it, including on a request that is then refused. Thresholds require at least one address in
alert_emails or an alert_slack_webhook.A Slack incoming-webhook URL on
hooks.slack.com. It is stored encrypted and never returned; reads show only a masked alert_slack_hint. Omit it on update to keep the stored one, send "" to remove it.user or model gives each user or model its own allowance of limit_amount instead of one shared pool.Common configurations
Team monthly spend cap with alerts
Team monthly spend cap with alerts
Per-user daily allowance on one model
Per-user daily allowance on one model
Per-key budget
Per-key budget
budget_usd_monthly is the allowance for whichever budget_reset_period you choose: daily, weekly, monthly (default) or never. In the console this is the Credit limit and Reset limit every… fields on Gateway → API Keys.Team token budget (self-hosted gateway.yaml)
Team token budget (self-hosted gateway.yaml)
monthly_token_budget to POST /admin/teams (organization Owner).Alerts
Alerts go to the budget’s email addresses and, if set, its Slack webhook. Email is sent through Resend, so it needssuite.resend_api_key and suite.resend_from on the deployment, plus suite.public_url for the link to the budgets screen; without a Resend key, thresholds are tracked and Slack still works but no email is sent. Without Redis, each gateway replica sends its own copy of a milestone. Check delivery before you rely on it:
slack_webhook (or budget_id, to test a saved budget’s stored webhook) and model_name. At most 20 addresses; at least one email or a webhook is required. On failure sent is false and error says which channel refused and why.
The budget API: GET/POST /admin/gateway/policies/budgets, GET/PUT/PATCH/DELETE .../budgets/{id}. Each budget is returned with the amount spent against it and the number of breaches (requests that reached it after it was spent) in the current window. Reading needs gateway.budgets.read; changing needs gateway.budgets.manage. PATCH takes {"enabled": bool}.
Usage reporting
| Endpoint | Returns |
|---|---|
GET /admin/usage?since=RFC3339 | usage[] rows of team_id, team_name, model, requests, prompt_tokens, completion_tokens, cost_usd. Default since: start of the month. Add group=user for chargeback per user (rows gain user_id, user_email) or group=key per API key (rows carry api_key_id, key_name and no model), and team_id= to narrow. |
GET /admin/usage/timeseries?bucket=hour|day&since=RFC3339 | buckets[] of ts, requests, prompt_tokens, completion_tokens, cost_usd. Default bucket: hour. Default since: 24 hours ago for hour, start of the month for day. |
Budgets are checked before each request against spend already recorded, so the request that crosses the line is served and the next one is refused. Spend is read from a short-lived counter seeded from the usage records (the in-memory counter refreshes about every 10 seconds; with Redis it is shared across replicas), so enforcement can briefly over-allow. Disabling a budget removes the limit; it is not a limit of zero.
Next steps
Rate limiting
Cap throughput per minute, hour or day.
Policies
Evaluation order and scopes.
Providers
Model entry fields, including pricing.
Auto Routing
Savings estimates per request.