Every request the gateway serves is priced from the model’s configured rates and written to a usage record. Budgets read those same records, so the limit you set, the spend the console shows and the chargeback you export can never disagree.

How cost is computed

cost_usd = base_input_tokens  / 1000 × input_price_per_1k
         + cached_input_tokens / 1000 × cached_input_price_per_1k   (unset → input price)
         + cache_write_tokens  / 1000 × cache_write_price_per_1k    (unset → input price)
         + output_tokens       / 1000 × output_price_per_1k
base_input_tokens is the prompt minus the tokens the provider reported as read from or written to its prompt cache. Prices come from the model that served the request, so a fallback or an Auto Routing escalation is billed at the target’s rate, not the requested name’s. Models you register in Gateway → Providers carry input_price_per_1k and output_price_per_1k (set together). When you omit them they are filled from the published list price if the gateway knows one. Edit the row for negotiated rates; the stored price is what bills. See Providers.
gateway.yaml
models:
  - logical_name: claude-sonnet
    provider: anthropic
    upstream_url: https://api.anthropic.com
    upstream_model: claude-sonnet-4-5
    upstream_api_key: ${ANTHROPIC_API_KEY}
    input_price_per_1k: 0.003
    output_price_per_1k: 0.015
    cached_input_price_per_1k: 0.0003
    cache_write_price_per_1k: 0.00375
  - logical_name: dall-e-3
    upstream_url: https://api.openai.com
    upstream_api_key: ${OPENAI_API_KEY}
    media_prices:
      image: 0.04               # per image, × request "n" on generations
      speech_per_1k_chars: 0.015
      transcription: 0.006      # flat per transcription or translation
A model with no price is metered at $0, and every budget measured in money ignores its traffic. Price every model you put a budget on.
Media endpoints (/v1/images/*, /v1/audio/*) are priced from the model’s media_prices and return the amount in an X-ManyLayers-Cost response header when it is above zero. For token endpoints, read cost from the usage API or the request trace.

Three kinds of budget

BudgetMeasuresWindowRefusal
Team monthly_token_budgetTokensCalendar month (UTC)429, insufficient_quota, budget_exhausted
API key budget_usd_monthlyUSDbudget_reset_period402, insufficient_quota, key_budget_exceeded
Budget policyUSDperiod402, insufficient_quota, budget_exceeded
Every refusal carries X-ManyLayers-Limit naming the rung that bound. None carries Retry-After: waiting does not refill a budget. The team token budget is checked against the prompt as well: a request whose prompt would take the team past its allowance is refused.

Budget policy structure

Budget policies live in Gateway → Policies → Budgets.
{
  "name": "Support monthly",            // optional label
  "scope": "workspace",                 // workspace (default) | organization
  "team_id": "team_9",                  // at most one of: team_id, service_account_id, user_id, api_key_id
  "model_name": "",                     // one logical model, or "" for all (create-only)
  "period": "monthly",                  // daily | weekly | monthly (default) | quarterly | never
  "limit_amount": 500,                  // must be > 0, in USD
  "currency": "USD",                    // default USD
  "applies_per": "",                    // "" shared | user | model (create-only)
  "enforcement": "enforce",             // enforce (default) | soft | audit
  "alert_thresholds": [50, 80, 100],    // ascending percentages 1-100
  "alert_emails": ["finops@example.com"],
  "alert_slack_webhook": "https://hooks.slack.com/services/...",
  "enabled": true
}
A budget’s identity is its subject, model and period: creating a second budget with the same three returns 409.

Key fields

period
string
Windows are calendar-aligned in UTC: a day from midnight, a week from Monday, a month from the 1st, a quarter from 1 January/April/July/October. never is a lifetime total that nothing resets.
limit_amount
number
Compared with spend in USD. currency is stored with the budget but does not convert anything, so leave it USD.
enforcement
string
enforce refuses once the allowance is spent. soft refuses only when no other budget matching the same request still has room. audit never refuses; it records and alerts, which is how you introduce a budget against existing traffic.
alert_thresholds
int[]
Each threshold is notified once per period, the first time spend crosses it, including on a request that is then refused. Thresholds require at least one address in alert_emails or an alert_slack_webhook.
alert_slack_webhook
string
A Slack incoming-webhook URL on hooks.slack.com. It is stored encrypted and never returned; reads show only a masked alert_slack_hint. Omit it on update to keep the stored one, send "" to remove it.
applies_per
string
user or model gives each user or model its own allowance of limit_amount instead of one shared pool.
A budget that names a service account cannot be measured, because spend is attributed to the keys the account owns. It never refuses and shows no spend. Put the budget on the account’s API key (api_key_id) or on its team instead.

Common configurations

curl -X POST https://app.manylayers.io/admin/gateway/policies/budgets \
  -H "Authorization: Bearer ml_pat_..." -H "Content-Type: application/json" \
  -d '{"name": "Support monthly", "team_id": "team_9", "period": "monthly",
       "limit_amount": 500, "alert_thresholds": [50, 80, 100],
       "alert_emails": ["finops@example.com"]}'
{"name": "GPT-4o per user", "model_name": "gpt-4o", "applies_per": "user",
 "period": "daily", "limit_amount": 5}
curl -X POST https://app.manylayers.io/admin/keys \
  -H "Authorization: Bearer ml_pat_..." -H "Content-Type: application/json" \
  -d '{"team": "engineering", "workspace_id": "ws_123", "name": "ci",
       "budget_usd_monthly": 50, "budget_reset_period": "weekly", "rpm_limit": 100}'
Despite its name, budget_usd_monthly is the allowance for whichever budget_reset_period you choose: daily, weekly, monthly (default) or never. In the console this is the Credit limit and Reset limit every… fields on Gateway → API Keys.
teams:
  - name: engineering
    monthly_token_budget: 1000000   # 0 = unlimited
Or send monthly_token_budget to POST /admin/teams (organization Owner).

Alerts

Alerts go to the budget’s email addresses and, if set, its Slack webhook. Email is sent through Resend, so it needs suite.resend_api_key and suite.resend_from on the deployment, plus suite.public_url for the link to the budgets screen; without a Resend key, thresholds are tracked and Slack still works but no email is sent. Without Redis, each gateway replica sends its own copy of a milestone. Check delivery before you rely on it:
curl -X POST "https://app.manylayers.io/admin/gateway/policies/budgets/test-alert?workspace_id=ws_123" \
  -H "Authorization: Bearer ml_pat_..." -H "Content-Type: application/json" \
  -d '{"emails": ["finops@example.com"], "name": "Support monthly", "limit_amount": 500, "period": "monthly"}'
# {"sent": true, "recipients": 1, "email": {"sent": true, "recipients": 1}}
The body also accepts slack_webhook (or budget_id, to test a saved budget’s stored webhook) and model_name. At most 20 addresses; at least one email or a webhook is required. On failure sent is false and error says which channel refused and why. The budget API: GET/POST /admin/gateway/policies/budgets, GET/PUT/PATCH/DELETE .../budgets/{id}. Each budget is returned with the amount spent against it and the number of breaches (requests that reached it after it was spent) in the current window. Reading needs gateway.budgets.read; changing needs gateway.budgets.manage. PATCH takes {"enabled": bool}.

Usage reporting

EndpointReturns
GET /admin/usage?since=RFC3339usage[] rows of team_id, team_name, model, requests, prompt_tokens, completion_tokens, cost_usd. Default since: start of the month. Add group=user for chargeback per user (rows gain user_id, user_email) or group=key per API key (rows carry api_key_id, key_name and no model), and team_id= to narrow.
GET /admin/usage/timeseries?bucket=hour|day&since=RFC3339buckets[] of ts, requests, prompt_tokens, completion_tokens, cost_usd. Default bucket: hour. Default since: 24 hours ago for hour, start of the month for day.
Budgets are checked before each request against spend already recorded, so the request that crosses the line is served and the next one is refused. Spend is read from a short-lived counter seeded from the usage records (the in-memory counter refreshes about every 10 seconds; with Redis it is shared across replicas), so enforcement can briefly over-allow. Disabling a budget removes the limit; it is not a limit of zero.

Next steps

Rate limiting

Cap throughput per minute, hour or day.

Policies

Evaluation order and scopes.

Providers

Model entry fields, including pricing.

Auto Routing

Savings estimates per request.