GET /metrics. Scrape the gateway replicas for request, token, policy, guardrail, cache and upstream health series; scrape Workspace and Deployer for process health.
How it works
/metricsis served by every service, next to/health,/healthz,/ready,/readyzand/version. Default ports are gateway:8180, workspace:8190, deployer:8200. It needs no authentication — restrict it at the network or ingress level.- Series are kept per process and cover every team and organization that process serves, so scrape them from a deployment you operate. Per-organization views belong in the console’s Analytics and Request Traces.
- Query with
sum by (...)to aggregate replicas. - The default Go runtime (
go_*) and process (process_*) collectors are included. - Request-path series are emitted by the service that handled the request — the gateway for all
/v1traffic.
prometheus.yml
Metric reference
Requests and tokens
| Metric | Type | Labels | Meaning |
|---|---|---|---|
manylayers_requests_total | counter | team, model, status | Completed requests. model is the model that served; status is the HTTP status code, e.g. 200, 429. |
manylayers_request_duration_seconds | histogram | model | End-to-end request duration. Buckets 0.05 s doubling to ~102 s. |
manylayers_tokens_total | counter | team, model, direction | Tokens processed. direction is input or output. |
OpenTelemetry GenAI conventions
These follow the OpenTelemetry GenAI metric conventions, so generic LLM dashboards work unchanged. All four carrygen_ai_operation_name (chat, embeddings, generate_content for /v1/responses), gen_ai_system (the serving provider type, or unknown), gen_ai_request_model and gen_ai_response_model. A failover shows up as different request and response models.
| Metric | Type | Extra labels | Meaning |
|---|---|---|---|
gen_ai_client_token_usage | counter | gen_ai_token_type | Tokens by type: input, output, cache_read, cache_creation, reasoning. The last three are subsets of the first two — do not sum all types. |
gen_ai_client_operation_duration_seconds | histogram | — | End-to-end duration of the operation. |
gen_ai_server_time_to_first_token_seconds | histogram | — | Time to first output token. Streamed responses only. |
gen_ai_server_time_per_output_token_seconds | histogram | — | Generation time per output token, excluding the wait for the first. |
Policy, firewall, guardrails and cache
| Metric | Type | Labels | Meaning |
|---|---|---|---|
manylayers_policy_decisions_total | counter | outcome, policy_type, code | Policy evaluations. outcome is allow or deny. On deny, policy_type is access, model_restriction, token_limit, rate_limit or budget, and code is e.g. RATE_LIMIT_EXCEEDED, BUDGET_EXCEEDED, POLICY_UNAVAILABLE. Both are empty on allow. |
manylayers_policy_evaluation_seconds | histogram | — | Time spent evaluating policy for one request (10 µs – 1 s buckets). |
manylayers_firewall_events_total | counter | team, mode | Prompt injection firewall trips; mode is audit or enforce. |
manylayers_guardrail_executions_total | counter | guardrail_type, hook, status, action | One per guardrail execution. hook: llm_input, llm_output. status: success, error, timeout, canceled, skipped. action: allowed, mutated, blocked, audited, ignored, skipped. |
manylayers_guardrail_execution_seconds | histogram | guardrail_type, hook | Time in one guardrail execution (50 µs – 5 s buckets). Skipped executions are not observed. |
manylayers_cache_events_total | counter | team, result | Cache lookups; result is hit (exact or semantic) or miss. |
Upstream routing
| Metric | Type | Labels | Meaning |
|---|---|---|---|
manylayers_upstream_healthy | gauge | model, upstream | 1 if the endpoint is in rotation, 0 if ejected. upstream is the endpoint’s base URL. |
manylayers_upstream_inflight | gauge | model, upstream | Requests currently in flight to the endpoint. |
manylayers_router_affinity_total | counter | result | Sticky-session and prefix-affinity outcomes: hit, miss, fallback (the pinned endpoint was no longer usable). |
team and model labels grow with the number of teams and models you run. Policy metrics deliberately carry no model, workspace or caller label; use the audit log to investigate a specific refusal.Example queries
Error rate by model
Error rate by model
p95 latency and time to first token
p95 latency and time to first token
Tokens per team per hour
Tokens per team per hour
Rate-limit and budget refusals
Rate-limit and budget refusals
Cache hit rate
Cache hit rate
Ejected upstreams and failovers
Ejected upstreams and failovers
min shows an endpoint any replica has ejected. In the second query, any row whose request and response models differ is traffic served by a fallback.Guardrail blocks and errors
Guardrail blocks and errors
Next steps
Tracing & OpenTelemetry Export
Per-request spans to your OTLP collector.
Request Logging & Audit
Drill from a metric spike to individual requests.
Self-Host the Gateway
Ports, probes and scaling for scrape targets.
Routing
What drives the upstream health gauges.