Each ManyLayers service exposes Prometheus metrics at GET /metrics. Scrape the gateway replicas for request, token, policy, guardrail, cache and upstream health series; scrape Workspace and Deployer for process health.

How it works

  • /metrics is served by every service, next to /health, /healthz, /ready, /readyz and /version. Default ports are gateway :8180, workspace :8190, deployer :8200. It needs no authentication — restrict it at the network or ingress level.
  • Series are kept per process and cover every team and organization that process serves, so scrape them from a deployment you operate. Per-organization views belong in the console’s Analytics and Request Traces.
  • Query with sum by (...) to aggregate replicas.
  • The default Go runtime (go_*) and process (process_*) collectors are included.
  • Request-path series are emitted by the service that handled the request — the gateway for all /v1 traffic.
prometheus.yml
scrape_configs:
  - job_name: manylayers-gateway
    metrics_path: /metrics
    static_configs:
      - targets: ["gateway-1:8180", "gateway-2:8180"]
  - job_name: manylayers-workspace
    static_configs:
      - targets: ["workspace:8190"]

Metric reference

Requests and tokens

MetricTypeLabelsMeaning
manylayers_requests_totalcounterteam, model, statusCompleted requests. model is the model that served; status is the HTTP status code, e.g. 200, 429.
manylayers_request_duration_secondshistogrammodelEnd-to-end request duration. Buckets 0.05 s doubling to ~102 s.
manylayers_tokens_totalcounterteam, model, directionTokens processed. direction is input or output.

OpenTelemetry GenAI conventions

These follow the OpenTelemetry GenAI metric conventions, so generic LLM dashboards work unchanged. All four carry gen_ai_operation_name (chat, embeddings, generate_content for /v1/responses), gen_ai_system (the serving provider type, or unknown), gen_ai_request_model and gen_ai_response_model. A failover shows up as different request and response models.
MetricTypeExtra labelsMeaning
gen_ai_client_token_usagecountergen_ai_token_typeTokens by type: input, output, cache_read, cache_creation, reasoning. The last three are subsets of the first two — do not sum all types.
gen_ai_client_operation_duration_secondshistogram—End-to-end duration of the operation.
gen_ai_server_time_to_first_token_secondshistogram—Time to first output token. Streamed responses only.
gen_ai_server_time_per_output_token_secondshistogram—Generation time per output token, excluding the wait for the first.

Policy, firewall, guardrails and cache

MetricTypeLabelsMeaning
manylayers_policy_decisions_totalcounteroutcome, policy_type, codePolicy evaluations. outcome is allow or deny. On deny, policy_type is access, model_restriction, token_limit, rate_limit or budget, and code is e.g. RATE_LIMIT_EXCEEDED, BUDGET_EXCEEDED, POLICY_UNAVAILABLE. Both are empty on allow.
manylayers_policy_evaluation_secondshistogram—Time spent evaluating policy for one request (10 µs – 1 s buckets).
manylayers_firewall_events_totalcounterteam, modePrompt injection firewall trips; mode is audit or enforce.
manylayers_guardrail_executions_totalcounterguardrail_type, hook, status, actionOne per guardrail execution. hook: llm_input, llm_output. status: success, error, timeout, canceled, skipped. action: allowed, mutated, blocked, audited, ignored, skipped.
manylayers_guardrail_execution_secondshistogramguardrail_type, hookTime in one guardrail execution (50 µs – 5 s buckets). Skipped executions are not observed.
manylayers_cache_events_totalcounterteam, resultCache lookups; result is hit (exact or semantic) or miss.

Upstream routing

MetricTypeLabelsMeaning
manylayers_upstream_healthygaugemodel, upstream1 if the endpoint is in rotation, 0 if ejected. upstream is the endpoint’s base URL.
manylayers_upstream_inflightgaugemodel, upstreamRequests currently in flight to the endpoint.
manylayers_router_affinity_totalcounterresultSticky-session and prefix-affinity outcomes: hit, miss, fallback (the pinned endpoint was no longer usable).
team and model labels grow with the number of teams and models you run. Policy metrics deliberately carry no model, workspace or caller label; use the audit log to investigate a specific refusal.

Example queries

sum by (model) (rate(manylayers_requests_total{status=~"5.."}[5m]))
  / sum by (model) (rate(manylayers_requests_total[5m]))
histogram_quantile(0.95, sum by (le, model) (rate(manylayers_request_duration_seconds_bucket[5m])))

histogram_quantile(0.95,
  sum by (le, gen_ai_response_model) (rate(gen_ai_server_time_to_first_token_seconds_bucket[5m])))
sum by (team, direction) (increase(manylayers_tokens_total[1h]))
sum by (policy_type, code) (rate(manylayers_policy_decisions_total{outcome="deny"}[5m]))
sum(rate(manylayers_cache_events_total{result="hit"}[15m]))
  / sum(rate(manylayers_cache_events_total[15m]))
min by (model, upstream) (manylayers_upstream_healthy) == 0

sum by (gen_ai_request_model, gen_ai_response_model) (
  rate(gen_ai_client_operation_duration_seconds_count[5m]))
Upstream gauges are per replica; min shows an endpoint any replica has ejected. In the second query, any row whose request and response models differ is traffic served by a fallback.
sum by (guardrail_type, hook, action) (
  rate(manylayers_guardrail_executions_total{action=~"blocked|ignored"}[5m]))
In Grafana, add the Prometheus data source and build one row per section above. Use a $model variable from label_values(manylayers_requests_total, model) and a $team variable from label_values(manylayers_tokens_total, team), and render histogram panels with histogram_quantile over sum by (le, ...) — never average quantiles across replicas.

Next steps

Tracing & OpenTelemetry Export

Per-request spans to your OTLP collector.

Request Logging & Audit

Drill from a metric spike to individual requests.

Self-Host the Gateway

Ports, probes and scaling for scrape targets.

Routing

What drives the upstream health gauges.