The gateway can send an OpenTelemetry trace of each request — authentication, policy, guardrails, routing, the model call — to any OTLP/HTTP endpoint you run. Use it to see gateway time next to your own services in Jaeger, Tempo, Honeycomb, Datadog or any OTLP-compatible backend.

How it works

Destinations are configured per Gateway workspace in the console (Destinations, under Observability) or the admin API; there is no global endpoint in gateway.yaml. While a request runs, the gateway records spans in memory. When it finishes, every destination of the request’s workspace is evaluated in this order:
enabled? → region accepted? → API key filter? → filter rules? → sampling? → export
Selected traces are batched and exported asynchronously through a bounded queue per destination, so an unreachable collector never slows or fails an LLM request; when a queue is full, spans are dropped. Destination changes reach every replica within 30 seconds. If the caller sent a W3C traceparent, the gateway continues that trace id, so your application’s spans and the gateway’s appear as one trace. The caller’s sampled flag is not consulted: each destination’s own sampling_rate decides. The gateway does not forward traceparent to providers. Every exported span carries the resource service.name manylayers-gateway, manylayers.destination.id, and deployment.environment.name when telemetry.environment is set.

Spans

SpanCovers
manylayers.requestThe whole request. The trace’s root, or a child of your span when you sent a traceparent.
manylayers.authenticationCredential resolution.
manylayers.policyAccess, model, rate, token and budget checks.
manylayers.guardrail.input / manylayers.guardrail.outputGuardrail hooks; present only when guardrails apply to the request.
manylayers.routingTarget selection; present only when a routing config applies.
{operation} {model}, e.g. chat gpt-4o-miniOne generation per target tried. {operation} is chat, embeddings or generate_content (/v1/responses). The last one carries usage, cost and content.
provider.requestThe outbound HTTP call to the provider, one per upstream attempt, inside the generation span.
A request refused before a model is called (blocked by a guardrail, denied by policy, served from cache) has no generation span; its usage and cost attributes go on the root span instead. Only 5xx responses mark the root span as an error; a 4xx refusal is the gateway working as configured, and a policy denial leaves the policy span’s status unset. A failed generation or provider attempt, such as one that fell over to the next target, has status Error and an error.type: the provider’s HTTP status, or timeout, canceled, stream_interrupted or _OTHER.

Attributes

LLM data follows the OpenTelemetry GenAI semantic conventions. The generation span always carries gen_ai.operation.name, gen_ai.provider.name (for example openai, anthropic, aws.bedrock), gen_ai.request.model, gen_ai.request.stream and http.response.status_code, plus gen_ai.response.id and gen_ai.response.finish_reasons when the provider returned them and, for streamed responses, gen_ai.response.time_to_first_chunk (seconds) and manylayers.time_to_last_token_ms. The root span always carries manylayers.request.id, the organization, workspace and API key ids, manylayers.credential.type / .id, manylayers.request.source (playground for the console) and the HTTP method, route and status code. Everything else is grouped by a destination’s metadata_config, all off by default. request_context and cost attributes go on the generation span; identity attributes go on the root span:
GroupAttributes
request_contextgen_ai.response.model, gen_ai.conversation.id (from X-Session-Id), manylayers.request.id, manylayers.routing.strategy, HTTP method, route and status, manylayers.guardrail.result (passed, blocked, modified, error) and manylayers.guardrail.name
costgen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.usage.cache_read.input_tokens, gen_ai.usage.cache_write.input_tokens, gen_ai.usage.reasoning.output_tokens, manylayers.cost.usd, manylayers.usage.estimated
raw_provider_usage (needs cost)manylayers.provider.raw_usage, the provider’s usage JSON verbatim
identityuser.id, manylayers.organization.id, .workspace.id, .team.id, .api_key.id, .service_account.id, .virtual_account.id, .project.id
Input tokens include cached ones; cache and reasoning counts are subsets, so do not add them. Cache and reasoning counts appear only when the provider reported usage; when the gateway had to estimate tokens, manylayers.usage.estimated is true. A value the gateway does not have is omitted, never exported as 0. The policy span adds manylayers.policy.result (allowed, denied), manylayers.policy.decision and the rate-limit or budget result. Guardrail spans add manylayers.guardrail.result and, for the guardrail that blocked or modified, manylayers.guardrail.name, .type and .action. The routing span adds manylayers.routing.strategy and, for Auto Routing, manylayers.routing.reason. Content is gen_ai.input.messages, gen_ai.output.messages and, for a Messages API system prompt, gen_ai.system_instructions, as JSON in the conventions’ role + parts shape. Content is taken after the gateway’s PII redaction. It is not governed by logging configs or X-ManyLayers-Store-Logs, which control the stored audit copy; use privacy_mode to keep content out of a destination. Content a guardrail flagged and left in place is withheld.

Configuration

Destinations belong to a Gateway workspace. Call the admin API on the console host with a personal access token (ml_pat_…) whose owner holds gateway.observability.manage:
POST https://app.manylayers.io/admin/gateway/observability/destinations?workspace_id={ws}
{
  "name": "tempo-prod",
  "type": "otel_collector",                       // the only type
  "endpoint": "https://otel.example.com:4318/v1/traces",
  "headers": "{\"Authorization\": \"Bearer abc\"}", // JSON object as a string; encrypted at rest
  "enabled": true,                                 // default true
  "sampling_rate": 0.25,                           // 0–1, default 1
  "privacy_mode": false,                           // true drops prompt and completion content
  "regions": ["eu"],                               // global | eu | us; empty = all
  "metadata_config": {
    "cost": true,                 // token counts, cache/reasoning tokens, USD cost
    "raw_provider_usage": false,  // provider's usage JSON (needs cost)
    "identity": true,             // org/workspace/team/key/user ids
    "request_context": true       // models, routing, HTTP, streaming, guardrail result
  },
  "filter_config": {              // groups are OR'd, rules in a group AND'd
    "groups": [
      {"rules": [{"field": "status", "operator": "equals", "value": "error"}]},
      {"rules": [{"field": "total_tokens", "operator": "greater_than", "value": "5000"},
                 {"field": "provider", "operator": "equals", "value": "anthropic"}]}
    ]
  },
  "api_key_filter": {"included": [], "excluded": ["<api-key-id>"]}
}
PUT takes the same body and replaces the destination: an omitted sampling_rate returns to 1, and an omitted privacy_mode or metadata_config returns to off. Only headers and enabled are kept when omitted.

Key fields

endpoint
string
required
An http or https OTLP/HTTP traces URL, at most 2048 characters. A URL with no path uses /v1/traces. gRPC endpoints and credentials in the URL are rejected — put credentials in headers.
headers
string
A JSON object of string values, up to 50 headers. Stored encrypted with the deployment’s credential key (MANYLAYERS_ENCRYPTION_KEY); without a key, saving headers fails. Reads return only headers_hint (names with masked values) and headers_set. On update, omit the field to keep the stored headers.
sampling_rate
number
default:"1"
Fraction of the traces your filters selected that are exported. The draw is derived from the trace id, so every span of a trace gets the same decision, and a destination at 0.5 receives a superset of one at 0.1.
privacy_mode
boolean
default:"false"
When true, the content attributes gen_ai.input.messages, gen_ai.output.messages and gen_ai.system_instructions are never attached. Identity and cost attributes are unaffected.
metadata_config
object
Optional attribute groups, all off by default. New groups added in later releases also start off, so an upgrade never widens what a destination receives.
regions
string[]
Matched against the deployment’s telemetry.region. Empty accepts every region.
api_key_filter
object
included and excluded lists of API key ids. Exclusion wins; a non-empty included list is an allow-list.
filter_config
object
Up to 20 groups of up to 20 rules; a rule value is at most 512 characters. Fields and operators come from GET /admin/gateway/observability/filter-fields.

Filter fields

TypeFieldsOperators
stringmodel, provider, api_key, user, team, organization, workspace, routing_strategy, selected_model, selected_provider, guardrail_name, request_path, environmentequals, not_equals, contains, starts_with, ends_with
numberinput_tokens, output_tokens, total_tokens, cost, http_statusequals, not_equals, greater_than, greater_than_or_equal, less_than, less_than_or_equal
enumstatus (success, error), guardrail_result (passed, blocked, modified, error)equals, not_equals
booleanstreamingis, is_not
api_key, user, team, organization and workspace match internal ids, not names. model is the model the caller requested and selected_model the one that served. status is error for any HTTP status of 400 or above. A rule on a field the request has no value for does not match, including not_equals.

Deployment settings

telemetry:
  region: eu              # global (default) | eu | us; unknown values read as global
  environment: production # resource attribute deployment.environment.name; empty omits it

Admin API

All routes are on the console host (https://app.manylayers.io) and take ?workspace_id=. Reading needs gateway.observability.read (every Gateway role); everything that writes or makes an outbound call needs gateway.observability.manage (Gateway admins).
MethodPathPurpose
GET/admin/gateway/observability/destinationsList destinations.
POST/admin/gateway/observability/destinationsCreate.
GET / PUT / DELETE/admin/gateway/observability/destinations/{id}Read, update, delete.
GET/admin/gateway/observability/destinations/{id}/headersDecrypted headers (manage only).
POST/admin/gateway/observability/destinations/test-connectionSend one synthetic span.
POST/admin/gateway/observability/destinations/send-traceSend a sample manylayers.request → chat sample-model trace with invented values.
GET/admin/gateway/observability/filter-fieldsFilter fields, operators and regions.
Both probes accept {"endpoint", "headers"} for an unsaved form, or {"id"} to test a saved destination with its stored headers. They answer {"ok": true} or {"ok": false, "error": "…"} within 15 seconds; send-trace also returns the trace_id to search for.

Correlating a request

X-ManyLayers-Trace-Id
string
The audit record id of this request. Open it in Request Traces or query /admin/audit?id=.
X-ManyLayers-Otel-Trace-Id
string
The W3C trace id of the exported trace. Present only when a destination is configured somewhere in the deployment, whether or not this request’s trace is exported; search your tracing backend for it.
The audit record also stores the OTel trace and root span ids, so you can move between the two in either direction.
One destination at sampling_rate: 1 with two groups: status equals error, and cost greater_than 0.50.
privacy_mode: true, metadata_config.identity: true, filter guardrail_result equals blocked.
Set MANYLAYERS_OTLP_DEBUG=1 on the gateway to print every OTLP export (URL, headers, decoded spans as JSON) to stderr. It prints prompt content for non-privacy destinations — never enable it in production.

Next steps

Metrics

Aggregate Prometheus series for dashboards and alerts.

Request Logging & Audit

The stored record behind X-ManyLayers-Trace-Id.

Guardrails

The checks behind the guardrail spans.

Routing

What the routing span decides.