The cache is switched on for a whole deployment by its operator, then each team opts in. If
X-ManyLayers-Cache on your responses always says disabled, one of those two switches is off — see Eligibility.How it works
The cache has two layers that are checked in order:- Exact match. A SHA-256 key built from the redacted request. A hit returns the stored completion immediately.
- Semantic match (optional). On an exact miss, the prompt is embedded and compared with earlier prompts for the same team and model. A close enough neighbour is returned.
id starting cache-, a single assistant message and finish_reason: "stop"; a streamed request is replayed as a stream.
Eligibility
A request uses the cache only when all of these hold:cache.enabled: truefor the deployment.- The calling team has
cache_enabled: true. - The endpoint is
/v1/chat/completions,/v1/completions,/v1/responsesor/v1/messages. Embeddings are never cached. - The request sets
temperatureto exactly0, or the model is configured withcache: true.
X-ManyLayers-Cache: disabled.
The model-level
cache: true flag is a gateway.yaml model setting. Models registered through Gateway → Providers do not carry it, so for them only requests with "temperature": 0 are cached.cache_enabled on POST /admin/teams (organization Owner; see Keys & Budgets).
Exact key composition
The exact key hashes:- the resolved logical model name;
- the text of every message (after redaction), then
prompt, theninput; - these parameters, when present:
temperature,top_p,max_tokens,n,stop,tools,tool_choice,response_format.
stream is not part of the key, so a streamed and a non-streamed request share one entry; a streamed request is replayed as a stream.
Configuration
Deployment-wide settings, fromgateway.yaml or the administration console under Settings → Cache:
Key fields
Turns the cache on for the deployment. Teams still opt in individually.
How long an entry lives, in both layers.
Capacity of the in-process LRU. Not used when Redis is configured.
A model the gateway calls on itself, as the calling identity, to embed prompts. It must be in the requesting team’s allowed models; if the embedding call fails for any reason, the request is treated as a plain miss.
Minimum cosine similarity for a semantic hit. Raise it to reduce false matches.
memory keeps a per-replica LRU. qdrant shares entries across replicas in the ml_semantic_cache collection and requires rag.qdrant.url.Backends
| Layer | Without Redis | With Redis (redis.url, redis.addr or REDIS_URL) |
|---|---|---|
| Exact | Per-replica LRU, max_entries | Shared across all replicas, keys prefixed mlcache:, TTL eviction |
| Semantic | store: memory per replica, or store: qdrant shared | Same — the semantic store is chosen by store, not by Redis |
Per-request control
no-cache skips the lookup; no-store skips writing the answer. Combine them as no-cache, no-store. Any other value is a 400. The header can only make the gateway cache less — it never makes an ineligible request cacheable — and it is removed before the request is forwarded.| Value | Meaning |
|---|---|
hit | Served from the exact layer. |
semantic | Served from the semantic layer. |
miss | Eligible, looked up, not found; the provider answered. |
bypass | Eligible, but the request sent no-cache. |
disabled | Not eligible (see Eligibility). |
strict_openai_compliance, which otherwise strips the other X-ManyLayers-* response headers.Monitoring
GET /admin/cache/stats aggregates lookups since ?since= (RFC 3339, default the last 24 hours). It is a deployment-operator endpoint: it needs an admin-role credential and is not available to workspace roles, so on the hosted service ask ManyLayers or use the Prometheus metric below.
latency_saved_ms sums the original generation time of each entry that was served.
| Prometheus metric | Labels | Values of result |
|---|---|---|
manylayers_cache_events_total | team, result | hit (exact or semantic), miss |
Next steps
PII Detection & Redaction
Redaction runs before the cache key is built.
Metrics
Chart hit rate alongside latency and spend.
Self-Host the Gateway
Add Redis to share the exact cache across replicas.
Routing
What happens on a cache miss.