The gateway can answer a request from a cache of earlier completions instead of calling the provider again. Use it for deterministic workloads — classification, extraction, FAQ bots — where the same prompt recurs and a fresh generation adds cost and latency but no value.
The cache is switched on for a whole deployment by its operator, then each team opts in. If X-ManyLayers-Cache on your responses always says disabled, one of those two switches is off — see Eligibility.

How it works

The cache has two layers that are checked in order:
  1. Exact match. A SHA-256 key built from the redacted request. A hit returns the stored completion immediately.
  2. Semantic match (optional). On an exact miss, the prompt is embedded and compared with earlier prompts for the same team and model. A close enough neighbour is returned.
On a miss the request goes to the provider. When it succeeds with HTTP 200, the completion is stored in the exact layer and, if an embedding was computed during lookup, in the semantic layer. Cache lookup runs after PII redaction and input guardrail mutations, so keys and stored entries never contain raw PII. The stored completion is the outbound-redacted text, so a replay never re-exposes PII even when the original response streamed through unredacted. A cache hit still passes through your input and output guardrails, and still counts against request-per-minute limits and spend policies, which are checked first; token-per-minute limits are not charged for a hit. Only the text of the completion (plus its token counts) is stored. A non-streamed replay is a freshly built response with an id starting cache-, a single assistant message and finish_reason: "stop"; a streamed request is replayed as a stream.

Eligibility

A request uses the cache only when all of these hold:
  • cache.enabled: true for the deployment.
  • The calling team has cache_enabled: true.
  • The endpoint is /v1/chat/completions, /v1/completions, /v1/responses or /v1/messages. Embeddings are never cached.
  • The request sets temperature to exactly 0, or the model is configured with cache: true.
A request that fails any condition is answered normally with X-ManyLayers-Cache: disabled.
The model-level cache: true flag is a gateway.yaml model setting. Models registered through Gateway → Providers do not carry it, so for them only requests with "temperature": 0 are cached.
Set a team’s opt-in with cache_enabled on POST /admin/teams (organization Owner; see Keys & Budgets).

Exact key composition

The exact key hashes:
  • the resolved logical model name;
  • the text of every message (after redaction), then prompt, then input;
  • these parameters, when present: temperature, top_p, max_tokens, n, stop, tools, tool_choice, response_format.
stream is not part of the key, so a streamed and a non-streamed request share one entry; a streamed request is replayed as a stream.
The exact key does not include the team, the workspace or the message roles. Any caller with caching enabled that sends the same redacted text with the same parameters to the same model name receives the same cached completion. Keep cache_enabled off for teams whose prompts must stay isolated, or use a distinct model alias. The semantic layer, by contrast, is scoped per team and model.

Configuration

Deployment-wide settings, from gateway.yaml or the administration console under Settings → Cache:
cache:
  enabled: true
  max_entries: 1024          # in-memory LRU capacity (ignored with Redis)
  ttl: 5m                    # lifetime of every entry, both layers
  semantic:
    enabled: false
    embedding_model: text-embedding-3-small   # required when enabled
    threshold: 0.92          # cosine similarity needed for a hit (0–1)
    max_entries: 512         # per (team, model), memory store only
    store: memory            # memory | qdrant

models:
  - logical_name: gpt-4o-mini
    cache: true              # cache even when temperature != 0

teams:
  - name: support-bot
    cache_enabled: true

Key fields

cache.enabled
boolean
default:"false"
Turns the cache on for the deployment. Teams still opt in individually.
cache.ttl
duration
default:"5m"
How long an entry lives, in both layers.
cache.max_entries
integer
default:"1024"
Capacity of the in-process LRU. Not used when Redis is configured.
cache.semantic.embedding_model
string
A model the gateway calls on itself, as the calling identity, to embed prompts. It must be in the requesting team’s allowed models; if the embedding call fails for any reason, the request is treated as a plain miss.
cache.semantic.threshold
number
default:"0.92"
Minimum cosine similarity for a semantic hit. Raise it to reduce false matches.
cache.semantic.store
string
default:"memory"
memory keeps a per-replica LRU. qdrant shares entries across replicas in the ml_semantic_cache collection and requires rag.qdrant.url.

Backends

LayerWithout RedisWith Redis (redis.url, redis.addr or REDIS_URL)
ExactPer-replica LRU, max_entriesShared across all replicas, keys prefixed mlcache:, TTL eviction
Semanticstore: memory per replica, or store: qdrant sharedSame — the semantic store is chosen by store, not by Redis
Redis errors are treated as cache misses, never as request failures. Backend changes take effect after a restart.

Per-request control

X-ManyLayers-Cache-Control
string
no-cache skips the lookup; no-store skips writing the answer. Combine them as no-cache, no-store. Any other value is a 400. The header can only make the gateway cache less — it never makes an ineligible request cacheable — and it is removed before the request is forwarded.
X-ManyLayers-Cache
string
ValueMeaning
hitServed from the exact layer.
semanticServed from the semantic layer.
missEligible, looked up, not found; the provider answered.
bypassEligible, but the request sent no-cache.
disabledNot eligible (see Eligibility).
Kept even when a routing config sets strict_openai_compliance, which otherwise strips the other X-ManyLayers-* response headers.
curl https://app.manylayers.io/v1/chat/completions \
  -H "Authorization: Bearer ml-..." \
  -H "X-ManyLayers-Cache-Control: no-cache" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4o-mini", "temperature": 0,
       "messages": [{"role": "user", "content": "Classify: refund request"}]}' -i

Monitoring

GET /admin/cache/stats aggregates lookups since ?since= (RFC 3339, default the last 24 hours). It is a deployment-operator endpoint: it needs an admin-role credential and is not available to workspace roles, so on the hosted service ask ManyLayers or use the Prometheus metric below.
curl "https://app.manylayers.io/admin/cache/stats?since=2026-10-01T00:00:00Z" \
  -H "Authorization: Bearer ml-..."
{
  "since": "2026-10-01T00:00:00Z",
  "hits": 1830, "misses": 4120, "hit_rate": 0.3076,
  "latency_saved_ms": 2196000,
  "stats": [{"team_id": "…", "model": "gpt-4o-mini", "hits": 1830, "misses": 4120, "latency_saved_ms": 2196000}]
}
latency_saved_ms sums the original generation time of each entry that was served.
Prometheus metricLabelsValues of result
manylayers_cache_events_totalteam, resulthit (exact or semantic), miss
Bypassed and ineligible requests are not counted.

Next steps

PII Detection & Redaction

Redaction runs before the cache key is built.

Metrics

Chart hit rate alongside latency and spend.

Self-Host the Gateway

Add Redis to share the exact cache across replicas.

Routing

What happens on a cache miss.