gateway.yaml is the primary configuration file for your ManyLayers deployment. Copy the bundled gateway.example.yaml to gateway.yaml and customize it for your environment.
Environment variable interpolation is supported everywhere: ${ENV_VAR}. HashiCorp Vault references also work: ${vault:secret/data/path#key} (requires VAULT_ADDR and VAULT_TOKEN). Unresolvable references cause startup to fail, preventing misconfigured deployments.

server

server:
  listen: ":8180"          # bind address
  max_body_bytes: 10485760 # 10 MiB request body limit
Changes to this section require a full restart.

database

database:
  url: postgres://postgres:test@localhost:5532/manylayers?sslmode=disable
  # url: ${MANYLAYERS_POSTGRES_URL}  # recommended for production
Postgres is the only supported database. Migrations run automatically on startup. Changes to this section require a full restart.

redis

redis:
  addr: ""                 # "" = in-memory limits (single replica)
  # password: ${MANYLAYERS_REDIS_PASSWORD}
  # db: 0
When configured, Redis shares rate limits, budgets, and the exact-match cache across all your gateway replicas. Recommended for any multi-replica deployment. Changes to this section require a full restart.

auth

auth:
  oidc:
    issuer: ""             # "" disables OIDC; set to your IdP URL to enable SSO
    # audience: manylayers
    # jwks_url: ""         # "" = auto-discovery via OIDC
    # team_claim: team     # JWT claim carrying the team name
    # team_mapping:        # map IdP group names to gateway teams
    #   idp-eng-group: engineering
    # role_claim: role
    # admin_role: admin
    # spa_client_id: ""    # OIDC client ID for dashboard SSO login
API-key authentication is always enabled. Configuring OIDC adds SSO JWT bearer token support for your users.

audit

audit:
  retention_days: 90       # auto-prune audit entries after N days
  log_bodies: true         # store request/response bodies (post-PII-redaction only)
The audit log uses a SHA-256 hash chain for tamper evidence. Verify integrity at any time via GET /admin/audit/verify.

pii

pii:
  detector: builtin        # builtin | http | off
  # http_url: "http://localhost:5002/analyze"  # Presidio-compatible sidecar
The builtin detector uses regex-based entity recognition (email, phone, SSN, credit card, etc.). The http mode calls a Presidio-compatible sidecar for production-grade NER with higher accuracy.

firewall

firewall:
  default_mode: audit      # off | audit | enforce
  threshold: 50            # summed rule score that triggers the firewall
Prompt-injection detection. In audit mode, detections are logged but not blocked. In enforce mode, requests exceeding the threshold are rejected. Individual teams can override this gateway default with their own firewall_mode.

cache

cache:
  enabled: false
  max_entries: 1024        # in-process LRU size (ignored when redis.addr is set)
  ttl: 5m
  # semantic:
  #   enabled: false
  #   embedding_model: text-embedding-3-small
  #   threshold: 0.92      # cosine similarity cutoff (0-1)
  #   max_entries: 512
  #   store: memory         # memory | qdrant
The exact-match cache keys on the full request body hash. The semantic cache uses embeddings to match near-duplicate prompts. Use store: qdrant to share the semantic cache across all gateway replicas.

queue

queue:
  backend: db              # db | kafka
  # kafka:
  #   brokers:
  #     - kafka:9092
  #   prefix: ""           # optional topic prefix, e.g. "prod."
The db backend uses Postgres-based polling and works without any additional infrastructure. The kafka backend provides durable topics with dead-letter queues and 5 automatic retries — recommended for high-throughput or multi-replica deployments. Changes to this section require a full restart.

router

router:
  strategy: least_inflight # least_inflight | round_robin | prefix_affinity
  # prefix_bytes: 1024
  # session_ttl: 10m       # sticky session lifetime
  # health_interval: 10s
  # failure_threshold: 3   # consecutive failures before ejecting an upstream
  # cooldown: 30s          # ejection duration before retrying
Controls how the gateway selects among multiple upstream endpoints for a model. You can override this per-request using named routing configs.

rag

rag:
  enabled: false
  vector_store: memory     # memory | qdrant
  # qdrant:
  #   url: http://localhost:6433
  #   api_key: ${QDRANT_API_KEY}
  # hybrid_search: false   # BM25 + vector via RRF
  # rerank_model: ""       # cross-encoder rerank via /v1/rerank
  # embedding_model: ""    # model for per-user memory semantic recall

connectors

connectors:
  encryption_key: ""       # AES-256-GCM passphrase for credential storage
Required if your team uses OAuth-based connectors (Confluence, Google Drive, SharePoint) or S3. Web crawl connectors work without an encryption key. This key is also used to encrypt MCP server auth headers. Changes to this key require a full restart.

suite

suite:
  enabled: false
  # session_ttl: 720h      # 30-day cookie sessions
  # local_users_enabled: true
  # user_approval_required: false
  # http_tool_allowlist:
  #   - api.example.com
Enables the workspace: chat, agents, workflows, evals, knowledge bases, and all /w/* endpoints. Set local_users_enabled: true to allow email and password logins.

deployer

deployer:
  mode: ""                 # "" (disabled) | kubernetes | dstack
  # namespace: manylayers
  # reconcile_interval: 5s
  # cold_start_timeout: 2m
  # dstack:
  #   server_url: http://localhost:3100
  #   token: ${DSTACK_TOKEN}
  #   project: main

models

Each model entry maps a logical name (what your clients use) to an upstream (where requests are sent):
models:
  - logical_name: gpt-4o
    provider: openai       # openai | anthropic | azure | gemini | vertex | bedrock | cohere | mistral | groq | deepseek | together | perplexity | xai | fireworks | openrouter | cerebras | sambanova | ai21 | ollama
    upstream_url: https://api.openai.com
    upstream_model: gpt-4o
    upstream_api_key: ${OPENAI_API_KEY}
    encoding: cl100k_base
    input_price_per_1k: 0.005
    output_price_per_1k: 0.015
    pii_output_mode: passthrough  # passthrough | buffer

aliases

A model may declare additional names that resolve to it. Aliases let callers keep sending a name they already use while you change what serves it, without either side editing code:
models:
  - logical_name: house-fast
    aliases: ["gpt-4o-mini", "fast"]
    provider: groq
    upstream_url: https://api.groq.com/openai
    upstream_model: llama-3.3-70b-versatile
    encoding: cl100k_base
Aliases appear in GET /v1/models, work on GET /v1/models/{model}, and a team policy naming either the alias or the model grants both. A model name always wins over an alias of the same name, so an alias can never reroute traffic away from a real model. See the Providers page for per-provider configuration examples, and OpenAI Compatibility for what each provider can honor.

teams

teams:
  - name: engineering
    monthly_token_budget: 1000000  # 0 = unlimited
    rpm_limit: 60
    tpm_limit: 100000
    cache_enabled: false
    firewall_mode: ""      # "" = use gateway default
    models: ["gpt-4o-mini"]
    api_keys:
      - name: dev
        key: ${MANYLAYERS_DEV_KEY}
        role: member       # admin | member
        # budget_usd_monthly: 50.00
        # rpm_limit: 100
        # tpm_limit: 20000
        # expires_at: "2027-01-01T00:00:00Z"
You can reload model and team configuration without downtime by sending SIGHUP to the gateway process.