ManyLayers ships as a single Go binary containing three integrated modules: the Gateway, the Workspace, and Deployment. You can run all three together or enable only what you need.

The three modules

Gateway is the core. Every LLM request — from your own applications, from Workspace, or from any tool that speaks the OpenAI API — flows through the gateway pipeline. The gateway handles authentication, rate limiting, budget enforcement, guardrails, caching, routing, and audit logging before forwarding the request to an upstream provider. Workspace is the built-in AI productivity layer. It gives your team a chat interface, RAG knowledge bases, web search, voice, agents, and workflow automation — all routed through the gateway, so the same controls apply. Deployment lets you host open-weight models (Llama, Mistral, Qwen, and others) on your own GPUs using dstack or Kubernetes. Once deployed, those models register as providers in the gateway and are callable through the same standard API.

Where configuration lives

Settings are stored in a database. On first boot, ManyLayers seeds initial configuration from a YAML file (gateway.yaml). After that, changes made in the admin UI persist in the database and take precedence. The YAML file acts as a baseline — useful for infrastructure-as-code deployments where you version-control your configuration.
You can always export your current running configuration back to YAML from the admin UI or the API, making it easy to keep your config file in sync.

Organizations

Every tenant is an organization. Organizations sign up independently and are fully isolated — separate teams, users, keys, conversations and knowledge bases. Operators can configure approval workflows, quota plans and billing webhooks. A single-team deployment is simply one organization, and a platform hosting many customers is many. It is the same gateway either way; there is no mode to choose. The mode: onprem key that used to select a single-tenant variant has been removed.

What happens to a request

Every request — regardless of which application sends it — passes through the same ordered pipeline in the gateway: authentication, RBAC, rate limits, budget check, firewall, PII scanning, cache lookup, guardrail checks, routing to the upstream, response scanning, and audit logging. See Request lifecycle for a step-by-step walkthrough of each stage and why it matters.