The pipeline at a glance
Step by step
Authentication
The gateway checks the
Authorization header and identifies who is making the request. It accepts API keys (ml-... prefixed), OIDC JWTs from your identity provider, session cookies from the workspace browser, and Personal Access Tokens.Why it matters: Only authenticated requests proceed. Unauthenticated requests are rejected with 401 before any processing occurs — protecting your upstream provider credits even if a key is leaked.RBAC — team and model access control
The API key is associated with a team. The gateway checks whether that team is allowed to call the requested model. Each team has an explicit model allow-list; if the model isn’t on it, the request is rejected with
403.Why it matters: You control exactly which teams can access which models. The engineering team can use gpt-4o; the customer support team might only have access to a smaller, cheaper model.Rate limiting
The gateway enforces per-team and per-key rate limits: requests per minute (RPM) and tokens per minute (TPM). Limits apply at both the team level and, optionally, at the individual API key level simultaneously.Why it matters: Prevents any single team or application from consuming all your upstream capacity and protects you from runaway loops or buggy clients.
Budget check
The gateway checks the team’s and key’s remaining token and USD budgets. Budgets can be set monthly at the team level and per-lifetime at the key level. If a budget is exhausted, the request is rejected with
402.Why it matters: Gives finance and engineering shared visibility into spend. You can set hard limits that prevent surprise bills.Firewall — prompt injection detection
The gateway scans the incoming prompt for prompt injection patterns. Each team can be in
monitor mode (log and pass through) or block mode (reject flagged requests).Why it matters: Protects against adversarial inputs that try to hijack your LLM’s behavior — especially important for customer-facing applications.PII redaction — input scan
Personally identifiable information in the request (names, email addresses, phone numbers, credit card numbers, and other configurable patterns) is redacted before the prompt is sent upstream or stored in the audit log.Why it matters: Keeps sensitive user data from leaving your perimeter. This runs before the request hits any external provider, so PII is never sent to OpenAI, Anthropic, or any other vendor.
Cache lookup
The gateway checks the cache for a matching response. It supports exact matching (byte-identical prompts) and semantic matching (prompts with the same meaning). A cache hit returns the stored response immediately without calling the upstream provider.Why it matters: Reduces latency and cost for repeated or similar queries — common in RAG pipelines and chat applications with templated prompts.
Guardrail checks — before phase
Before-phase guardrails run on the request. These can be regex pattern matches, webhook calls to external classifiers, or PII policy checks. If a guardrail triggers in block mode, the request is rejected with a configurable error message.Why it matters: Lets you enforce content policies, domain restrictions, or custom business rules before wasting upstream API calls on requests you’d reject anyway.
Router — upstream selection
The gateway selects which upstream provider and model endpoint to use. The routing strategy — failover, canary, load-balanced, or conditional — is configured per team or per request via the
X-ManyLayers-Config header. If the primary upstream fails, failover automatically routes to the next configured backend.Why it matters: Gives you resilience across providers and the ability to safely migrate from one model to another using canary traffic splits.Upstream provider call
The gateway translates the OpenAI-format request into the upstream provider’s format (Anthropic, Bedrock, Vertex, Gemini, Cohere, and more), forwards it, and handles streaming responses. Your application always receives an OpenAI-format response — the translation is transparent.
PII scan — output
The response from the upstream is scanned for PII before it is returned to your application or stored. Any sensitive data in the model’s output can be redacted or flagged.Why it matters: Models sometimes echo back PII from the prompt in unexpected ways. Output scanning closes this gap.
Guardrail checks — after phase
After-phase guardrails run on the completed response. These can check for unwanted content, off-topic answers, or policy violations in the model’s output. Triggered guardrails can block the response and return an error instead.Why it matters: Input guardrails can’t catch everything — the model might still produce disallowed content. Output guardrails are your last line of defense.
Audit log and metering
The complete request and response (post-PII-redaction) are written to the tamper-evident audit log. Token counts and cost are recorded for billing and analytics. The audit entry includes a SHA-256 hash chained to the previous entry.Why it matters: Gives you an immutable record of all AI activity for compliance, security investigation, and cost attribution. The hash chain means any tampering is detectable.
Next steps
- Configure guardrails for your teams
- Set up keys and budgets to control spend
- Explore routing strategies for resilience and canary deploys
- Monitor traffic with metrics and audit logs