llm_input) and the model’s response before your client does (llm_output). A guardrail can validate content — allow it or block it — or mutate it, replacing what it found before the request or response continues.
Every guardrail runs inside the gateway process. There are no external moderation services, no remote guardrail servers and no third-party calls on the request path.
Guardrails decide what safety checks apply. Policies decide whether a request is allowed at all. Policy runs first; a request refused by a policy never reaches a guardrail.
Concepts
| Concept | What it is |
|---|---|
| Guardrail type | One of the seven built-in implementations, such as pii or secrets. |
| Guardrail | One configuration of a type: its name, operation, priority, enforcement strategy and settings. pii-strict and pii-audit can be two guardrails over the same type. |
| Guardrail group | A named container of guardrails. A guardrail is referred to as <group>/<guardrail>. |
| Rule | Which requests a set of guardrails applies to, and on which hook. |
"scope": "organization" — to the organization, where they apply to every Gateway workspace. A resource in one workspace is invisible from every other.
In the console, open Guardrails in the Gateway sidebar: the Registry tab holds groups and their guardrails, the Rules tab holds rules.
The built-in guardrails
| Type | Operations | Default priority | What it does |
|---|---|---|---|
pii | mutate | 5 | Detects and masks PII/PHI: names, email, phone, address, SSN, passport, credit card, date of birth, medical record numbers, health information, IP addresses, bank accounts and Aadhaar numbers. |
secrets | validate, mutate (default) | 10 | Detects credentials: cloud and provider API keys, GitHub and Slack tokens, JWTs, bearer, OAuth and refresh tokens, private keys (RSA, SSH, PGP), database connection strings, generic key assignments, passwords and high-entropy secrets. |
prompt_injection | validate | 10 | Detects instruction overrides, system prompt extraction, jailbreaks, instruction-hierarchy attacks, indirect injection and obfuscated payloads. |
regex | validate (default), mutate | 20 | Matches preset patterns (SSN, email, phone, national IDs and passports, card numbers, cloud keys, IPs, URLs) and your own custom patterns. |
content_moderation | validate | 30 | Grades hate, self-harm, sexual and violent content against per-category severity thresholds. |
code_safety | validate | 40 | Flags unsafe Python, shell, JavaScript and SQL: eval, exec, os.system, child_process, rm -rf, mkfs, dd, DROP, GRANT, and more. |
sql_sanitizer | validate (default), mutate | 1 | Blocks destructive statements, DELETE/UPDATE without WHERE and injection patterns; in mutate mode strips SQL comments and blocks anything it cannot make safe. |
GET /admin/gateway/guardrail-types returns the same catalog with each type’s operations, default operation, default priority and default settings.
Operations
Validate inspects content and never changes it. If the guardrail finds a violation, the verdict is block. Mutate may rewrite content. The rewritten content is what every later guardrail, and the model, receives:Enforcement strategies
A violation (the guardrail ran and rejected the content) and an error (the guardrail could not complete, or timed out) are different facts, and strategies treat them differently:| Strategy | Violation | Error or timeout |
|---|---|---|
enforce | block | block |
enforce_but_ignore_on_error | block | allow |
audit | allow, and record | allow, and record |
audit mode a mutating guardrail records what it would have changed but leaves the content untouched.
How a request flows
Streaming
With"stream": true, input guardrails run as usual. When any output guardrail applies, the gateway holds the provider’s stream, runs the output guardrails on the assembled completion, and then re-emits the guarded completion to your client as a stream — or returns the refusal. Nothing the model generated reaches your client before the output guardrails have passed it. The cost is time to first token. Such responses carry X-ManyLayers-Guardrails-Output: buffered-streaming.
Message scope
Each guardrail has amessage_scope: all (default), last, or a number N for the latest N messages. System and developer messages are never inspected — they are your instructions, not caller content — and are never rewritten. Text parts of a multimodal message are inspected and rewritten individually; image parts are left as they were.
Rules
A rule targets requests and binds guardrails to hooks:- Within a rule, every dimension it names must match; a dimension matches when the request’s value is any listed value. An empty dimension matches everything.
usersmatches the person a request belongs to — the session’s user, or the person an API key was issued to.teamsmatches the request’s team, or a team its person belongs to.keysmatches the API key the request arrived on.modelsmatches the requested model or the model an alias points at.metadatamatches values from theX-ManyLayers-Metadatarequest header. A key listed with no values matches any value for that key.- Every matching rule applies. Resolution is a union, never first-match-wins. A guardrail selected by several rules runs once.
PATCH with {"enabled": false}).
Creating a group in the console also lets you say where it applies (applies: named, workspace or conditions). That is saved as a rule the group owns, listed under Rules.
Selecting guardrails per request
A request can add guardrails with theX-ManyLayers-Guardrails header:
<group>/<guardrail> or guardrail ids; the header holds at most 32 names and 8 KB. The header only adds guardrails — it cannot remove one a rule applies. Naming a guardrail requires gateway.guardrails.read, and only guardrails visible from the caller’s workspace resolve; anything else is unknown_guardrail. A group restricted to named users, teams or keys can be selected only by them (403 guardrail_access_denied); this limits who can choose a group, never whether a rule’s guardrails apply. The header is never forwarded to a provider.
Errors
| Status | Code | Meaning |
|---|---|---|
| 403 | guardrail_block | A guardrail blocked the request (request blocked by guardrail "…") or the response (response blocked by guardrail "…"). The message names the guardrail and the categories that tripped — never the content. |
| 503 | guardrail_error | An enforce guardrail could not complete. |
| 503 | guardrail_timeout | An enforce guardrail did not finish within its timeout. |
| 503 | guardrails_unavailable | The guardrail configuration could not be loaded. |
| 400 | invalid_guardrails_header / unknown_guardrail | The X-ManyLayers-Guardrails header is malformed or names a guardrail the caller cannot use. |
| 403 | guardrail_access_denied | The header names a guardrail in a group the caller has not been granted. |
Configuring a guardrail
The admin API is served by the console athttps://app.manylayers.io. Authenticate with a personal access token (ml_pat_…) and name the Gateway workspace with ?workspace_id=.
A group and its guardrails are usually written together, in one call. Either the whole group is stored or none of it is:
PUT /admin/gateway/guardrail-groups/{id} with a guardrails list replaces the group’s guardrails in the same way: an entry with an id updates that guardrail, an entry without one creates a guardrail, and a guardrail you leave out is deleted — which also removes it from every rule. Omit guardrails to rename a group without touching what it holds. A group holds at most 100 guardrails. A failure is reported against the entry it belongs to, for example param: "guardrails[1].operation".
Guardrails can also be managed one at a time:
| Field | Default | Notes |
|---|---|---|
name | required | Unique within the group; at most 128 bytes; no /. |
type | required | Fixed once created. |
operation | the type’s default | Must be one the type supports. |
priority | the type’s default | 0–10000; lower runs first. |
enforcement_strategy | enforce | |
message_scope | all | all, last, or 1–1000. |
timeout_ms | 0 (1000 ms) | Up to 10000. |
config | the type’s defaults | Type-specific settings, below. Unknown settings are refused. |
enabled | true |
400 invalid_guardrail_config, naming the offending field, so no request is ever the first to find out.
POST …/guardrails/{id}/test with {"text": "...", "hook": "llm_input"} (text up to 64 KB; hook defaults to llm_input) runs a stored guardrail over sample text and returns blocked, mutated, status, verdict, action, findings and — for a mutating guardrail — the rewritten text. Nothing is sent to a model and nothing is recorded.
Settings by type
secrets
secrets
categories limits detection to a subset (empty means all): aws_access_key, aws_secret_key, openai_api_key, anthropic_api_key, google_api_key, github_token, slack_token, jwt, bearer_token, oauth_token, refresh_token, private_key, rsa_private_key, ssh_private_key, pgp_private_key, database_connection_string, generic_api_key, password, high_entropy_secret. For key=value assignments only the value is redacted.pii
pii
redaction is mask (one mask_char per character, keeping the shape) or placeholder ({category} expands to the category). Names are detected from context cues (“my name is”, honorifics, “call …”). Aadhaar numbers are redacted when labelled (“Aadhaar”, “UID”) in any 4-4-4 form, or unlabelled when they pass the UIDAI format and Verhoeff checksum.prompt_injection
prompt_injection
sensitivity is low (block at score 80), medium (50) or high (30); a threshold of 1–100 overrides it, and 0 uses the sensitivity. Categories: direct_injection, instruction_override, system_prompt_extraction, jailbreak, instruction_hierarchy, indirect_injection, obfuscation. Signals in tool results weigh more, since that is where indirect injection arrives.content_moderation
content_moderation
low, medium, high, critical). thresholds overrides severity_threshold per category.code_safety
code_safety
min_severity blocks. disabled_rules names rules to skip, such as python.eval or shell.chmod_777.sql_sanitizer
sql_sanitizer
validate at priority 1. In mutate mode comments are stripped; if anything blocking remains, the guardrail blocks rather than rewriting the SQL into something it cannot vouch for.regex
regex
ssn, email, us_phone, india_aadhaar, india_pan, passports for India, US, UK, Germany, France, Netherlands, Canada, Australia, China and Japan (india_passport, us_passport, …), visa, mastercard, amex, discover, aws_access_key, aws_secret_key, github_token, slack_token, generic_api_key, ipv4, ipv6, url. Families expand to several: pii, passports, government_id, financial, credentials, network. Patterns use RE2 syntax — matching is linear-time, so there is no catastrophic backtracking — and patterns that are too complex, or that match the empty string, are refused.Observability
Every guardrail execution is recorded with the request id, organization, workspace, user, team, model, guardrail, type, hook, operation, priority, enforcement strategy, status, verdict, action, execution time, whether a mutation was applied, and its normalized findings. Execution records are kept foraudit.retention_days (default 90).
- In the console, Request Traces → Guardrail executions lists them; a single request’s trace shows its guardrail verdicts.
GET /admin/gateway/guardrail-executions?request_id=…returns one request’s guardrail trace in execution order. It also acceptsguardrail_id,action,q,since,until,limitandoffset, and requiresgateway.audit.read.- A
guardrail executionstructured log line is written per execution. manylayers_guardrail_executions_total(labelsguardrail_type,hook,status,action) andmanylayers_guardrail_execution_seconds(guardrail_type,hook) are exported; see Metrics.- Executions that blocked, mutated or found something appear in the request’s audit row, and each hook is a span in exported traces.
Guardrail management API
| Method | Path | Permission |
|---|---|---|
GET | /admin/gateway/guardrail-types | gateway.guardrails.read |
GET, POST | /admin/gateway/guardrail-groups | read / manage |
GET, PUT, DELETE | /admin/gateway/guardrail-groups/{id} | read / manage |
GET, POST | /admin/gateway/guardrail-groups/{id}/guardrails | read / manage |
GET, PUT, DELETE | /admin/gateway/guardrail-groups/{id}/guardrails/{guardrailID} | read / manage |
POST | /admin/gateway/guardrail-groups/{id}/guardrails/{guardrailID}/test | read |
GET | /admin/gateway/guardrail-groups/{id}/access | read |
POST | /admin/gateway/guardrail-groups/{id}/access | manage, or group manager |
DELETE | /admin/gateway/guardrail-groups/{id}/access/{bindingID} | manage, or group manager |
GET | /admin/gateway/guardrail-groups/{id}/activity | read |
GET, POST | /admin/gateway/guardrail-rules | read / manage |
GET, PUT, PATCH, DELETE | /admin/gateway/guardrail-rules/{id} | read / manage |
GET | /admin/gateway/guardrail-executions | gateway.audit.read |
GET | /admin/gateway/guardrail-executions/filters | gateway.audit.read |
?workspace_id=. Organization-wide resources require gateway.guardrails.manage across the Gateway product to write.
Gateway-wide checks that run before guardrails
Two older checks run on every request independently of guardrails, and guardrails see the request after both:| Check | What it does | Page |
|---|---|---|
| Prompt injection firewall | Scores request text; per-team off / audit / enforce. | Prompt Injection Firewall |
| PII detection and redaction | Replaces PII before the cache, provider and audit log; optionally reversible. | PII Detection & Redaction |
pii or prompt_injection guardrail adds per-workspace, per-rule control on top of these.