Guardrails are the safety checks that run on AI traffic inside the gateway: they inspect a request before the model sees it (llm_input) and the model’s response before your client does (llm_output). A guardrail can validate content — allow it or block it — or mutate it, replacing what it found before the request or response continues. Every guardrail runs inside the gateway process. There are no external moderation services, no remote guardrail servers and no third-party calls on the request path.
Guardrails decide what safety checks apply. Policies decide whether a request is allowed at all. Policy runs first; a request refused by a policy never reaches a guardrail.

Concepts

ConceptWhat it is
Guardrail typeOne of the seven built-in implementations, such as pii or secrets.
GuardrailOne configuration of a type: its name, operation, priority, enforcement strategy and settings. pii-strict and pii-audit can be two guardrails over the same type.
Guardrail groupA named container of guardrails. A guardrail is referred to as <group>/<guardrail>.
RuleWhich requests a set of guardrails applies to, and on which hook.
Groups, guardrails and rules belong to a Gateway workspace, or — with "scope": "organization" — to the organization, where they apply to every Gateway workspace. A resource in one workspace is invisible from every other. In the console, open Guardrails in the Gateway sidebar: the Registry tab holds groups and their guardrails, the Rules tab holds rules.

The built-in guardrails

TypeOperationsDefault priorityWhat it does
piimutate5Detects and masks PII/PHI: names, email, phone, address, SSN, passport, credit card, date of birth, medical record numbers, health information, IP addresses, bank accounts and Aadhaar numbers.
secretsvalidate, mutate (default)10Detects credentials: cloud and provider API keys, GitHub and Slack tokens, JWTs, bearer, OAuth and refresh tokens, private keys (RSA, SSH, PGP), database connection strings, generic key assignments, passwords and high-entropy secrets.
prompt_injectionvalidate10Detects instruction overrides, system prompt extraction, jailbreaks, instruction-hierarchy attacks, indirect injection and obfuscated payloads.
regexvalidate (default), mutate20Matches preset patterns (SSN, email, phone, national IDs and passports, card numbers, cloud keys, IPs, URLs) and your own custom patterns.
content_moderationvalidate30Grades hate, self-harm, sexual and violent content against per-category severity thresholds.
code_safetyvalidate40Flags unsafe Python, shell, JavaScript and SQL: eval, exec, os.system, child_process, rm -rf, mkfs, dd, DROP, GRANT, and more.
sql_sanitizervalidate (default), mutate1Blocks destructive statements, DELETE/UPDATE without WHERE and injection patterns; in mutate mode strips SQL comments and blocks anything it cannot make safe.
GET /admin/gateway/guardrail-types returns the same catalog with each type’s operations, default operation, default priority and default settings.

Operations

Validate inspects content and never changes it. If the guardrail finds a violation, the verdict is block. Mutate may rewrite content. The rewritten content is what every later guardrail, and the model, receives:
"My SSN is 123-45-6789"   →   pii (mutate)   →   "My SSN is ***********"
Mutating guardrails run one at a time, in ascending priority — lower numbers first — because each one’s output is the next one’s input. Validating guardrails run after every mutation, concurrently, over the final mutated content.

Enforcement strategies

A violation (the guardrail ran and rejected the content) and an error (the guardrail could not complete, or timed out) are different facts, and strategies treat them differently:
StrategyViolationError or timeout
enforceblockblock
enforce_but_ignore_on_errorblockallow
auditallow, and recordallow, and record
In audit mode a mutating guardrail records what it would have changed but leaves the content untouched.

How a request flows

request
  → policy
  → prompt injection firewall         (gateway-wide, see below)
  → PII redaction                     (gateway-wide, see below)
  → rule resolution                   (union of every matching rule)
  → input mutations                   (sequential, by priority)
  → input validations  ┐
  → model call         ┘ concurrently — a blocking validation cancels the model call
  → output mutations                  (sequential, by priority)
  → output validations
  → client
Input validation runs alongside the model call, so an allowed request pays no extra latency for it. Nothing the model returns reaches your client until input validation has allowed the request; if it blocks, the upstream call is cancelled and the response is discarded. A cache hit is held to the same guardrails: input validation must allow the request, and a non-streamed replay passes the output hook.

Streaming

With "stream": true, input guardrails run as usual. When any output guardrail applies, the gateway holds the provider’s stream, runs the output guardrails on the assembled completion, and then re-emits the guarded completion to your client as a stream — or returns the refusal. Nothing the model generated reaches your client before the output guardrails have passed it. The cost is time to first token. Such responses carry X-ManyLayers-Guardrails-Output: buffered-streaming.

Message scope

Each guardrail has a message_scope: all (default), last, or a number N for the latest N messages. System and developer messages are never inspected — they are your instructions, not caller content — and are never rewritten. Text parts of a multimodal message are inspected and rewritten individually; image parts are left as they were.

Rules

A rule targets requests and binds guardrails to hooks:
{
  "name": "engineering on gpt-5",
  "conditions": {
    "teams": ["<team id>"],
    "models": ["gpt-5"],
    "users": [],
    "keys": [],
    "metadata": { "env": ["prod"] }
  },
  "input_guardrail_ids": ["<prompt-injection id>", "<pii id>"],
  "output_guardrail_ids": ["<secrets id>", "<code-safety id>"]
}
  • Within a rule, every dimension it names must match; a dimension matches when the request’s value is any listed value. An empty dimension matches everything.
  • users matches the person a request belongs to — the session’s user, or the person an API key was issued to.
  • teams matches the request’s team, or a team its person belongs to.
  • keys matches the API key the request arrived on.
  • models matches the requested model or the model an alias points at.
  • metadata matches values from the X-ManyLayers-Metadata request header. A key listed with no values matches any value for that key.
  • Every matching rule applies. Resolution is a union, never first-match-wins. A guardrail selected by several rules runs once.
A workspace rule may bind its own workspace’s guardrails and organization-wide ones; an organization-wide rule may bind only organization-wide guardrails. A rule can also be disabled without deleting it (PATCH with {"enabled": false}). Creating a group in the console also lets you say where it applies (applies: named, workspace or conditions). That is saved as a rule the group owns, listed under Rules.

Selecting guardrails per request

A request can add guardrails with the X-ManyLayers-Guardrails header:
curl https://app.manylayers.io/v1/chat/completions \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -H 'X-ManyLayers-Guardrails: {"input_guardrails": ["production/pii"], "output_guardrails": ["production/secrets"]}' \
  -d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "..."}]}'
Names are <group>/<guardrail> or guardrail ids; the header holds at most 32 names and 8 KB. The header only adds guardrails — it cannot remove one a rule applies. Naming a guardrail requires gateway.guardrails.read, and only guardrails visible from the caller’s workspace resolve; anything else is unknown_guardrail. A group restricted to named users, teams or keys can be selected only by them (403 guardrail_access_denied); this limits who can choose a group, never whether a rule’s guardrails apply. The header is never forwarded to a provider.

Errors

StatusCodeMeaning
403guardrail_blockA guardrail blocked the request (request blocked by guardrail "…") or the response (response blocked by guardrail "…"). The message names the guardrail and the categories that tripped — never the content.
503guardrail_errorAn enforce guardrail could not complete.
503guardrail_timeoutAn enforce guardrail did not finish within its timeout.
503guardrails_unavailableThe guardrail configuration could not be loaded.
400invalid_guardrails_header / unknown_guardrailThe X-ManyLayers-Guardrails header is malformed or names a guardrail the caller cannot use.
403guardrail_access_deniedThe header names a guardrail in a group the caller has not been granted.

Configuring a guardrail

The admin API is served by the console at https://app.manylayers.io. Authenticate with a personal access token (ml_pat_…) and name the Gateway workspace with ?workspace_id=. A group and its guardrails are usually written together, in one call. Either the whole group is stored or none of it is:
curl -X POST "https://app.manylayers.io/admin/gateway/guardrail-groups?workspace_id=$WS" \
  -H "Authorization: Bearer $ML_PAT" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "production",
    "guardrails": [
      {"name": "pii", "type": "pii", "priority": 1, "config": {"categories": ["email", "ssn"]}},
      {"name": "prompt", "type": "prompt_injection", "enforcement_strategy": "enforce_but_ignore_on_error"}
    ]
  }'
PUT /admin/gateway/guardrail-groups/{id} with a guardrails list replaces the group’s guardrails in the same way: an entry with an id updates that guardrail, an entry without one creates a guardrail, and a guardrail you leave out is deleted — which also removes it from every rule. Omit guardrails to rename a group without touching what it holds. A group holds at most 100 guardrails. A failure is reported against the entry it belongs to, for example param: "guardrails[1].operation". Guardrails can also be managed one at a time:
# A group
curl -X POST "https://app.manylayers.io/admin/gateway/guardrail-groups?workspace_id=$WS" \
  -H "Authorization: Bearer $ML_PAT" \
  -H "Content-Type: application/json" \
  -d '{"name": "production"}'

# A guardrail in it
curl -X POST "https://app.manylayers.io/admin/gateway/guardrail-groups/$GROUP/guardrails?workspace_id=$WS" \
  -H "Authorization: Bearer $ML_PAT" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "pii",
    "type": "pii",
    "operation": "mutate",
    "priority": 10,
    "enforcement_strategy": "enforce",
    "message_scope": "all",
    "timeout_ms": 1000,
    "config": {"categories": ["email", "phone", "ssn"]}
  }'
FieldDefaultNotes
namerequiredUnique within the group; at most 128 bytes; no /.
typerequiredFixed once created.
operationthe type’s defaultMust be one the type supports.
prioritythe type’s default0–10000; lower runs first.
enforcement_strategyenforce
message_scopeallall, last, or 1–1000.
timeout_ms0 (1000 ms)Up to 10000.
configthe type’s defaultsType-specific settings, below. Unknown settings are refused.
enabledtrue
Every configuration is validated — and every regular expression compiled — when it is written. A configuration the gateway cannot run is refused with 400 invalid_guardrail_config, naming the offending field, so no request is ever the first to find out. POST …/guardrails/{id}/test with {"text": "...", "hook": "llm_input"} (text up to 64 KB; hook defaults to llm_input) runs a stored guardrail over sample text and returns blocked, mutated, status, verdict, action, findings and — for a mutating guardrail — the rewritten text. Nothing is sent to a model and nothing is recorded.

Settings by type

{"redaction_text": "***REDACTED***", "categories": []}
categories limits detection to a subset (empty means all): aws_access_key, aws_secret_key, openai_api_key, anthropic_api_key, google_api_key, github_token, slack_token, jwt, bearer_token, oauth_token, refresh_token, private_key, rsa_private_key, ssh_private_key, pgp_private_key, database_connection_string, generic_api_key, password, high_entropy_secret. For key=value assignments only the value is redacted.
{"categories": ["name", "email", "phone", "address", "ssn", "passport", "credit_card",
                "date_of_birth", "medical_record_number", "health_information",
                "ip_address", "bank_account", "aadhaar"],
 "redaction": "mask", "mask_char": "*", "placeholder": "[REDACTED:{category}]"}
redaction is mask (one mask_char per character, keeping the shape) or placeholder ({category} expands to the category). Names are detected from context cues (“my name is”, honorifics, “call …”). Aadhaar numbers are redacted when labelled (“Aadhaar”, “UID”) in any 4-4-4 form, or unlabelled when they pass the UIDAI format and Verhoeff checksum.
{"sensitivity": "medium", "threshold": 0, "categories": []}
sensitivity is low (block at score 80), medium (50) or high (30); a threshold of 1–100 overrides it, and 0 uses the sensitivity. Categories: direct_injection, instruction_override, system_prompt_extraction, jailbreak, instruction_hierarchy, indirect_injection, obfuscation. Signals in tool results weigh more, since that is where indirect injection arrives.
{"categories": ["hate", "self_harm", "sexual", "violence"],
 "severity_threshold": "medium", "thresholds": {"violence": "high"}}
Content blocks when its detected severity is at or above the category’s threshold (low, medium, high, critical). thresholds overrides severity_threshold per category.
{"languages": ["python", "shell", "javascript", "sql"], "min_severity": "medium", "disabled_rules": []}
Any finding at or above min_severity blocks. disabled_rules names rules to skip, such as python.eval or shell.chmod_777.
{"block_destructive_statements": true, "strip_sql_comments": true,
 "block_delete_update_without_where": true, "require_parameterization": false,
 "block_injection_patterns": true}
Defaults to validate at priority 1. In mutate mode comments are stripped; if anything blocking remains, the guardrail blocks rather than rewriting the SQL into something it cannot vouch for.
{"preset_patterns": ["pii"],
 "custom_patterns": [{"name": "employee-id", "pattern": "EMP-\\d{6}", "redaction_text": "[EMPLOYEE]", "enabled": true}],
 "redaction_text": "[REDACTED]"}
Presets: ssn, email, us_phone, india_aadhaar, india_pan, passports for India, US, UK, Germany, France, Netherlands, Canada, Australia, China and Japan (india_passport, us_passport, …), visa, mastercard, amex, discover, aws_access_key, aws_secret_key, github_token, slack_token, generic_api_key, ipv4, ipv6, url. Families expand to several: pii, passports, government_id, financial, credentials, network. Patterns use RE2 syntax — matching is linear-time, so there is no catastrophic backtracking — and patterns that are too complex, or that match the empty string, are refused.

Observability

Every guardrail execution is recorded with the request id, organization, workspace, user, team, model, guardrail, type, hook, operation, priority, enforcement strategy, status, verdict, action, execution time, whether a mutation was applied, and its normalized findings. Execution records are kept for audit.retention_days (default 90).
  • In the console, Request Traces → Guardrail executions lists them; a single request’s trace shows its guardrail verdicts.
  • GET /admin/gateway/guardrail-executions?request_id=… returns one request’s guardrail trace in execution order. It also accepts guardrail_id, action, q, since, until, limit and offset, and requires gateway.audit.read.
  • A guardrail execution structured log line is written per execution.
  • manylayers_guardrail_executions_total (labels guardrail_type, hook, status, action) and manylayers_guardrail_execution_seconds (guardrail_type, hook) are exported; see Metrics.
  • Executions that blocked, mutated or found something appear in the request’s audit row, and each hook is a span in exported traces.
Detected values are never recorded. A finding holds a category, a severity and a byte range — never the matched text. When a guardrail flags content it does not redact, that request’s body is left out of the audit row.

Guardrail management API

MethodPathPermission
GET/admin/gateway/guardrail-typesgateway.guardrails.read
GET, POST/admin/gateway/guardrail-groupsread / manage
GET, PUT, DELETE/admin/gateway/guardrail-groups/{id}read / manage
GET, POST/admin/gateway/guardrail-groups/{id}/guardrailsread / manage
GET, PUT, DELETE/admin/gateway/guardrail-groups/{id}/guardrails/{guardrailID}read / manage
POST/admin/gateway/guardrail-groups/{id}/guardrails/{guardrailID}/testread
GET/admin/gateway/guardrail-groups/{id}/accessread
POST/admin/gateway/guardrail-groups/{id}/accessmanage, or group manager
DELETE/admin/gateway/guardrail-groups/{id}/access/{bindingID}manage, or group manager
GET/admin/gateway/guardrail-groups/{id}/activityread
GET, POST/admin/gateway/guardrail-rulesread / manage
GET, PUT, PATCH, DELETE/admin/gateway/guardrail-rules/{id}read / manage
GET/admin/gateway/guardrail-executionsgateway.audit.read
GET/admin/gateway/guardrail-executions/filtersgateway.audit.read
Every route takes ?workspace_id=. Organization-wide resources require gateway.guardrails.manage across the Gateway product to write.

Gateway-wide checks that run before guardrails

Two older checks run on every request independently of guardrails, and guardrails see the request after both:
CheckWhat it doesPage
Prompt injection firewallScores request text; per-team off / audit / enforce.Prompt Injection Firewall
PII detection and redactionReplaces PII before the cache, provider and audit log; optionally reversible.PII Detection & Redaction
A pii or prompt_injection guardrail adds per-workspace, per-rule control on top of these.