The prompt injection firewall is a fast, always-available check that scores inbound request text against a table of weighted heuristics before anything else inspects it. Use it as a deployment-wide safety net for instruction overrides, system-prompt exfiltration and jailbreak framing, with per-team control over whether a hit is only recorded or actually blocked.

How it works

  1. After authentication and policies admit the request, the gateway collects every message’s text content plus prompt and input.
  2. Each rule is matched across that text. A rule contributes its score once per request, however many times or segments it matches.
  3. The scores are summed. If the total reaches firewall.threshold, the firewall trips.
  4. The team’s mode decides what happens: nothing (off), a recorded finding (audit), or a 403 (enforce).
The firewall runs before PII redaction, so it sees the request as sent. Findings record only the rule name, score and byte offsets — never the matched text.

Rules

RuleScoreSignals
instruction_override60”ignore/disregard/forget … previous/system … instructions”, “your new instructions are”
prompt_exfiltration60 / 50”reveal/print/repeat … system prompt / hidden instructions”, “what is your system prompt”
role_confusion60 / 50Chat-template tokens (<|im_start|>, [INST], <<SYS>>), smuggled "role": "system"
jailbreak_persona50”DAN mode”, “do anything now”, “jailbreak”, “evil mode”, “no longer an AI”
safety_bypass50”without/bypass … restrictions/filters/guardrails”, “you are now unrestricted”
invisible_chars50Zero-width characters, bidi overrides, Unicode tag characters and other invisible format characters
system_line25A line starting system:
developer_mode25”developer mode”, “god mode”
base64_payload25A base64 blob of 80+ characters that decodes
With the default threshold of 50, any one strong rule trips the firewall on its own, while the 25-point rules only trip in combination.

Configuration

firewall:
  default_mode: audit   # off | audit (default) | enforce
  threshold: 50         # summed score that trips the firewall (default 50)

teams:
  - name: support-bot
    firewall_mode: enforce   # "" = use firewall.default_mode

Key fields

firewall.default_mode
string
default:"audit"
Mode for every team that does not set its own. Any value other than off, audit or enforce stops startup.
firewall.threshold
integer
default:"50"
The summed score at which a request trips. Lower is stricter.
teams[].firewall_mode
string
Per-team override: off, audit or enforce. Empty inherits default_mode.
ModeOn a trip
offThe firewall does not run.
auditThe request continues. Rule hits are stored on the request’s audit record and the metric is incremented.
enforceThe request is refused with 403 before PII redaction, guardrails or any provider call.
These values are read from gateway.yaml by the gateway process. The per-team override is stored on the team and is set through the teams API below (or teams[].firewall_mode in gateway.yaml).

What a block returns

HTTP/1.1 403 Forbidden

{
  "error": {
    "message": "request blocked by prompt-injection firewall",
    "type": "permission_error",
    "code": "prompt_injection"
  }
}
The blocked request is still written to the audit log with status 403, its findings (source firewall, the rule name and score), and a redacted copy of the body when bodies are recorded. No tokens are consumed.

Common configurations

Keep default_mode: audit, watch manylayers_firewall_events_total and the audit findings for a week, then switch high-risk teams to enforce.
curl -X POST https://app.manylayers.io/admin/teams \
  -H "Authorization: Bearer ml_pat_..." \
  -H "Content-Type: application/json" \
  -d '{"name": "support-bot", "cache_enabled": false, "firewall_mode": "enforce"}'
POST /admin/teams creates the team or, if an organization team with that name exists, updates it. The update replaces monthly_token_budget, rpm_limit, tpm_limit, cache_enabled, firewall_mode and default_config_id with what you send, so include the team’s current values; models changes only when you send it. It requires org.teams.manage and refuses (409) a name used by a workspace team.
Set firewall_mode: off on that team only; everyone else keeps the default.
Lower threshold to 25 so every single signal, including a bare system: line or a long base64 blob, trips the firewall. Expect more false positives on technical prompts.

Metrics

MetricLabelsMeaning
manylayers_firewall_events_totalteam, modeRequests that tripped the firewall. mode is audit or enforce.
sum by (team, mode) (rate(manylayers_firewall_events_total[5m]))

Firewall vs. the prompt_injection guardrail

Firewallprompt_injection guardrail
ScopeWhole deployment, per-team modePer workspace or organization, applied by rules
ScoringSummed across the whole requestPer message segment
TuningOne thresholdsensitivity (low 80, medium 50, high 30), threshold 1–100, categories
Modesoff, audit, enforceenforce, enforce_but_ignore_on_error, audit
Block response403 prompt_injection403 guardrail_block
They are independent detectors and can run together: the firewall is a cheap first line that runs on every inference request; the guardrail gives finer, per-workspace control and a richer detector set (indirect injection, obfuscation, instruction hierarchy).
The firewall inspects request text only. It does not scan model output or tool results returned in later turns unless they are sent back as message content.

Next steps

Guardrails

Configure the prompt_injection guardrail per workspace.

PII Detection & Redaction

The step that runs right after the firewall.

Metrics

Every Prometheus series the gateway exports.

Request Logging & Audit

Find blocked requests and their findings.