A rule in a prompt or in CLAUDE.md is guidance. Claude reads it and usually follows it, but nothing stops a tool call that breaks it. A Claude Code mod is code that runs around the session, and every tool call passes through it before it runs. That makes a mod the place for rules that must hold every time: no deploys from a session, no reading .env, no refunds through an MCP tool. A good guard blocks only the one action and tells Claude why, so Claude finds another way to finish the job instead of stopping. Here is how the layers fit, a small guard mod we validated and tested, and where the limits are.
A prompt guides, a mod guards
Anthropic is direct about this. Claude Code's docs say Claude treats CLAUDE.md files as "context, not enforced configuration," and that there is "no guarantee of strict compliance."1 The same docs split the work in two: settings for technical enforcement, CLAUDE.md for behavioral guidance.1 That is not a flaw in the model. A prompt is text the model weighs against everything else in the conversation. A long session, a vague rule, or two rules that pull in different directions can all tip it the other way. That is why people find their CLAUDE.md rules ignored mid session: the rule was read, then outweighed.
A mod works at a different level. It is a plugin whose JavaScript or TypeScript functions run inside Claude Code. When Claude is about to use any tool, Claude Code fires an event called tool.call and passes it through each mod's hook first. That includes MCP tools and calls a subagent makes.2 The hook can let the call through, change it, or refuse it. If the hook refuses, the tool never runs, and no permission prompt appears. The model does not get a vote, because the check is not in the model.
New to mods? Start with Claude Code mods, explained, then come back here.
The same request, with a CLAUDE.md rule and with a mod
Say your project's CLAUDE.md has the line "Never deploy from this session, deploys go through CI." You ask Claude to fix a bug and ship it. Most of the time Claude respects the rule. But "ship it" is a direct instruction from you, it arrives later than the CLAUDE.md line, and the deploy command is one tool call away. Nothing in the session checks the command against the rule. This is an illustration of the shape, not a recorded incident:
Now the same request with a guard mod loaded. The deploy call reaches the mod's tool.call hook, the hook refuses it, and Claude reads the refusal as that tool's result. The session doesn't stop. Claude takes the route the reason points to:
The rule now lives in code. It holds every time the pattern matches, and the reason turns a dead end into a next step. The CLAUDE.md line is still worth keeping: it saves Claude a wasted attempt. The mod is what makes it a rule.
Claude Code safety layers: permissions, settings hooks, mods
A mod is not the first guard Claude Code has. It sits on top of what Anthropic already ships, and for many rules the built-in layers are enough. Here is what each one does, from softest to strongest.
| Layer | What it can do | Pick it when |
|---|---|---|
| CLAUDE.md or a prompt | Tell Claude what you want. Not enforced. | Conventions, context, preferences |
| Permission rules (allow, ask, deny) | Allow, prompt for, or block a tool or a command pattern such as Bash(rm *) or Read(./.env). Deny is checked first and wins. | The rule is a fixed command, path, or tool name |
| Settings hooks (PreToolUse, PostToolUse) | Run a script on each call. PreToolUse can deny with a reason Claude sees and can change the input. PostToolUse can replace the output. | You want a script you already have to decide |
| Mods (tool.call, tool.check) | Everything a hook can, inside Claude Code, plus state across calls, asking you mid call, answering a call itself, fail closed handling, UI, and tests. | The rule needs memory, judgment, a question to you, or must not fail open |
| Managed settings | The same layers, set by an admin, which users can't override | A rule must hold for the whole team |
Start with a deny rule when the rule is fixed. Rules are checked in the order deny, ask, allow, and the first match in that order decides, so an allow rule can't carve an exception out of a deny rule.3 They also understand shell operators, so a deny on Bash(rm *) still catches rm inside cd /tmp && ... or a subshell.3 Anthropic's own guidance for a fixed command or path is a permission rule, not code.2
Settings hooks already block with a reason. To be fair to them: a PreToolUse hook can return permissionDecision: "deny" with a permissionDecisionReason that Claude sees, or exit with code 2 and a message on stderr. It can also rewrite the tool's input, and a PostToolUse hook can replace the output.4 If a shell script does what you need, keep it.
What a mod adds is mostly about rules that are not just a pattern:
- Memory across calls. A mod's hooks share the variables in its file.5 One hook can note that Claude connected to production, and another can refuse every write after that.
- A question in the middle of a call. A
tool.callhook can hold the call, ask you with$.ui.ask, and only then run or refuse it.2 - An answer without the tool. A hook can return its own
result, so the tool never runs and Claude reads what the hook returned.2 - Decisions from live state. The
tool.checkevent fires after the permission rules and settings hooks have decided, and can turn their answer into allow, ask, or deny based on something true right now, such as the current Git branch.2 - Fail closed. A
PreToolUsesettings hook that times out does not block the call, which goes on through the normal permission check.4 A mod hook that throws is skipped too, unless you attach a.catchhandler that answers with a deny instead.2 - Tests.
claude plugin testfires tool calls through your hooks with no session, so you can prove the guard denies what it should.6
Build a small guard mod
This mod, house-rules, guards three things: deploy commands and forced recursive deletes in Bash, secret files in Read, Edit, and Write, and money-moving tools on an MCP server named payments. Each refusal says why and what to do instead. It is three small files plus a test:
A few things to notice. The second argument to on is a matcher: a string, an array, or a regular expression, so /^mcp__payments__/ picks out one MCP server's tools.2 Returning { deny } without calling next ends the call there. Calling next(e) passes it on to the permission check and the tool. The .catch on the Bash hook means a bug in the check refuses the command rather than letting it through.
Then a test, which fires tool calls through the hooks with a stub standing in for Claude Code:
We ran both commands on Claude Code 2.1.287. claude plugin validate prints what the mod hooks and calls, which is also what a reviewer should read before installing anyone's mod.7
The one warning is about the manifest, not the code: plugin.json has no author field, which only matters for attribution and doesn't stop the mod from loading.
The tests earned their keep. Our first version matched .env with a looser pattern, and the test showed it also blocked .env.example, the very file the deny message tells Claude to read. A guard that blocks its own escape route turns a nudge into a dead end. The validate and test runs above are real.
To try it, put the folder anywhere and start a session with claude --plugin-dir ./house-rules. Mods need Claude Code 2.1.287 or later.5
What to guard
Good candidates share one trait: the damage is hard to undo, and there is a safe way to get the same answer. Block the risky action, and point Claude at the safe one.
| Guard | Block | Point Claude at |
|---|---|---|
| Secrets | Reading or editing .env, keys, credential files; secrets leaving through a tool | .env.example, a vault reference, or asking you |
| Production databases | INSERT, UPDATE, DELETE, DROP, ALTER on a prod connection | A SELECT, or the exact SQL for a person to run |
| Payments and bank details | Refund, payout, and transfer tools; card or IBAN numbers in prompts or tool input | Read only tools, and a written description of the action |
| Specific MCP tools | Write tools on a server you only want to read from | The same server's read tools |
| Deploys | terraform apply, kubectl apply, prod deploy commands | A commit and a hand off to CI |
| Deleting files | Forced recursive deletes, deletes outside the project | A list of what to remove, for you to confirm |
Two of these need more than a deny. Secrets and payment data are often a question of what text enters the session or leaves through a tool, so the fix is to rewrite it, not refuse it. That is the subject of redacting sensitive data in Claude Code. Production databases have their own traps, and they get a full walk through in keeping Claude Code read only on your production database.
For deletes, Anthropic's sample mod blast-radius takes a softer line: it holds a risky shell command such as rm -rf or a force push, shows what it would change, and gives you buttons to proceed or cancel.5 Holding and asking is often better than a flat no when the action is sometimes right.
Write the reason for Claude, not for you
The deny text is the only thing Claude learns about why its call failed. Anthropic's docs put it plainly: write it as an instruction Claude can act on.2 A guard that says "blocked" makes Claude guess, retry with another spelling, or give up. A guard that says what to do next keeps the work moving.
- Name the rule. "Deploys run from CI" tells Claude this is policy, not a broken tool, so it doesn't retry.
- Give the safe route. "Use .env.example," "use a SELECT," "describe the refund for a person to make."
- Prefix the mod's name. When you read the transcript later, you know which guard fired.
- Block narrowly. Refuse the one action, not the whole tool. A guard that blocks all of Bash ends most tasks.
Enforce it for a whole team
A guard on your own laptop is a habit. For a team, it has to come from managed settings, the configuration an admin deploys by file, MDM, or the claude.ai admin console.8 Three pieces matter.
The built-in guard. Claude Code ships a mod called sec-default (listed in /plugin as cc-plugin-sec-default). It loads ahead of every user's mod when the machine has managed settings or the user is signed in with a Team or Enterprise plan, and users can't turn it off.8 It protects what the organization manages: a user's mod can't change what managed hooks receive or decide, the system prompt, managed CLAUDE.md, or managed MCP servers' tools. Where it loads, a user's mod also can't approve a call that a deny rule refuses.8 It adds no other limits, so it is not a guardrail for your own rules. That part is yours.
Your own mod, first in line. A mod counts as the organization's when managed settings enable it from a marketplace directory on the machine, by absolute path. List it in prependPlugins and it sees every event before any user's mod.8 Set allowManagedModsOnly on the built-in guard and users' own mods don't load at all, while their settings hooks and status lines keep working.8
Managed hooks, still on top. A PreToolUse hook in managed settings runs before any mod sees the call, and its block is final. If a mod rewrites the call, the managed hooks run again on the new version.8 So the strongest setup layers all three: managed deny rules for fixed patterns, a managed hook or mod for the rules that need logic, and allowManagedModsOnly so nobody's personal mod can approve around them.
Honest limits
A mod is a strong layer. It is not a sandbox, and it is not a guarantee. Know where it stops:
- It only sees tool calls. A guard on Bash reads the command text.
./scripts/release.shcan runterraform applyinside it, and the guard never sees that line. - Patterns miss spellings. Our
rmrule catchesrm -rfbut notrm -r -f. Anthropic's docs point out the same kind of gap in their own example, which missesgit push -f.2 Prefer allow lists on what is safe over deny lists of what is dangerous where you can, and test every pattern. - It can be turned off.
claude --safe-modestarts a session without installed mods, and that includes mods your organization deploys through managed settings. Managed settings hooks and permission rules still apply in safe mode, and built-in mods such assec-defaultkeep running.589disableAllHooksin a user's own settings turns off the mods they installed in every session.5 If the worker that runs installed mods crashes three times, Claude Code unloads every mod that isn't built in, yours included, until/reload-pluginsor a new session.8 - Mods aren't sandboxed. A mod runs as you, with your files, processes, and network. If you turn on sandboxing, it isolates the Bash commands Claude runs, and a process a mod starts runs outside it.5 Read
claude plugin validateoutput before you install anyone's mod. - Other mods can approve too. Without managed settings, a user's mod that approves calls can approve one an
askrule would prompt for, or one a non managedPreToolUsehook blocked.8 - The real fix is often upstream. A read only database user, a protected branch, and a payments key without refund scope stop the action for every client, not just Claude Code. Use the mod to give Claude a clear reason. Use least privilege to make the action impossible.
Where Fixter fits
Guardrails decide what Claude may do. They don't tell you what happened in production after the code shipped. Fixter is monitoring for teams that build with coding agents: you send your logs and traces over standard OpenTelemetry, and Fixter finds the issues and bugs in your system and tells you what broke and why. You pull them into Claude Code over MCP and fix them there. A read only MCP server for your telemetry is also a nice fit for the guard pattern above: Claude can look at production all day without touching it.
Key takeaways
- CLAUDE.md is context, not enforcement; use it to guide, not to forbid
- A mod's
tool.callhook sees every tool call, MCP and subagent calls included, before it runs - Fixed rules belong in deny rules; settings hooks already deny with a reason
- A mod adds state across calls, asking mid call, fail closed handling, and tests
- Block one action and say what to do instead, so Claude still finishes the task
- For teams, deploy the mod through managed settings and set
allowManagedModsOnly
Frequently asked questions
Is a rule in CLAUDE.md enforced?
No. Anthropic's docs say Claude treats CLAUDE.md as context, not enforced configuration, and that there is no guarantee of strict compliance. Claude usually follows a clear rule, but nothing stops a tool call that breaks it. For anything that must never happen, use a deny rule, a settings hook, or a mod, which run in code around the session.
Why does Claude Code ignore my CLAUDE.md rules?
Because CLAUDE.md is context, not configuration. Claude Code delivers it as a user message after the system prompt, and Anthropic's docs say there is no guarantee of strict compliance, especially for vague or conflicting instructions. If two rules contradict each other, Claude may pick one arbitrarily. Run /context to check the file loaded, then make the rule specific and remove conflicts. If the rule must hold every time, move it into a deny rule, a settings hook, or a mod.
What is a Claude Code guardrail mod?
A mod is a plugin with JavaScript or TypeScript hooks that run inside Claude Code. A guardrail mod handles the tool.call event, which fires before every tool Claude uses, including MCP tools and calls from subagents. It can refuse the call with a reason Claude reads, change the arguments, hold the call while it asks you, or let it through.
Can't a settings hook already block a tool call?
Yes. A PreToolUse settings hook can deny a call with a reason Claude sees, and can change the tool's input. For a fixed rule, a deny rule in settings is simpler still. A mod adds things on top: state shared across calls, holding a call while it asks you, answering a call without running the tool, a fail closed handler when the check itself breaks, and automated tests with claude plugin test.
Does a guard stop Claude from finishing the task?
Not if you write it well. A deny refuses one tool call, not the session. Claude reads the deny text as that tool's result, so if the text says what to do instead, such as use a SELECT, commit and hand off the deploy, or read .env.example, Claude usually takes that route and keeps going.
Can my team enforce guardrail mods for everyone?
Yes, through managed settings. An admin can install the organization's own mod from a directory on each machine, run it first with prependPlugins, and set allowManagedModsOnly on the built-in sec-default guard so users' own mods don't load. A PreToolUse hook in managed settings runs before every mod and its block is final.
Are guardrail mods a security boundary?
Treat them as a strong layer, not a guarantee. Mods aren't sandboxed, a user can start a session without installed mods using --safe-mode (that includes mods deployed through managed settings, while managed hooks and permission rules still apply), a hook that throws is skipped unless you add a fail closed handler, and pattern matching on command text misses other spellings. Pair a mod with deny rules, sandboxing, and least privilege credentials.
Sources
- Claude Code docs, How Claude remembers your project (CLAUDE.md is context, not enforced configuration; settings for enforcement)
- Claude Code docs, React to events with a mod (tool.call, deny, matchers, $.ui.ask, tool.check, .catch)
- Claude Code docs, Configure permissions (deny, ask, allow order; compound commands)
- Claude Code docs, Hooks reference (PreToolUse deny with a reason, updatedInput, PostToolUse output, timeouts)
- Claude Code docs, Mods overview (what a mod can reach, turning mods off, sample mods, built-in mods)
- Claude Code docs, Test a mod (claude plugin test, stubs)
- Claude Code docs, Create a mod (files, claude plugin validate)
- Claude Code docs, Manage mods for your organization (sec-default, allowManagedModsOnly, prependPlugins, managed hooks, safe mode)
- Claude Code docs, CLI reference (what --safe-mode turns off and what managed settings still apply)