Back to all posts
Security

How to Give Claude Code Your Business Context, Not Just Your Conventions

How to give Claude Code your business logic and security rules: what CLAUDE.md, rules, skills and hooks each carry, and how to mine, verify and deliver them.

On this page
  1. What does "context" mean for a coding agent?
  2. What are the limits of CLAUDE.md, AGENTS.md and skills?
  3. CLAUDE.md: read once, followed when it fits
  4. AGENTS.md: the same file, shared by more agents
  5. Skills: loaded when the agent thinks of them
  6. What none of them does
  7. Why does the agent lose the rule before it writes?
  8. CLAUDE.md, rules, skills, MCP or hooks: which one carries what?
  9. How do you deliver a rule at the moment of the edit?
  10. Where does the do-it-yourself version stop?
  11. What does the complete method look like?
  12. Step 1: mine the rules from your own code
  13. Step 2: have a human verify every rule
  14. Step 3: deliver the right rule at the edit
  15. Step 4: check the result at the end of the chain
  16. What changes when the rule arrives at the edit?
  17. Where do you start this week?
  18. Frequently asked questions
  19. How do I give Claude Code context about my codebase?
  20. What is the difference between skills and hooks in Claude Code?
  21. What is the difference between rules and skills in Claude Code?
  22. Does Claude Code support rules?
  23. Are Claude Code hooks useful?
  24. Why does my PreToolUse hook run but Claude ignores its output?
  25. What is context engineering for coding agents?
  26. Can Claude Code generate a CLAUDE.md automatically?
  27. Does Claude Code read AGENTS.md?

Four steps in a row: rules mined from the code, verified by a human, delivered to Claude Code at the moment of the edit, then checked at the end of the chain.

Claude Code already knows how to write a refund endpoint. What it cannot know is that yours answers 422 REFUND_EXCEEDS_CAPTURED, writes an audit event with the actor's id, and never crosses from one tenant to another. That is context: your business logic, your names for things, your security rules. The model's intelligence is no longer the bottleneck. Getting that context to it, at the moment it writes, is. This guide covers where each kind of context belongs, a hook you can copy today, and the four-step method CybeDefend uses to do it for a whole codebase: mine the rules from your code, have a human verify them, deliver the right one at the edit, check the result.

What does "context" mean for a coding agent?

Everything the model needs to know about your system that it cannot infer from the code in front of it. Anthropic calls the discipline context engineering, and Birgitta Böckeler, writing about coding agents on Martin Fowler's site, quotes the shortest definition going: "Context engineering is curating what the model sees so that you get a better result." For a coding agent working on a real product, that context comes in three layers.

Conventions

How code looks here: the framework patterns, the folder layout, the test command. A model picks up most of this from the surrounding code, and a short CLAUDE.md covers the rest.

Business logic

What your software must do that no model can infer: a refund never exceeds the captured amount, a manager approves above a threshold, an order status is called SHIPPED and not DISPATCHED. Your nomenclature lives here too.

Security rules

The rules with a regulator or a customer behind them: which fields are personal data, what may appear in a log, which query must be scoped to the tenant, which action needs an audit event.

The first layer is the one every guide to CLAUDE.md is about. The second and third are where an agent's mistake costs money, and where no scanner looks, because a refund that credits too much is valid code. That class of flaw is covered in business logic flaws in AI-generated code. This guide is about preventing it at the source: making sure the agent knows the rule when it writes.

What are the limits of CLAUDE.md, AGENTS.md and skills?

Each was built for a real job, and each does it. None was built to carry a company's business rules to the moment of the edit. Here is where each one stops, in its vendor's own words.

CLAUDE.md: read once, followed when it fits

Anthropic's documentation is candid about it. CLAUDE.md files are loaded "at the start of every conversation", and "Claude treats them as context, not enforced configuration". The same page asks you to target "under 200 lines per CLAUDE.md file", because "longer files consume more context and reduce adherence": two hundred lines for your build commands, your conventions and every business rule you have. And "if two rules contradict each other, Claude may pick one arbitrarily", while nothing checks the file against the code it describes.

Users report the result in the issue tracker. Issue #2901, opened in July 2025 with 31 comments, reads: "Claude Code frequently violates explicit project and user instructions defined in CLAUDE.md files". Issue #33603, still open, is titled "CLAUDE.md hard rules and persistent memory instructions consistently ignored".

AGENTS.md: the same file, shared by more agents

AGENTS.md is the open format that Codex, Cursor, Copilot and others read, used by over 60,000 open-source projects. Its own site sums up its nature: "AGENTS.md is just standard Markdown", "a README for agents". The standard settles where the file lives, not whether it is followed.

The only rigorous test so far is not flattering. In February 2026, Thibaud Gloaguen, Martin Vechev and three co-authors published Evaluating AGENTS.md: context files "do not generally improve task success rates, while increasing inference cost by over 20% on average". The instructions in them were "well followed", and the authors keep the files for "specifying non-standard coding practices", while repository overviews are "not helpful".

Three practical limits come on top:

  • Claude Code skips it when a CLAUDE.md exists. With both files in the repository, Claude Code reads "your CLAUDE.md files only" by default. A team that keeps both has its AGENTS.md skipped by Claude Code, while the other agents on the team read it.
  • Codex caps it. Codex stops adding instruction files once their combined size reaches 32 KiB by default.
  • Copilot asks for it short. The prompt GitHub suggests for writing your instructions says they "must be no longer than 2 pages" and "must not be task specific". A business rule is task specific by nature.

Skills: loaded when the agent thinks of them

A skill is a folder of instructions and scripts that loads only when it is used. Anthropic's skills documentation explains how Claude chooses: it reads the skill's description to "decide when to apply the skill". A rule packaged as a skill therefore applies only if the agent recognises the moment.

The documentation lists the ways that recognition fails. The description is "truncated at 1,536 characters". With many skills installed, Claude Code "drops some descriptions to fit the listing's character budget, which removes the keywords Claude needs to match your request". And after a long session is compacted, "older skills can be dropped entirely". Skills are excellent for procedures. As a carrier for rules that must hold every time, they depend on the one thing you are trying not to depend on. And a skill someone else wrote runs code in your session, which is its own risk: see are Claude Code skills safe?.

What none of them does

All three assume the hard part is already done. None of them:

  • finds your rules. Each one holds only what someone remembered to write down.
  • knows which rule applies to this edit. They are loaded by session, by path or by description, never by what the code being written actually does.
  • notices when a rule stops matching the code. A stale value stays in the file until someone happens to read it.
  • checks the result. Whether the rule was honoured is left to code review.
  • works the same across your agents. A team on five agents maintains five dialects.

Why does the agent lose the rule before it writes?

Because a model's attention is not uniform, and a rule read at the start is far from the edit by the time it matters. Anthropic's engineering team puts it plainly: context "must be treated as a finite resource with diminishing marginal returns", and it names the effect context rot, "as the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases".

Independent work measured it. Chroma tested 18 models in its Context Rot report and found that "model performance degrades as input length increases, often in surprising and non-uniform ways". The older Lost in the Middle study showed performance "significantly degrades when models must access relevant information in the middle of long contexts". And IFScale, which piles instructions on top of each other, found that at 500 simultaneous instructions "even the best frontier models only achieve 68% accuracy".

Security gets no special treatment. In the SusVibes benchmark, 57% of the solutions produced by SWE-Agent with Claude Sonnet 4 were functionally correct and only 11.8% were secure, and "augmenting the feature request with vulnerability hints" did not fix it. Telling an agent to be careful is not the same as giving it your rules.

Our own measurement took the hardest case, rules where the agent already has a plausible answer of its own. A realistic rules section in CLAUDE.md produced 7 exact rule specifics out of 55, exactly what no file at all produced. The full protocol is in does Claude Code follow CLAUDE.md?. Two forces explain it. Dilution: by the thirtieth edit, the file is an old fragment under tens of thousands of tokens of code and output. And habit: the code in front of the agent shows it how things are done here, and it compiles.

Here is what that looks like on one prompt. On 27 September 2026 we asked Claude Code (Haiku model) to write a refund function in an empty repository, once without any rule and once with the hook shown further down, carrying our refund rule. One run each: an illustration, not a measurement.

Same prompt, one run each
No rule delivered
Rule delivered at the edit
Over-refund check
Present
Present
Error returned
A prose message with the amounts
The team's own error code
Audit event
None
The exact event, with the actor

The lines the rule changed, in the second run:

if (amountCents > availableForRefund) {
  throw new RefundError("REFUND_EXCEEDS_CAPTURED");
}
order.refundedCents += amountCents;
await audit.log("refund.created", {
  orderId: order.id, amountCents, actorId,
});

Both versions are reasonable code. Only one is your code. That gap, the right behaviour with the wrong specifics, is the one a reviewer misses and an API client breaks on.

CLAUDE.md, rules, skills, MCP or hooks: which one carries what?

Each mechanism reaches the model at a different moment, and the moment decides what it is good for.

MechanismWhen it reaches the modelGood forWeak for
CLAUDE.md or AGENTS.mdOnce, at session startBuild commands, layout, conventions, non-standard practicesA specific value needed thirty edits later
.claude/rules/ with pathsWhen Claude reads a file matching the patternRules tied to one directory or file typeRules triggered by what the code does rather than where it sits
SkillsWhen Claude decides a skill matches the taskProcedures and playbooks you would otherwise pasteRules that must apply whether or not the agent thinks of them
MCP toolsWhen the agent decides to call themLarge corpora, search, live dataAnything that depends on the agent remembering to ask
HooksOn an event, every time: session start, prompt, before or after a toolThe right rule at the edit; blocking a commandNothing, except that you write the matching logic

The column that matters is the second one. A file, a skill or an MCP tool all depend on something happening earlier or on the model choosing to look. A hook is run by the harness, whether or not the model thinks of it. Path-scoped rules sit in between: Anthropic's documentation says they "trigger when Claude reads files matching the pattern, not on every tool use", which is a real improvement over one big file, and still keyed on a path rather than on the operation.

For the safety side of skills, which run code from someone else's repository, see are Claude Code skills safe?.

How do you deliver a rule at the moment of the edit?

With a PreToolUse hook on the write tools, which looks up the rules for the file being written and returns them as additionalContext. Three files. First, register the hook in .claude/settings.json:

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Edit|Write",
        "hooks": [
          { "type": "command", "command": "node \"$CLAUDE_PROJECT_DIR/.claude/hooks/rules-at-edit.mjs\"" }
        ]
      }
    ]
  }
}

Then write each rule as a small file in .claude/edit-rules/, with the paths it applies to on its first line:

applies-to: src/billing/**, src/api/refunds/**
Refunds: never refund more than the captured amount.
Answer 422 with the code REFUND_EXCEEDS_CAPTURED, never a prose message.
Every refund writes an audit event: audit.log("refund.created", { orderId, amountCents, actorId }).

And the hook itself, .claude/hooks/rules-at-edit.mjs (Node 22.5 or later, for path.matchesGlob):

import { readFileSync, readdirSync } from "node:fs";
import path from "node:path";

const root = process.env.CLAUDE_PROJECT_DIR ?? process.cwd();
const input = JSON.parse(readFileSync(0, "utf8"));
const file = path.relative(root, input.tool_input?.file_path ?? "");
const dir = path.join(root, ".claude/edit-rules");

const rules = readdirSync(dir)
  .map((name) => readFileSync(path.join(dir, name), "utf8"))
  .filter((text) => {
    const globs = text.match(/^applies-to:\s*(.+)$/m)?.[1] ?? "";
    return globs.split(",").some((g) => path.matchesGlob(file, g.trim()));
  });

if (rules.length > 0) {
  // JSON, not plain text: on PreToolUse, Claude Code only reads additionalContext.
  process.stdout.write(JSON.stringify({
    hookSpecificOutput: {
      hookEventName: "PreToolUse",
      additionalContext: `Rules for ${file}:\n\n${rules.join("\n---\n")}`,
    },
  }));
}

That is the setup behind the right-hand column of the refund comparison above. Claude Code places hook context "next to the tool result", so the agent reads the rule with the result of its write to that file and corrects it on the spot. In our run it wrote the function, read the rule, and answered "Updated the function to" before listing the error code and the audit event.

The other agents have their own dialects. OpenAI's Codex documents the same shape on PreToolUse: "to add model-visible context without blocking, return hookSpecificOutput.additionalContext". Cursor's preToolUse can allow, deny or rewrite a call, and its postToolUse hook adds additional_context "after the tool result". GitHub Copilot's repository instructions scope a file with an applyTo pattern, and the nearest AGENTS.md in the tree takes precedence.

Where does the do-it-yourself version stop?

At the rules themselves. The hook fixes the moment, and it is worth installing today. It leaves the rest of the list above standing: the rules still come from someone's memory and still go stale, nothing checks whether they were honoured, and each agent in the team needs its own version. It also adds a limit of its own, because paths are not intent: a refund can be written in utils.ts, and a glob cannot tell a refund from a formatter. Past ten rules and one agent, keeping this up by hand becomes a job of its own.

What does the complete method look like?

It starts where every file stops: with no rule written at all. This is the method CybeDefend runs with VibeDefend, from the codebase to the diff, and it is built for a development team rather than for one person's laptop. The rules live in one place per project, every developer's agent receives the same verified set, and each step exists because one of the limits above does.

Mine the rules from your codeA human verifies each oneDelivered at the editChecked at the end
The rule starts in your code, passes a human, reaches the agent at the write, and is checked at the end of the chain.

Step 1: mine the rules from your own code

Your code already contains your rules, written as repetition. Once a repository is connected and parsed, a miner reads its code graph and looks for five kinds of regularity:

  • Presence patterns: a field or a call that almost every instance carries. Every entity has an organizationId, every write to orders calls the audit logger.
  • Value sets: the literal names your code uses for statuses, roles and error codes. This is your nomenclature, and an agent that invents a sixth order status breaks every consumer of the other five.
  • Co-occurrence: guards and decorators that always travel together, like an authentication check and a rate limit.
  • Required calls before sensitive operations: permission checks, tenant scoping, rate limits, feature flags.
  • Conventions named by a model from clusters of similar code, with the outliers listed.

Each proposal comes with its evidence, where the pattern was found and how often, and a confidence. The outliers are the most useful part. One endpoint without tenant scoping in a codebase where forty others have it is either an exception someone meant or a flaw nobody noticed. That is exactly where business logic and security meet.

Step 2: have a human verify every rule

Mining proposes, it never decides. A pattern can be a habit rather than a rule, and a rule is policy: whoever writes it steers every agent in the team, which is why instruction files are an attack surface in their own right. In VibeDefend, each proposal is accepted, edited or rejected in the dashboard, or from inside the agent when a session starts. A rule suggested by the agent itself is refused unless you confirmed it. A weekly drift check proposes an update when the code moves away from an accepted rule, which is how the rules stay true without anyone maintaining a file.

Step 3: deliver the right rule at the edit

This is the hook from the previous section, with retrieval instead of globs. Before each write, the file path and the start of the code being written are sent as the intent, and up to five business rules and five security rules relevant to that change come back as context for that file. The installer wires it into Claude Code, Cursor, Codex, Windsurf and GitHub Copilot in VS Code, each with the limits of its own hook system.

Delivery informs the agent, it does not force it. For what must never happen, a separate guard checks every command before it runs and can block it. The rule in the context is a strong suggestion; the guard on the action is the control. We would rather say which one is which.

Step 4: check the result at the end of the chain

Three checks, from the session to the architecture:

  • At the end of a session, a review asks one question: did this work reveal a durable rule that is written nowhere? It proposes one at most, often none, and nothing is recorded without your yes.
  • Before the commit, the diff is scanned for vulnerabilities, infrastructure misconfigurations and secrets, so the agent fixes them in the same session.
  • On the business logic itself, BLSA, our Business Logic Security Analysis, reads the whole architecture segment by segment to find what no pattern can: a refund that can be replayed, a tenant that can read another's data, personal data that leaks through an export, an operation that is not idempotent. BLSA is a research track run with the CNRS and the CRIStAL laboratory, live with design partners and not generally available yet. See what it looks for in fintech.

Put side by side, the difference is less about the agent than about the team around it.

For a development team
Rules files and skills
Mine, verify, deliver, check
Where rules come from
Someone's memory
Your own code
Who approves them
Whoever edits the file
A human, rule by rule
When the agent sees them
At session start, or by chance
At the edit they govern
When the code changes
Nobody notices
A weekly drift check
Across agents
One file per tool
One set, five agents
Result checked
In review, if spotted
Session, diff, business logic

What changes when the rule arrives at the edit?

Most rules stop being lost. Our controlled study ran 30 developer tickets through three autonomous agents, on one codebase with one model. Over the 19 tasks of its first phase, the agent with no rules and the agent with a realistic rules file each implemented 7 of 55 rule specifics exactly. The agent that received the rules at the edit implemented 46 of 53.

13%

exact with no rules at all (7 of 55)

13%

exact with a realistic CLAUDE.md rules section (7 of 55)

87%

exact with the rule delivered at the edit (46 of 53)

What this measures, and what it does not. It measures delivery at the edit: the 49 rules were written by us, not mined, so it says nothing yet about how good mining is. It is one codebase, one run per arm per task, graded by blind model auditors. And it contains its own counter-example: on one task a rule was delivered thirteen times and the agent still broke it, which is why the guard exists. The protocol and every number are in the study.

Where do you start this week?

With ten rules, not a platform.

  1. List ten rules your team repeats in code review. The comments you have typed more than twice.
  2. Sort them by one question: could a competent engineer new to the team guess it? If yes, it belongs in CLAUDE.md. If no, it has to reach the agent at the edit.
  3. Scope what is tied to a directory with .claude/rules/ and a paths field.
  4. Add the hook above for the arbitrary ones, and check that it returns JSON.
  5. Turn anything that must never happen into a deny rule or a guard, not a sentence.
  6. Test it where it fails: one feature that touches a rule, thirty turns into a session.

When the list stops fitting in your head, that is the moment to mine it instead of writing it.

Frequently asked questions

How do I give Claude Code context about my codebase?

In three layers. Put build commands, layout and conventions in a short CLAUDE.md, which Claude reads at the start of every session. Scope directory-specific rules in .claude/rules/ with a paths field, so they load when Claude reads a matching file. Deliver business and security rules the model cannot guess, error codes, thresholds, audit formats, at the moment of the edit, with a PreToolUse hook that returns JSON additionalContext.

What is the difference between skills and hooks in Claude Code?

A skill is something Claude chooses to load, a hook is something Claude Code runs whether Claude chooses or not. A skill's description sits in context and its body loads when Claude, or you, invoke it, which makes it good for procedures. A hook fires on an event, such as a session start, a prompt or a tool call, which makes it the right place for a rule that must reach the agent every time, or for blocking a command.

What is the difference between rules and skills in Claude Code?

Rules are instructions, skills are procedures. Files in .claude/rules/ load into context unconditionally, or when Claude reads a file matching their paths pattern. A skill packages a multi-step procedure, with optional scripts, that loads only when it is used. Use rules for what must always be true of the code, skills for how to carry out a task.

Does Claude Code support rules?

Yes. Markdown files in .claude/rules/ are loaded as instructions, recursively, one topic per file. A rule with a paths field in its frontmatter only loads when Claude reads a file matching one of its glob patterns; a rule without one loads in every session. Personal rules in ~/.claude/rules/ apply to every project on your machine.

Are Claude Code hooks useful?

For anything that has to happen every time, they are the only reliable mechanism. Anthropic's own documentation says that to block an action "regardless of what Claude decides", you should use a PreToolUse hook rather than an instruction. Hooks can also add context at a precise moment, such as the rules for the file being written, which a file read at session start cannot do.

Why does my PreToolUse hook run but Claude ignores its output?

Because plain text printed by a PreToolUse hook goes to the debug log, not to the model. Claude Code only adds plain-text stdout to the context for UserPromptSubmit, UserPromptExpansion, SessionStart and PostModelSwitch. Return JSON instead, with hookSpecificOutput.hookEventName set to PreToolUse and your text in additionalContext. We checked the difference on Claude Code 2.1.282 on 27 September 2026.

What is context engineering for coding agents?

Deciding what information reaches the model's context window, and when, so it can do the task right. For a coding agent that means less about writing a longer prompt and more about timing: the stable facts at session start, the relevant rules at the moment of the edit, and nothing else competing for attention. Anthropic's engineering team uses the term for the discipline as a whole.

Can Claude Code generate a CLAUDE.md automatically?

Yes, /init analyses your codebase and writes a starting CLAUDE.md with build commands, test instructions and the conventions it discovers. That is a good start for the first layer of context. It is not the same as mining business rules: the Evaluating AGENTS.md study found that model-generated context files did not generally improve task success, and a generated file still arrives once, at session start, and still needs someone to verify it.

Does Claude Code read AGENTS.md?

Yes, when there is no CLAUDE.md. Anthropic's documentation says Claude Code reads AGENTS.md as your project instructions if there is no CLAUDE.md or CLAUDE.local.md in the working directory or above it. With both files present, it reads your CLAUDE.md files only by default, unless your CLAUDE.md imports AGENTS.md or you change the setting. Either way it is the same mechanism, a file loaded at session start, with the same limits for rules needed thirty edits later.

Install VibeDefend in 5 seconds.

One command wires every coding agent on your machine to CybeDefend: your business rules, your compliance frameworks, and guards that block destructive calls before they fire.

Install in 5 secondsNode 18.17+
npx -y @cybedefend/vibedefend@latest install
Auto-detects
  • Claude CodeClaude Code
  • CursorCursor
  • OpenAI CodexOpenAI Codex
  • WindsurfWindsurf
  • GitHub CopilotVS Code Copilot