Skip to content
← Articles

How I Split Claude Code Work Across Three Agents

Revised 8 min read

By Jake Bauman

claude-code / ai-agents / multi-agent / build-notes / context-engineering

Dark editorial graphic: Three Claude agents working as a team on AI infrastructure

Three weeks ago I was building a feature for Personably and hit the same wall I hit every time. One Claude Code session does everything. It plans the architecture, writes the code, reviews the output, and fixes its own mistakes. As the session grows, it becomes harder to distinguish current instructions from earlier decisions and discarded approaches.

I had read about the agent teams pattern before. Strategy agent in Opus, execution agent in Sonnet, review agent in another Sonnet. Each one gets a clean context window and a narrow job. But I had been ignoring it because spinning up three sessions sounded like more work, not less.

Last week I finally tried it on a real feature. On that feature, the review agent found issues I might have missed. This is one experience, not a measured comparison of rework across projects.

Here is how I set it up, an illustrative token-cost calculation, and the CLAUDE.md template I use to give all three agents shared instructions.

The problem with one agent doing everything

The standard Claude Code workflow looks like this: you open a session, describe the feature, Claude plans it, writes the code, you review, Claude fixes, repeat. One context window absorbing every decision, every mistake, every correction.

Three problems compound inside that single window:

Context degradation. As a session grows, earlier decisions compete for attention with later revisions. The LangChain 2026 State of Agent Engineering survey reports that 32% of respondents cited quality as a top production barrier. It does not establish that context length caused those quality problems.

Self-review blindness. Asking one agent to review its own code is like asking a writer to proofread their own article. It reads what it meant to write, not what it actually wrote. The same blind spots that produced the bug will also miss the bug.

Context cost compounding. Every wasted token in the planning phase carries forward into execution and review. Long transcripts can make later steps carry context that is no longer useful. Measure your own token use before assigning a percentage.

The three-agent architecture

I adapted a plan, execution, and independent review pattern for my own stack. The handoff files and role boundaries below are my implementation.

Strategy Agent (Opus 4.6, $5/$25 per M tokens)
    |
    |-- produces: plan.md, architecture decisions, file list
    |
    +--> Execution Agent (Sonnet, $3/$15)
    |       |
    |       |-- produces: implemented code, passing build
    |       |
    |       +--> Review Agent (Sonnet, $3/$15)
    |               |
    |               |-- produces: review comments, diff of fixes

Each agent gets exactly one input file and produces exactly one output file. No agent edits another agent's work directly. The files are the interface.

Strategy agent receives: the feature description and relevant codebase context. It produces a plan.md with architecture decisions, a file list of every file that needs to change, and the order of changes.

Execution agent receives: plan.md and the current codebase. It implements every change in order, runs the build, and produces a done.md listing every file changed and the build result.

Review agent receives: plan.md, done.md, and the full diff. It checks for alignment with the plan, obvious bugs, and violations of project conventions. It produces review.md with required fixes and optional suggestions.

I do the final merge. The agents never push. The review agent's output goes back to the execution agent for fixes if needed, or to me for approval if clean.

What a sample token calculation costs

The cost math is the part that made me finally try this. I was afraid it would triple my Claude bill. It does not.

Here is an illustrative calculation using assumed token counts and the published Opus 4.6 and Sonnet 4.6 standard API rates checked on September 29, 2026. This is not an invoice or measured per-feature average. Check the provider's current pricing and your actual usage before budgeting:

| Agent | Model | Tokens In | Tokens Out | Cost | |-------|-------|-----------|------------|------| | Strategy | Opus 4.6 | 8,000 | 1,500 | ~$0.08 | | Execution | Sonnet | 12,000 | 2,000 | ~$0.07 | | Review | Sonnet | 10,000 | 800 | ~$0.04 | | Total | | 30,000 | 4,300 | ~$0.19 |

The arithmetic shown is $0.1855, or about $0.19, for these assumed tokens at the cited standard rates. Tool use, caching, other billable tokens, and different model selections can change a real bill. The earlier $3.50 headline and $0.34 calculation used inconsistent and outdated rates and had no supporting run log, so I removed them. I have not published a measured cost or rework comparison for this workflow.

A smaller model may be sufficient for some steps; choose by task quality and current price. The token counts above show how to do the calculation, not how much every feature will use.

The shared brain: one CLAUDE.md for three agents

The split can give each role a narrower context. But narrow is not the same as uninformed. If all three agents do not share the same rules, conventions, and project context, the review agent will flag things as errors that are actually intentional choices the strategy agent made.

The solution is a shared CLAUDE.md that every agent reads at session start. One shared instruction file can help the roles follow the same project conventions.

Here is the template I use. Copy this, strip what you do not need, and save it as CLAUDE.md in your project root.

# CLAUDE.md -- Project Name

## Identity
You are an AI coding agent working on [PROJECT NAME], a [WHAT IT IS].
You serve [WHO THE USERS ARE].

## Goals
1. [Concrete outcome 1]
2. [Concrete outcome 2]
3. [Concrete outcome 3]

## Rules
- Never push to main. Open a PR.
- Run `npm run lint` and `npm run build` before declaring done.
- No em dashes in copy. Use periods, commas, or sentence breaks.
- [Any project-specific conventions]

## Workspace
- `src/app/` -- Next.js App Router pages
- `src/components/` -- Shared React components
- `src/lib/` -- Library code, types, utilities
- `content/` -- MDX articles and data files
- `public/images/` -- Static assets

## When You Are the Strategy Agent
- Your ONLY job is to produce `plan.md`.
- Read the feature request and the relevant source files.
- Output: architecture decisions, file list, change order.
- Do NOT write code. Do NOT review anything.

## When You Are the Execution Agent
- Your ONLY job is to implement `plan.md`.
- Follow the change order exactly.
- Run the build. If it fails, fix it before declaring done.
- Output: `done.md` with every file changed and build result.

## When You Are the Review Agent
- Your ONLY job is to review the diff against `plan.md`.
- Check for: alignment with architecture, missed files, convention violations, obvious bugs.
- Output: `review.md` with required fixes and optional suggestions.
- Do NOT rewrite code. Flag issues for the execution agent.

The role-specific sections are the part I had to figure out myself. I iterated on these role instructions during my own use. Adapt them to your repository and review the resulting work.

What broke and what I fixed

Problem 1: The review agent rewrote code instead of flagging it. The first few runs, the review agent would produce a full corrected implementation instead of a review document. I added the explicit instruction: "Do NOT rewrite code. Flag issues for the execution agent." That fixed it.

Problem 2: The strategy agent tried to write code. Opus is smart. It sees a problem and wants to solve it. I had to add "Do NOT write code. Do NOT review anything." to the strategy section. The strategy agent produces a plan. That is its entire job. Giving it permission to stop at the plan is counterintuitive but necessary.

Problem 3: Context gating. My first attempt gave the execution agent the full CLAUDE.md plus the plan plus the full codebase. Context burn. Now I give each agent only what its role needs. The strategy agent gets the feature description and relevant source files. The execution agent gets the plan and the files to change. The review agent gets the plan, the change list, and the diff.

When to use this (and when not to)

The three-agent pattern makes sense for features that touch more than three files or involve architectural decisions. For a single-file bug fix, it is overkill. One Sonnet session is fine.

Here is my decision heuristic:

| Feature size | Pattern | Why | |-------------|---------|-----| | 1 file, no architecture change | Single Sonnet session | Overhead not worth it | | 2-3 files, simple changes | Single Sonnet, self-review | Marginal benefit from split | | 3+ files, architecture decisions, or data model changes | Three-agent team | Context cost of single session exceeds team overhead | | Cross-cutting changes across 5+ files | Three-agent team + one review round | Worth the extra review cycle |

The threshold is not about lines of code. It is about context complexity. If the agent needs to hold six files in its head to make a decision, split the work. The strategy agent can hold all six files, make the decision, and hand a clean plan to the execution agent that only needs to see one file at a time.

The copyable part

The CLAUDE.md template above is yours to use. Strip the sections that do not apply to your project. Add your own conventions. The role-specific sections at the bottom are the part I had to figure out through trial and error, and they are the part that makes the pattern actually work.

If you try this, start with one feature. Do not reorganize your entire workflow at once. Pick a feature that touches three or four files, write the plan by hand once so you know what good looks like, then turn the agents loose and compare their output to yours. Compare the additional setup and review time with the defects each approach catches.

I am still iterating on the handoff format between agents. The current version uses markdown files in a .claude/plans/ directory. I may eventually build a small CLI tool that manages the three files and routes them automatically. But for now, the files are the interface, and that simplicity is keeping the whole thing from collapsing into a context management problem of its own.

For a starting workflow, see the free Growth Agent Starter Kit. For an explicit independent review step, see the Quality Gate.

Get the free Starter Kit.

Enter your email to see the download link on this page. Build Notes is optional and has a separate checkbox.

Your files appear here after signup. If email delivery is available, you may also receive your requested file by email. The newsletter is optional. Privacy.

Related reading