Context Engineering Tutorial and System DesignApplication
Chapter 09
9 / 10

Multi-Agent Context Isolation

⏱️ 20 min

Should sub-agent context be shared / isolated / partial sharing? The Anthropic Multi-Agent Research System case + Claude Code's Agent tool implementation + JR omni-report's 17-routine async handoff via git commit

CHAPTER SYSTEM DECISION
01Engineering question

Should sub-agent context be shared / isolated / partial sharing? The Anthropic Multi-Agent Research System case + Claude Code's Agent tool implementation + JR omni-report's 17-routine async handoff via git commit

02Reviewable output

One inspectable section of the Context System Spec, with its source, rule and boundary recorded.

03Definition of done

Run the decision against a real case, preserve the trace and record the condition for continuing.

Chapter 6 solved "how does a single agent run a long task without blowing up context". Production hits another scale — a main agent dispatching sub-agents. The choice between shared, isolated, or partially shared context decides whether the system can scale.

Three Multi-Agent Context Topologies

1. Shared Context — One LLM Runs the Whole Thing

Every step runs in the same messages array.

messages = [
  user: "调研 Kubernetes 三个最流行的 ingress controller 各自优劣"
  assistant: 我先列三个候选 → ingress-nginx, traefik, contour
  tool_use: WebFetch ingress-nginx docs
  tool_result: ...
  tool_use: WebFetch traefik docs
  tool_result: ...
  ... (累积 50K context)
  assistant: 综合分析后, ingress-nginx 优在生态成熟, traefik 优在...
]

Pro: easiest to implement (no state management); zero information loss.

Con: context grows linearly, the Lost in the Middle problem bites around step 30; every step pays the accumulated token cost.

Fits: under 10 steps with strong dependencies between steps.

2. Isolated Context — Sub-Agents Fully Independent

The sub-agent gets a fresh context — just task description + needed reference material. After running, it returns a final answer + short summary, not the process.

主 agent context:
  user: "调研 K8s 三个 ingress controller"
  assistant: 我会派 3 个 sub-agent 并行调研
  tool_use: spawn_subagent(name="research_ingress_nginx", task="...")
  tool_result: { summary: "ingress-nginx 优在生态成熟...", refs: [...] }
  tool_use: spawn_subagent(name="research_traefik", task="...")
  tool_result: { summary: "traefik 优在配置灵活...", refs: [...] }
  ...
  assistant: 综合三个 sub-agent 的总结...

# 每个 sub-agent 自己的 context(独立):
  user: "调研 ingress-nginx:架构、性能、生态、坑"
  tool_use: WebFetch ingress-nginx docs
  ... (自己累积 30K context, 完成后丢弃)
  assistant: 给主 agent 一段 500 字摘要

Pro: main agent context stays small; sub-agents can run in parallel; one blowing up doesn't affect others.

Con: lossy handoff; main can't see intermediate reasoning, has to trust the result; needs spawn / coordinate infrastructure.

Fits: sub-tasks independent and parallelizable; main agent needs to scale past 30 steps.

3. Partial Sharing — Summary Plus Key Raw Data

Sub-agents return a summary, plus the key supporting evidence verbatim. Main agent gets summary + a few key quotes and can verify the sub-agent.

Anthropic's Multi-agent Research System blog pattern — main is Lead Researcher, sub-agents are Search Subagents. Handoff includes source URL + key quotes for verification.

Cost: handoff payload 5-10× larger than pure summary, but trust goes way up.

Anthropic Multi-Agent Research System — Official Case

Anthropic's 2025-04 blog describes the multi-agent research feature they built into Claude.ai:

Architecture:

LeadResearcher (主)
  └─ 拆解用户 query → 决定派几个 sub-agent
  └─ 派 SearchSubagent 1: "找 X 主题的最新 paper"
  └─ 派 SearchSubagent 2: "找 Y 公司的官方文档"
  └─ 派 SearchSubagent 3: "找用户评测和实际案例"
  └─ 收集所有 sub-agent 摘要 → 综合写最终 report

Anthropic's engineering takeaways (direct quotes):

  • "Sub-agent context can't be shared — parallel sub-agents not knowing what others do actually avoids duplicate work and group think"
  • "Optimal sub-agent count is 3-5; past 5, lead coordination cost outweighs the gain"
  • "Sub-agents handing raw search results to lead is wrong — they have to distill into findings first"
  • "The whole system uses 4× more tokens than a single-agent baseline, but the quality lift far outweighs the token cost"

Claude Code's Agent Tool — Sub-Agent in Production

Claude Code's built-in Agent tool is the sub-agent pattern (this course was very likely written by Claude Code — main agent dispatches an Explore subagent to look up prompt-master configs. The main agent doesn't read the 200-line config raw, just the structured summary returned).

Config:

  • subagent_type — different sub-agents get different tool sets (Explore can only read, code-architect can design but not edit)
  • isolation: "worktree" — high-blast-radius tasks dispatch to a sub-agent in a separate git worktree, then merge or discard
  • Description and Prompt — sub-agent only sees its prompt + own tool list at startup

"Isolated context + lossy handoff" productized — main delegates, protects its own context.

JR Real Case: Isolation Across omni-report's 17 Routines

JR Academy's omni-report runs 17 independent routines (AI Visibility / Competitor Weekly / Marketing Topics / Growth Playbook / Daily Jobs ×4 / various daily reports). Each is its own cron job — independent context, git commit, Notion sync.

Why not merge into a single "omni-report master agent"?

Cross-routine flow: async via git commits. Marketing Topics runs Monday and writes to marketing-topics/$DATE.md. Growth Playbook runs Tuesday and Reads last week's report into its own context. "Filesystem as sub-agent handoff channel" — no message queue, git is enough.

Shared / Isolated / Partial — Trade-off

DimensionShared contextIsolated + summaryPartial sharing
Main agent context growthLinear, painful past step 30Almost flatSlow growth (summary + a little raw)
Information loss0Large (only summary)Medium (summary + key citations)
Sub-agent parallelismNo (serial only)Perfect parallelPerfect parallel
Infrastructure0 (one LLM)spawn / coordinate / summarizer+ citation extraction
Token costMedium (linear growth)High (multiple LLMs)Higher (summary + citations)
Trust (can main verify sub)PerfectWeakMedium (can see raw evidence)
Fits< 10 steps, strong dependency30+ steps, parallelizableSerious reasoning, traceability needed

JR's internal rule: under 10 steps single LLM. 10-30 partial sharing. 30+ isolated sub-agent. The curve Anthropic verified building their own research system.

Takeaway

Multi-agent isn't "split as fine as possible". Optimal sub-agent count is 3-5 — past 5, coordination cost eats the gain. Sub-agents have to distill findings before handing back; raw data still blows up the main agent. Anthropic's own research system uses 4× tokens but the quality lift far outweighs cost — proving "more tokens but layered" beats "fewer tokens but tangled".


References

  1. Anthropic. (2025-04-15). How we built our multi-agent research system — Lead + Search subagent architecture and the 3-5 sub-agent rule of thumb.
  2. Anthropic. Claude Code documentation — Agent tool — sub-agent type / isolation / handoff implementation.
  3. Anthropic. (2024-12-20). Building Effective Agents — original orchestrator-workers pattern.
  4. LangGraph. Multi-agent supervisor pattern — open-source equivalent.
  5. AutoGen. GitHub — Microsoft's multi-agent framework, comparative reference.

Production case: JR Academy omni-report — 17 independent routines using git commits for async cross-agent handoff, filesystem as sub-agent communication channel.

📚 Related resources

Common questions

Open a question to review the practical answer.

Should sub-agent context be isolated or shared?

Pick by step count + parallelism: < 10 steps with strong info dependencies use shared (one LLM runs all), 30+ steps that parallelize use isolated (sub-agent has own context, returns summary), serious reasoning needing traceability uses partial sharing (summary + verbatim key quotes). Anthropic Multi-Agent Research System uses partial sharing.

Is more sub-agents always better?

No. Anthropic's own research system data: 3-5 sub-agents is the sweet spot, past 5 the lead agent's coordination overhead outweighs the gains. The whole system costs 4× the tokens of single-agent, but the quality jump on hard problems is far bigger than the token bill.

How do sub-agents communicate with each other?

JR omni-report uses filesystem as the sub-agent handoff channel: 17 independent routines pass info asynchronously via git commits (Marketing Topics finishes Monday → writes marketing-topics/$DATE.md, Growth Playbook on Tuesday Reads it into its own context). 10× simpler than a message queue, enough for low-frequency cross-agent traffic.

What does a multi-agent system cost per day to run?

Anthropic's own research-system numbers: one complex query with lead agent + 4 sub-agents burns ~80K input + ~15K output tokens = $0.30/query on Sonnet. 100 queries/day = $30/day ≈ $900/mo. 4× the single-agent cost, but the accuracy lift on hard problems far outweighs it.

I am not on Claude Code's Agent tool — can I implement multi-agent with LangGraph instead?

Yes: LangGraph is LangChain's official multi-agent framework — nodes = single agents, edges = handoff protocols, state dict = configurable shared / partial / isolated context. OpenAI Swarm (lightweight) and CrewAI (role-driven) are also common picks. All three are model-agnostic.

What is the most common multi-agent failure mode?

Sub-agent info loss: the lead splits a task to sub-agents, sub-agents return only "done," and the lead cannot reason further. Enforce structured returns from every sub-agent (JSON: result + key_findings + sources); refuse free-text responses. This is a hard constraint in Anthropic's Multi-Agent Research System.