Context Engineering Tutorial and System DesignEngineering
Chapter 06
6 / 10

Agent Memory — Three Layers and Their Tool Stacks

⏱️ 25 min

Three memory layers by lifecycle (scratchpad / working / persistent) + Mem0 (arXiv:2504.19413) / Letta / LangChain tool comparison + Letta's counterintuitive benchmark (plain filesystem at 74% beats specialized vector stores)

CHAPTER SYSTEM DECISION
01Engineering question

Three memory layers by lifecycle (scratchpad / working / persistent) + Mem0 (arXiv:2504.19413) / Letta / LangChain tool comparison + Letta's counterintuitive benchmark (plain filesystem at 74% beats specialized vector stores)

02Reviewable output

One inspectable section of the Context System Spec, with its source, rule and boundary recorded.

03Definition of done

Run the decision against a real case, preserve the trace and record the condition for continuing.

An agent runs 5 steps fine. Run 50 and the context blows — most people's first instinct is "switch to a 1M context model." But running 50 steps on 1M costs 100× more than 5 steps on 200K, and Lost in the Middle gets worse, not better, at 1M.

The right move is giving the agent layered memory, not buying more context budget.

Three layers of memory — split by lifetime

Borrowing OS terminology. The Letta project (GitHub) is built on this OS metaphor:

LayerLifetimeOS analogyWhat goes inSize
ScratchpadOne taskCPU registertool-call intermediate results, temp vars1-10K
WorkingOne sessionRAMconversation history, task plan, user input5-50K
PersistentAcross sessionsDiskuser prefs, learned facts, past task summariesUnbounded

Each layer has a different toolchain — mixing them up is why context blows.

Layer 1: Scratchpad — task-internal scratch

What goes in: search-tool return value, intermediate markdown table from step two, step three's reasoning picking top 3.

Scratchpad uses no memory service — lives inside the task's messages array. After every tool call, the result becomes a tool_result message and the next LLM turn sees it.

  • Once the task ends, scratchpad gets thrown out
  • Token ceiling 10K; over that, split into sub-tasks (chapter 9)
  • If a task ran 30 steps and scratchpad still bloats — task needs splitting, not memory upgrading

JR omni-report's "Phase 0/1/2/3/4" is exactly the scratchpad pattern — every phase commits, the next re-reads from git instead of the previous phase's scratchpad. Avoids contamination.

Layer 2: Working memory — session-scoped rolling

What goes in: every turn of the conversation, current multi-step task plan, session-level config (language preference, verbosity).

Simple: rolling window

Keep only the last N turns; truncate older ones. LangChain's ConversationBufferWindowMemory:

# tested: 2026-04-26 · langchain@0.3.x
from langchain.memory import ConversationBufferWindowMemory
memory = ConversationBufferWindowMemory(k=10)  # 只留最近 10 轮

Good for: short conversations, Q&A assistants. Bad for: recalling 30-turn-old details.

Advanced: rolling window + history summary

Recent N turns verbatim; compress earlier with cheap model (Haiku), stuff into system prompt. LangChain's ConversationSummaryBufferMemory, LlamaIndex's ChatSummaryMemoryBuffer both do this.

Catch: summaries lose info. Turns with exact numbers, citations, code snippets can't be compressed.

Anthropic native: Message Batches API

100K+ turns (rare) — Anthropic Message Batches API async-processes the whole conversation history as RAG-on-history.

Layer 3: Persistent memory — cross-session fact store

What goes in: user prefs (language, style, expertise), learned facts (works at X, runs Y, follows Z), summaries of past task outcomes (not raw output).

The biggest gap between production agents and toy demos. Only persistent memory lets an agent "remember" across sessions.

Mem0 — arXiv:2504.19413

Two-phase extract-update: at the end of every conversation, extraction pulls salient facts; update phase asks an LLM to add / update / delete / no-op a memory entry.

Mem0 benchmarks: vs OpenAI default thread memory, accuracy 26% higher, p95 latency 91% lower, token cost 90%+ lower.

# tested: 2026-04-26 · mem0ai@0.1.x
from mem0 import Memory
m = Memory()
m.add("用户在悉尼工作,做 AI engineer", user_id="alice")
results = m.search("alice 在哪个城市", user_id="alice")
# → "用户在悉尼工作"

Letta — GitHub

Heavier OS three-layer (core + archival + recall). Letta benchmark: plain filesystem (one markdown file per user) hit 74% accuracy and beat plenty of specialized vector store memory libraries — simple usually works.

OpenAI Threads / Anthropic Conversation API

OpenAI Assistants thread comes with persistent context built in. Anthropic has no thread equivalent, but prompt caching 5-min TTL + your own conversation-history storage equals the same effect.

DIY — Filesystem / DB

One markdown file per user; at conversation end, LLM appends summary. On retrieval, read whole file into context. Letta benchmark proves this beats vector store when user-level state ≤ 10K characters.

Real JR case: skills-data-manager and prod-state.json

JR Academy skills-data-manager (tools/skills-data-manager/) handles bootcamp curriculum sync — edit locally, diff prod, sync one click. Persistent memory:

curriculum/{bootcamp-slug}/public/prod-state.json
  └─ 每次从 prod 拉数据回来时缓存的「最后已知 prod 状态」
  └─ next diff 用本地内容 - prod-state 算变化
  └─ 跨 session 持久化,避免每次都拉 prod

"Filesystem as persistent memory" in the wild — one JSON file is the persistent memory, no vector store, no Mem0. Simple, debuggable, git-trackable. Only when entry count balloons past 1000+ and semantic recall becomes necessary do you upgrade to Mem0/Letta. Don't pre-empt for "we might need it someday."

Mem0/Letta service vs DIY filesystem — Trade-off

DimensionManaged (Mem0 / Letta)DIY filesystem
Onboarding< 1 day1-3 days
Recall quality26% above baselineTied < 100 entries; falls behind > 1000
Latency100-500ms5-50ms
Cost$0.01-0.10/MAU + tokenNear zero
DebuggabilityBlack box100% readable
Data complianceLeaves companyControllable
Multi-userBuilt-in user_id isolationDIY
Best for100+ user × frequent sessions< 100 user / internal / POC

JR rule: start filesystem, run 3 months, watch entry count + recall quality, migrate to Mem0 only after crossing threshold. Letta's 74% filesystem benchmark is the basis.

Takeaway

Memory isn't one thing — it's three: scratchpad (within-task) / working (within-session) / persistent (cross-session). Mix them up and the context blows. Production agents must layer — scratchpad lives in the messages array, working uses rolling window + summary, persistent starts on filesystem and graduates to Mem0.


References

  1. Mem0 team. (2025-04-28). Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. arXiv:2504.19413.
  2. Mem0 research benchmark. Benchmarking Mem0's token-efficient memory algorithm — 26% acc / 91% latency / 90% token 节省 vs OpenAI baseline.
  3. Letta. Letta documentation — OS-inspired 三层 memory(core / archival / recall)+ filesystem 74% benchmark surprise.
  4. LangChain. Migrating from ConversationSummaryBufferMemory — 滚动 + 摘要 working memory pattern.
  5. Anthropic. Message Batches API — 异步处理超长 conversation.
  6. mem0ai. GitHub — Mem0 开源实现.

Production case: JR Academy skills-data-manager(tools/skills-data-manager/)— prod-state.json 作为 filesystem persistent memory.

📚 Related resources

❓ Common questions

Open a question to review the practical answer.

Why can't agents just use a 1M context model and call it a day?

1M context does not fix it: 200K → 1M is 5× the cost, and Lost in the Middle is worse at 1M (the blind midsection grows wider). The right direction is layered memory on the agent (scratchpad / working / persistent), not more context budget.

Mem0 vs Letta — which fits my project?

Pick by user count: < 100 users use filesystem, > 1000 users use Mem0. Mem0 uses a two-stage extraction-update flow, vs OpenAI baseline scores 26% higher accuracy, 91% lower p95 latency, 90% token cost savings. Letta uses OS-style three layers (core/archival/recall) but its own benchmarks show plain filesystem at 74% beating the vector store.

Where is the line between working memory and persistent memory?

Split by lifecycle: Working = rolling context inside current session (use LangChain ConversationSummaryBufferMemory to summarize earlier turns), Persistent = cross-session fact store (user prefs, learned facts, past task summaries). Session ends → working clears, persistent survives to the next.

For a single-user assistant product, do I need all 3 memory layers?

No: a single-user assistant needs only working + persistent. Use LangChain ConversationSummaryBufferMemory for working (auto-summary), use SQLite + a hand-written schema for persistent (user prefs / facts). Scratchpad only matters when an agent does multi-step reasoning; pure chat skips it.

Are Mem0 / Letta expensive to run monthly?

Mem0 self-hosted is free (Apache 2.0 OSS); Mem0 Cloud free tier gives 1K memories + 1K searches/month, enough for demo. Letta is 100% OSS self-hosted, just PostgreSQL + one Python service — cloud cost is the PG instance (~$15/mo on RDS t4g.micro). Neither blows up your bill.

What is the most common memory failure mode?

Dumping every LLM output into persistent memory — three months in, the store balloons to 50K+ entries and every search returns stale/contradictory content. Correct: gate writes through an LLM-as-judge "should-remember" check, add expiration (30/90 days), periodically run dedup + contradiction merging. Mem0 ships these by default; if you roll your own, add them.