Long prompt, unstable output
Rules, evidence and task instructions compete in the same unstructured block.
A longer prompt will not fix stale retrieval, noisy memory or overloaded tools. Learn to give the model the right evidence at the right step — with a budget, trace and evaluation gate.
These failures often look like model problems. They are usually selection, ordering, memory or permission problems.
Rules, evidence and task instructions compete in the same unstructured block.
The model still follows a louder distractor or misses evidence in the middle.
Old decisions and stale facts enter every turn without an expiry or write policy.
Schemas consume attention while permissions and failure boundaries stay implicit.
Five sources enter the call. Selection and budget decide what survives. Trace and evaluation decide whether the system is safe to release.
Stable rules, identity and non-negotiable boundaries.
Only the capabilities required for this step, with permission limits.
Session state and durable facts with lifecycle rules.
Fresh, attributable passages selected for the current question.
The immediate request, constraints and definition of done.
Use one real assistant, RAG flow or agent throughout the track. Every chapter adds reviewable evidence to the same system specification.
Complete this artifact inside the chapter workspace.
02 · InventoryComplete this artifact inside the chapter workspace.
03 · SelectComplete this artifact inside the chapter workspace.
04 · BudgetComplete this artifact inside the chapter workspace.
05 · GroundComplete this artifact inside the chapter workspace.
06 · RememberComplete this artifact inside the chapter workspace.
07 · ActComplete this artifact inside the chapter workspace.
08 · EvaluateComplete this artifact inside the chapter workspace.
Follow the core gates in order. The tool comparison and agent chapters extend the design after the foundation is stable.
Separate the task from the context that supports it.
Allocate space, retrieve broadly and select deliberately.
Control what persists and which tools can act.
Inspect real agent strategies, then validate your own system.
The finished spec explains what enters context, why it is trusted, how much space it receives, what can act, and how failures are detected.