Module III — Context Engineering for Agents
Status: outline. Lecture body not authored.
Context stops being a prompt and becomes a system resource.
Topics
- context budgets and token allocation
- retrieval and relevance filtering
- compaction and structured summaries
- context contamination and leakage
- context inheritance across handoffs
- working context vs durable state
- memory admission gates and freshness policies
Memory, Context, State & Evidence
A rigorous harness enforces strict architectural separation between four distinct primitives:
MEMORY ≠ CONTEXT ≠ STATE ≠ EVIDENCE
- Context (Active Inference Context): The finite token window assembled specifically for an active inference pass. Subject to attention degradation and strict budget constraints.
- Memory (Persistent Informational Storage): Long-term, non-authoritative storage of past events, observations, or semantic facts across sessions. Memory *informs*, but cannot dictate execution.
- State (Operational Ground Truth): The authoritative, deterministic machine state (phase pointer, active step, variables, locks). State dictates *what the system is doing right now*.
- Evidence (Verifiable Observation): Verifiable records and artifacts used to substantiate claims about execution (tool outputs, process exit codes, hashes, test assertions).
The Memory Paradox
Core Principle:
More memory ≠ better reasoning.
Uncontrolled or incorrectly scoped memory can degrade inference.
Injecting retrieved memory into an LLM context is not cost-free. A retrieved memory item may be historically authentic and semantically similar, yet operationally toxic to the current reasoning step.
Failure Modes of Unbounded Memory
- Context Rot: As context expands with historical data, model attention across middle positions degrades (*lost in the middle*), obscuring immediate task constraints.
- Context Contamination: Precedents from previous iterations inject obsolete assumptions into fresh tasks.
- Excessive Context: High signal-to-noise ratio drops reasoning fidelity below single-shot zero-memory baselines.
- Stale Memory: Facts that were true at time $T_0$ become false at $T_1$ after code, environment, or requirement mutations.
- Cognitive Traps & Reasoning Fixation: The model anchors on past intermediate attempts, repeating flawed logic rather than exploring clean solutions.
- Belief Distortion: Semantic similarity in vector space measures embedding proximity, not operational authority, logical validity, or execution correctness.
Fundamental Axiom:
Retrieval relevance does not imply operational authority.
OLD VALID MEMORY (True at T0)
↓
RETRIEVAL (High semantic similarity)
↓
CURRENT CONTEXT (Injected without gate)
↓
ANCHORING (Fixation on obsolete premise)
↓
WRONG CURRENT DECISION
Memory Admission
Before any retrieved memory candidate is admitted into active working context, the harness must apply an admission filter:
MEMORY CANDIDATE
↓
scope check (Does this memory apply to the active task domain?)
↓
freshness check (Is this memory invalidated by newer events or clock?)
↓
authority check (Who produced this memory? Is it an unverified claim or evidence?)
↓
contradiction check (Does it conflict with active state or verified evidence?)
↓
WORKING CONTEXT (Admitted into current inference window)
*(Control plane enforcement and automated policy gating are developed in Phase 7 · Module IX).*
Unique process — Context Fidelity Test
Evaluate context degradation and mutation across chained agent handoffs:
Agent A
↓ handoff
Agent B
↓
Agent C
The test measures five distinct vectors across handoff boundaries:
1. Information Survival: Percentage of essential task parameters retained. 2. Information Distortion: Semantic drift, hallucinated additions, or altered constraints. 3. Stale-State Propagation: Leakage of superseded or invalidated states downstream. 4. Contamination: Injection of irrelevant intermediate reasoning into clean agent contexts. 5. Authority Preservation: Maintenance of the boundary between verified facts and speculative inferences.
Retaining more tokens does not guarantee preserving operational fidelity. A harness that prunes aggressively often outperforms one that accumulates historical context indiscriminately.