All modules Module II — Harness Anatomy →

Module I — From Model to Agent System

Phase 1 · HARNESS — Module 1
Status: Authored & Verified.
Canonical Equation: $\text{AGENT SYSTEM} = \text{MODEL} + \text{HARNESS} + \text{ENVIRONMENT}$
Lab: ../labs/L1-harness-swap.md (Student Package: hefesto-lab1.zip)

Executive Positioning

The object of study in HEFESTO is not prompt crafting, nor is it waiting for model providers to release larger foundation checkpoints. The object of study is:

The system that controls the agents: harnesses, runtimes, governance, and control planes.

Holding the model fixed and changing the harness changes performance materially. The empirical result of an autonomous system depends on $\text{Model} \times \text{Harness}$, not on either one in isolation.

       ┌──────────────────────────────────────────────────────────┐
       │                       AGENT SYSTEM                       │
       │                                                          │
       │   ┌───────────────┐     ┌───────────────┐     ┌──────┐   │
       │   │     MODEL     │  ×  │    HARNESS    │  ×  │ ENV  │   │
       │   │ (Probabilistic│     │(Deterministic │     │(Host │   │
       │   │  Hypotheses)  │     │ Control Plane)│     │ /OS) │   │
       │   └───────────────┘     └───────────────┘     └──────┘   │
       └──────────────────────────────────────────────────────────┘

Component 00 — Presentation, Vision & Macro Roadmap

1. The Prototype Collapse

A naive while True: loop wrapped around a long prompt looks convincing in a demo. However, upon encountering the first runtime exception in production, ungoverned systems inevitably collapse into stochastic panic.

2. The Agent Necessity Test

More agents do not equal more intelligence. Each additional agent introduces quadratic coordination latency, token burn, and correlated failure modes.

3. Ten-Phase Curricular Path

1. MODEL → AGENT → HARNESS (Anatomy, Control Loop, Environment) 2. CONTEXT + STATE + TOOLS (Admission Firewalls, Memory Tiers, Recovery) 3. ROLE ENGINEERING (Separation of Role, Model, Harness, Provider) 4. MULTI-AGENT ORCHESTRATION (Topologies, Handoff Contracts, Anti-patterns) 5. VERIFICATION + FAILURE ENGINEERING (Independent Judges, Fault Injection) 6. MULTI-HARNESS (Adapters, Unified Control) 7. CONTROL PLANE (Design of Deterministic Planes) 8. SECURITY + OBSERVABILITY (Capability Gates, Forensic Lineage) 9. COST / TOKEN OPTIMIZATION (Equal-Compute Benchmarking, SAS vs. MAS) 10. CAPSTONE (Governed Multi-Harness Enterprise System)


Component 01 — Probabilistic Model (The Latent Engine)

1. The Nature of the Latent Engine

From a formal mathematical standpoint, a foundation language model is a conditional probability density function over a closed vocabulary:

$$P(w_t \mid w_1, w_2, \dots, w_{t-1})$$

The model does not compile code, compute logic, or verify database integrity. It projects statistical likelihoods conditioned exclusively on its active context window.

2. Stochastic Nature vs. Deterministic Software

Absolute determinism cannot be guaranteed in distributed LLM clusters, even at temperature = 0.0. Non-associative floating-point operations across parallel GPU tensors and speculative decoding algorithms produce minute variations in logits, which can alter greedy token selection at close decision boundaries.

3. The Three Structural Voids

A foundation model in isolation suffers from three systemic deficiencies: 1. Statelessness (Zero Persistence): The model retains no state between independent API invocations. 2. Absence of an Internal Clock: The model executes synchronously upon receiving an input tensor and shuts down immediately upon emitting the end-of-sequence token. 3. Absence of Native I/O: The model cannot open network sockets, inspect filesystems, or execute operating system syscalls.

4. The Commodity Inversion

In amateur development, the model is the center of the architecture. In senior systems engineering:


Component 02 — The Agent Loop (Agent Loop & FSM)

1. The Formal Finite State Machine

The canonical agent loop is modeled as a formal 6-tuple $M = \langle S, \Sigma, A, \delta, s_0, F \rangle$:

 [Observation] ──> [Proposal] ──> [Contract Validation]
                         │                 │ (Fail)
                         │ (Pass)          └──> [Structured Feedback]
                         ▼
                   [Isolated Dispatch]
                         │
                         ▼
               [Observation Sanitizer] ──> [State Transition]

2. Stochastic Entrapment & Collapse Attractors

When an ungoverned loop feeds raw stack traces into context upon tool failure, recency bias in self-attention inflates the conditional probability of repeating the exact same failed action. As action entropy $H(A)$ collapses toward zero, the system becomes trapped in a cyclic collapse attractor ($P(\text{auto-recovery}) < 5\%$ after 3 unmitigated failures).

3. Mitigation Oracles

1. Action Similarity Oracles: Cryptographic hashing and Levenshtein distance metrics detect syntactic and semantic loops before dispatching LLM API calls. 2. Cycle Budgets: Hard ceilings on loop turns ($K_{\max}$) and execution timeout windows. 3. Circuit Breaker Pattern: Trips open after $N$ consecutive tool failures, halting dispatch and forcing replanning or human escalation.


Component 03 — The Environment (Sandboxing & Blast Radius)

1. Operational Asymmetry

The model produces only text tokens; the environment executes physical state mutations on host filesystems, networks, and databases. An uncontained model connected directly to an OS shell invites catastrophic system corruption.

2. Blast Radius Containment (Saltzer & Schroeder)

Adhering to the Principle of Least Privilege, execution operates across three concentric zones:

3. Ephemeral Git Worktrees & Copy-on-Write

The host workspace state $S_{\text{host}}$ is never directly mutated:

$$S_{\text{host}}^{(t+1)} = S_{\text{host}}^{(t)} + \mathcal{O}_{\text{verify}}(\Delta_{\text{worktree}})$$


Component 04 — Tools & Dispatcher (4-Stage Pipeline & Idempotency)

1. Actuation Interface

Dumping raw shell strings from an LLM into eval() or a bash subshell is an architectural anti-pattern that permits shell injection and untyped errors.

2. The 4-Stage Dispatcher Pipeline

1. Proposal: Grammar-guided JSON payload generated via strict JSON schema envelopes. 2. Contract Validation: In-memory validation against Pydantic/Zod schemas with extra='forbid'. 3. Isolated Dispatch: Execution in unprivileged subprocesses with argument vectors decoupled from the shell interpreter. 4. Observation Sanitization: Compressing stderr/stdout and isolating return codes.

3. The Law of Idempotency ($f(f(x)) = f(x)$)


Component 05 — Context Management (Attention Economics)

1. Context Window as Volatile RAM

The context window is finite, expensive, and volatile memory with quadratic attention latency. Treating it as an infinite storage dump degrades cognitive fidelity.

2. Attention Degradation & Lost-in-the-Middle

In scaled sequence lengths, the softmax denominator $\sum_{j} \exp(q \cdot k_j^T / \sqrt{d_k})$ inflates, dispersing attention weights away from keys located in the middle third of the context.

3. The 3-Tier Admission Firewall

1. Tier 1 (Structural Gatekeeper): Rejects malformed payloads and enforces strict byte limits. 2. Tier 2 (Deduplication & State Hashing): Eliminates duplicate tool outputs across cycles. 3. Tier 3 (Relevance Masking): Extracts strictly impacted AST nodes and lines of interest.

4. Active Compaction & Pruning Pipeline

Transient tool observations are deterministically pruned or replaced with structured entity hashes once their outcome has been reduced into the authoritative state.


Component 06 — Memory & State Persistence (4 Tiers & ACID State)

1. The Persistence Duality

2. Deterministic State Transitions

$$S_{t+1} = \mathcal{T}(S_t, A_t, \Omega_t)$$

Natural language summaries incur stochastic drift ($\Delta H(S) > 0$). Industrial state is maintained in strongly-typed schemas (SQLite / Pydantic) enforcing ACID transaction boundaries.

3. The Four Memory Tiers

1. Ephemeral (Scratchpad): Task-level working buffer flushed upon milestone completion. 2. Episodic: Chronological, immutable audit log of actions, observations, and decisions. 3. Semantic: Long-term associative knowledge base (domain entities, indexed documentation). 4. Procedural: Immutable policies, tool contracts, and behavioral playbooks.

4. Checkpointing ($C_k$) & Event Sourcing

Checkpoints bind system state tuples $\langle S_k, \text{git\_sha}_k, \text{env\_hash}_k \rangle$. State is derived by executing reducers over an append-only event ledger, enabling instantaneous, deterministic rollback upon verification failure.


Lab L1 — Bridge to Practice

Theoretical comprehension culminates in the Harness Swap Experiment: