← Module XI — Security All modules Module XIII — Cost & Tokens →

Module XII — Observability & Evidence

Phase 9 · OBSERVABILITY & EVIDENCE — Module XII
Status: Authored & Empirically Verified.
Lecture Components: 6 FHD 1080p master videos + Lab L12.
Canonical Core Axioms:

AGENT TRAJECTORY + ARTIFACT LINEAGE + DECISION EVIDENCE
STATE + EVIDENCE + ADMITTED CONTEXT + RETRIEVED MEMORY ──> DECISION ──> ARTIFACT / ACTION
OBSERVABILITY MUST CAPTURE NOT ONLY WHAT ENTERED THE CONTEXT, BUT ALSO MATERIAL CANDIDATES REJECTED FOR GOVERNANCE REASONS.

1. Beyond Traditional Logging: The Agent Trajectory Paradigm

In classical software engineering, logging systems such as syslog, ELK, or Datadog were engineered under the assumption of deterministic code branches: every log line reflects a pre-compiled, predictable code path.

When applied to autonomous agent architectures, this classical mindset completely collapses:

In HEFESTO, we replace unstructured text logs with the formal Agent Trajectory Paradigm: an immutable, causally-linked directed acyclic graph (DAG) that models autonomous execution as a sequence of atomic state transitions.

┌────────────────────────────────────────────────────────────────────────┐
│                        AGENT TRAJECTORY (RUN ID)                       │
│    RunID · HarnessID · AgentID · ModelID · Timestamp · TraceID         │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │ Sequenced Monotonic Turns (0..N)
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                        ATOMIC TURN SPECIFICATION                       │
│   PriorState ──> Observation ──> Proposal ──> Policy ──> ResultState   │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │ OpenTelemetry GenAI Hierarchy
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                        CAUSAL SPAN TOPOLOGY                            │
│   Parent Span (Orchestrator) ──> Child Spans (Subagents, Tools, AST)   │
└────────────────────────────────────────────────────────────────────────┘

The Six Cardinal Coordinates of an Execution Turn

Every turn within an agent trajectory captures: 1. Model & Provider Coordinates: Foundation model identifier, provider checkpoint, sampling parameters. 2. Harness Containment Context: Active runtime policies, token quotas, and capability mask. 3. Filtered Observation: Sanitized environment feedback, tool returns, and compiler outputs. 4. Syntactic Action Proposal: The candidate tool invocation or code modification emitted by the model. 5. Capability Gate Verdict: Deterministic evaluation (ALLOW / DENY) enforcing least privilege. 6. Resulting Ground-Truth State: Environment state delta after physical execution.


2. Decision Lineage: Authority Reconstruction & The Cardinal Equation

Traditional distributed tracing tracks network latency across microservices: $$\text{INPUT} \longrightarrow \text{MODEL} \longrightarrow \text{OUTPUT}$$

Governed multi-agent systems require Decision Lineage: a mathematical audit trail reconstructing the hierarchy of authority that justified every resolution.

The Cardinal Lineage Formulation

$$\text{STATE} + \text{EVIDENCE} + \text{ADMITTED CONTEXT} + \text{RETRIEVED MEMORY} \longrightarrow \text{DECISION} \longrightarrow \text{ARTIFACT / ACTION}$$

Every decision is sealed cryptographically via SHA-256 content digests across all four contributing authority tiers:

type DecisionRecord struct {
    DecisionID            string         `json:"decision_id"`
    RunID                 string         `json:"run_id"`
    TurnIndex             int            `json:"turn_index"`
    StateInputs           []*LineageNode `json:"state_inputs"`
    EvidenceInputs        []*LineageNode `json:"evidence_inputs"`
    ContextInputs         []*LineageNode `json:"context_inputs"`
    MemoryInputs          []*LineageNode `json:"memory_inputs"`
    DecisionOutcome       string         `json:"decision_outcome"`
    DerivedAction         string         `json:"derived_action"`
    DerivedArtifactDigest string         `json:"derived_artifact_digest"`
    LineageDigest         string         `json:"lineage_digest"`
    Timestamp             time.Time      `json:"timestamp"`
}

The Four Authority Tiers

1. Tier 1 — Canonical State (Top Authority): Verified Git commits, active branches, verified environment variables. Cannot be overridden by historical memory. 2. Tier 2 — Empirical Evidence: Compiler outputs, AST parser trees, deterministic oracle test results ($100\%$ PASS receipts). Preempts speculative claims. 3. Tier 3 — Admitted Context: Sanitized user prompts and system instructions bounded by token quotas. 4. Tier 4 — Historical Memory (Advisory): Vector store retrievals and prior session memories. Tentative suggestions strictly subordinate to Tiers 1–3.

$$\text{Strict Precedence:} \quad \text{STATE (T1)} > \text{EVIDENCE (T2)} > \text{CONTEXT (T3)} > \text{MEMORY (T4)}$$


3. Memory & Context Observability: Lifecycle Events & Minimal Provenance

When historical memory informs an execution, the harness observability stack must reconstruct the governed information path by which that information reached the model.

Observable Memory Lifecycle Events

Minimal Memory Provenance Profile

To guarantee verifiable audit trails without requiring proprietary schemas, governed systems record seven minimal provenance attributes: 1. Memory Reference: Cryptographic SHA-256 content hash or immutable storage pointer. 2. Originating Source: Tool return, human interaction, or agent inference. 3. Originating Role/Agent: Creator identity and cryptographic signature. 4. Scope Tags: Multi-tenant partition ID, task ID, and environment tier. 5. Retrieval Timestamp & Age: UTC retrieval wall-clock time and duration relative to maximum TTL. 6. Authority Class: Classification of the record (evidence, state, memory, hypothesis). 7. Governance Outcome & Rule Citation: ADMITTED, DEMOTED, or REJECTED, citing the exact policy rule ID.


4. Auditability of Rejected Candidates & Data Minimization

The Unseen Inputs Principle:
Observability must capture not only what entered the context, but also material candidates rejected for governance reasons when that rejection is operationally relevant.

Recording governance rejections enables post-mortem audits of:

Data Minimization vs Context Bloat

Archiving $100\text{K}$ raw tokens on every loop turn inflates telemetry storage costs and violates GDPR/SOC 2 retention mandates. HEFESTO enforces structured data minimization:


5. Artifact Lineage & Causal Provenance Graphs

In software engineering workflows, a code commit generated by an agent must be backed by a causal provenance graph linking generated files to verification receipts:

┌─────────────────┐       ┌─────────────────┐       ┌─────────────────┐
│   USER INTENT   │ ───>  │   TURN INDEX    │ ───>  │ CAPABILITY GATE │
└─────────────────┘       └────────┬────────┘       └────────┬────────┘
                                   │                         │
                                   ▼                         ▼
                          ┌─────────────────┐       ┌─────────────────┐
                          │ DETERMINISTIC   │ ───>  │ ARTIFACT NODE   │
                          │ TEST ORACLE     │       │ (SHA-256 DIGEST)│
                          └─────────────────┘       └─────────────────┘

Content Addressing (SHA-256)


6. Telemetry, Granular Cost Attribution & OpenTelemetry Tracing

Autonomous multi-agent systems introduce real-time financial and latency risks. HEFESTO models tokens and dollars as first-class system resources bounded by rate cards:

Dynamic Pricing Rate Cards (USD per 1M Tokens)

$$\text{Turn Cost (USD)} = \frac{(\text{PromptTokens} \times \text{Rate}_{\text{in}}) + (\text{CompletionTokens} \times \text{Rate}_{\text{out}})}{1,000,000}$$

OpenTelemetry GenAI Semantic Conventions

Every execution span emits standardized attributes:

Real-Time Budget Enforcement


7. Forensic Post-Mortem & Deterministic Trajectory Replay

When an autonomous system fails in production, re-running the agent with the same prompt fails due to stochastic model drift and mutated environment state.

The Four-Phase RCA Protocol

1. Incident Isolation: Extract immutable RunID and export the captured trajectory JSON. 2. Gate Auditing: Inspect Capability Gate decisions and memory rejection audit logs. 3. Lineage Inspection: Verify SHA-256 node digests across all four authority tiers. 4. Causal Classification: Categorize failure as Model Hallucination, Context Truncation, External Tool Fault, or Policy Misconfiguration.

Deterministic Replay in Sandbox