← Module X — State & Recovery All modules Module XII — Observability →

Module XI — Security & Governance

Phase 8 · SECURITY & GOVERNANCE — Module XI
Status: Authored & Empirically Verified.
Lecture Components: 6 FHD 1080p master videos + Lab L11.
Canonical Core Axioms:

CAPABILITY GATE: NO AGENT RECEIVES A CAPABILITY ONLY BECAUSE ITS MODEL CAN USE IT.
PERSISTED ≠ TRUSTED · RETRIEVED ≠ SAFE
THE OS KERNEL IS THE ONLY AUTHORITATIVE BOUNDARY; HARNESS INVARIANTS MUST REMAIN SOVEREIGN OVER MODEL OBEDIENCE.

1. Least Privilege and the Capability Gate Doctrine

In naive multi-agent frameworks, developers routinely grant broad system access to agent runtimes, trusting that prompt instructions (such as *"you are a safe assistant; do not touch production databases"*) will constrain behavior. This is an egregious architectural error: foundation models are stochastic text generators, not security enforcement boundaries.

The Capability Gate doctrine mandates that access rights are physical, deterministic constraints enforced at the harness layer before any model-generated action reaches an execution dispatcher.

┌────────────────────────────────────────────────────────────────────────┐
│                        AGENT CONTROL HARNESS                           │
│     Prompt Ingestion · Context Window · Stochastic Tool Selection      │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │ Emits Unverified Action Candidate
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                      THE CAPABILITY GATE (GO KERNEL)                   │
│   Token Verification (TTL) · O(1) Bitmask Gating · Escalation Traps    │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │ Authorized Actions Only
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                      PHYSICAL EXECUTION SANDBOX                        │
│    Ephemeral Worktrees · SafePath Containment · Zero-Trust Egress      │
└────────────────────────────────────────────────────────────────────────┘

O(1) Atomic Bitmask Evaluation

Privileges are encoded as discrete bits within an atomic uint64 capability bitmask:

const (
    CapReadDisk        uint64 = 1 << 0
    CapWriteIsolated   uint64 = 1 << 1
    CapCompileCode     uint64 = 1 << 2
    CapExecuteBinary   uint64 = 1 << 3
    CapNetworkEgress   uint64 = 1 << 4
    CapInspectAST      uint64 = 1 << 5
    CapReadMemory      uint64 = 1 << 6
    CapWriteMemory     uint64 = 1 << 7
    CapAdmin           uint64 = 1 << 8
)

Authorization evaluates in constant $O(1)$ CPU register time with zero heap allocations: $$\text{Authorized} \iff (\text{Token.Bitmask} \ \& \ \text{Required}) == \text{Required}$$

Non-Transitive Delegation Ceiling

When an orchestrator delegates tasks to subagents, the child's capability mask cannot exceed the parent's assigned scope. Any attempt to elevate privileges trips an immediate ErrPrivilegeEscalation trap: $$(\text{ChildMask} \ \& \ \sim\!\text{ParentMask}) \neq 0 \implies \text{ABORT}$$


2. Sandboxing, Process Isolation, and Network Egress Policies

Interpreted language wrappers (e.g. Python monkey-patching or string inspection filters) provide zero meaningful containment against malicious payloads or prompt injections. The operating system kernel is the only authoritative security boundary.

Physical Isolation Primitives

1. Go Process Silos: Independent sub-processes executed under dedicated, unprivileged operating system accounts. 2. Cgroups V2: Hard ceilings on RAM consumption (e.g. 512 MB per worker) and CPU quotas (e.g. 50% of 1 core) to neutralize Economic Denial of Service (EDoS) attacks and compiler fork-bombs. 3. Ephemeral Ramdisks (tmpfs): Mutations occur in volatile memory mounted strictly for the task duration, unmounted and scrubbed upon completion with zero residual storage footprint. 4. SafePath Path Containment: Strict canonical path resolution prohibiting ../ traversal or symbolic link escapes targeting host filesystems: go cleanRel := filepath.Clean(relPath) targetPath := filepath.Join(s.rootDir, cleanRel) rel, err := filepath.Rel(s.rootDir, targetPath) if err != nil || strings.HasPrefix(rel, "..") { return "", ErrPathEscape }

Zero-Trust Network Egress Interceptor

Outbound network access is disabled by default. When egress is explicitly granted (CapNetworkEgress):


3. Memory as a Trust Boundary and Memory Poisoning

Persistent memory (vector databases, semantic stores, key-value caches) is an independent attack and failure surface. In HEFESTO, persistence alone does not establish credibility:

PERSISTED ≠ TRUSTED
RETRIEVED ≠ SAFE

Memory Threat Taxonomy

Memory Poisoning: Accidental vs Adversarial

1. Accidental Contamination: LLM hallucinations or flawed intermediate deductions persisted without oracle verification. 2. Adversarial Poisoning: Deliberate trojan payloads ingested via external repositories, documentation, or tool outputs designed to hijack future agent planning.

Pre-Admission Firewalls & Non-Contradiction Oracles

Every persistent candidate must pass pre-admission filtering (pkg/memory):

func (e *MemoryEnclave) Admit(record *MemoryRecord) error {
    if !record.Verified { return ErrUnverifiedMemory }
    if activeVal, exists := e.activeInvariants[record.InvariantKey]; exists {
        if record.InvariantValue != activeVal {
            return fmt.Errorf("%w: key=%s active=%s candidate=%s", ErrMemoryContradiction, ...)
        }
    }
    e.records[record.ID] = record
    return nil
}

4. Prompt Injection, Tool Poisoning, and Agent-to-Agent Cascades

While direct prompt injections from end-users are relatively easy to sanitize, indirect prompt injection is the primary threat in autonomous software workflows. It occurs when an agent ingests untrusted environmental data: source code comments, README files, bug reports, or tool outputs containing adversarial override directives.

The Von Neumann Conflation in Transformers

Because LLMs concatenate instructions and data into a single token stream, relying on prompt instructions to ignore injected text is fundamentally flawed. HEFESTO replaces stochastic LLM judges with Deterministic Defense-in-Depth Oracles:

┌────────────────────────────────────────────────────────────────────────┐
│                        UNTRUSTED EXTERNAL DATA                         │
│       GitHub Issues · Scraped Web Docs · Command Stdout / Stderr       │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │ Raw Byte Stream
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                     LAYER 1: PAYLOAD SANITIZER                         │
│     ANSI Control Code Stripping · Override Regex Pattern Filtering     │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │ Sanitized Structured Envelope
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                     LAYER 2: CRYPTOGRAPHIC CANARY                      │
│     HMAC-SHA256 Token Ingestion · Instant Prompt Leakage Detection     │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │ Clean AST Generation
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                     LAYER 3: GO AST IMPORT ORACLE                      │
│      go/parser Static Validation · Unauthorized Package Blocking       │
└────────────────────────────────────────────────────────────────────────┘

1. HMAC Canary Engine: Secret cryptographic canary tokens seeded inside system prompts; any response or tool echo containing the canary triggers an immediate ErrCanaryLeaked alarm. 2. Observation Sanitizer: Strips terminal ANSI control escape sequences and blocks override patterns (ignore previous instructions, developer mode enabled). 3. Go AST Import Oracle: Inspects agent-generated code using standard library go/parser, verifying all imports against an approved package whitelist before compilation.


5. Autonomous Git, Supply Chain Security, and Secret Brokerage

Allowing an autonomous agent to execute direct git commit or git push operations against central branches represents an existential compliance risk.

Ephemeral Git Worktrees

The central upstream repository is mounted strictly read-only. All agent mutations occur in isolated Git Worktrees (git worktree add -b sandbox/task-101 /tmp/wt-101 HEAD). If deterministic verification tests fail, the worktree is scrubbed with zero residue on upstream branches.

Neutralizing Generative Package Typosquatting

Foundation models frequently hallucinate plausible third-party package names. Adversaries monitor LLM outputs and register these squatting packages on npm or PyPI with embedded malware.

Credential Brokerage Pattern

Agents must never view raw production secrets. The runtime implements a local mediation proxy:


6. Human Approval Gates and Cryptographic Governance Receipts

Human-in-the-loop (HITL) is an audit and escalation control, never a replacement for least privilege or sandboxing. Designing naive confirmation dialogs leads to Approval Fatigue, causing human operators to rubber-stamp complex actions without adequate review.

Deterministic Escalation Triggers

Human intervention is invoked exclusively upon deterministic, high-impact thresholds: 1. Irreversible Mutations: Production schema migrations, direct cloud deployments, or CI/CD workflow alterations. 2. Budget Ceilings: Token consumption or compute expenditures exceeding assigned task allowances. 3. Contradictory Oracles: Independent verification oracles yielding conflicting verdicts (e.g. unit tests PASS, but AST security linters FAIL).

Cryptographic Governance Receipts

Every security-relevant decision emits an immutable, non-repudiable receipt sealed via HMAC-SHA256 in pure Go:

type Receipt struct {
    ReceiptID      string    `json:"receipt_id"`
    TaskID         string    `json:"task_id"`
    Role           string    `json:"role"`
    Action         string    `json:"action"`
    CapabilityMask uint64    `json:"capability_mask"`
    Verdict        string    `json:"verdict"` // APPROVED | BLOCKED
    Reason         string    `json:"reason"`
    IssuedAt       time.Time `json:"issued_at"`
    HMACSignature  string    `json:"hmac_signature"`
}

Receipt signatures are verified in constant time via hmac.Equal, providing mathematical proof for SOC 2 Type II, HIPAA, and ISO 27001 audits that no agent action bypassed corporate access control policies.