Module XI — Security & Governance
Phase 8 · SECURITY & GOVERNANCE — Module XI
Status: Authored & Empirically Verified.
Lecture Components: 6 FHD 1080p master videos + Lab L11.
Canonical Core Axioms:
CAPABILITY GATE: NO AGENT RECEIVES A CAPABILITY ONLY BECAUSE ITS MODEL CAN USE IT.
PERSISTED ≠ TRUSTED · RETRIEVED ≠ SAFE
THE OS KERNEL IS THE ONLY AUTHORITATIVE BOUNDARY; HARNESS INVARIANTS MUST REMAIN SOVEREIGN OVER MODEL OBEDIENCE.
1. Least Privilege and the Capability Gate Doctrine
In naive multi-agent frameworks, developers routinely grant broad system access to agent runtimes, trusting that prompt instructions (such as *"you are a safe assistant; do not touch production databases"*) will constrain behavior. This is an egregious architectural error: foundation models are stochastic text generators, not security enforcement boundaries.
The Capability Gate doctrine mandates that access rights are physical, deterministic constraints enforced at the harness layer before any model-generated action reaches an execution dispatcher.
┌────────────────────────────────────────────────────────────────────────┐
│ AGENT CONTROL HARNESS │
│ Prompt Ingestion · Context Window · Stochastic Tool Selection │
└───────────────────────────────────┬────────────────────────────────────┘
│ Emits Unverified Action Candidate
▼
┌────────────────────────────────────────────────────────────────────────┐
│ THE CAPABILITY GATE (GO KERNEL) │
│ Token Verification (TTL) · O(1) Bitmask Gating · Escalation Traps │
└───────────────────────────────────┬────────────────────────────────────┘
│ Authorized Actions Only
▼
┌────────────────────────────────────────────────────────────────────────┐
│ PHYSICAL EXECUTION SANDBOX │
│ Ephemeral Worktrees · SafePath Containment · Zero-Trust Egress │
└────────────────────────────────────────────────────────────────────────┘
O(1) Atomic Bitmask Evaluation
Privileges are encoded as discrete bits within an atomic uint64 capability bitmask:
const (
CapReadDisk uint64 = 1 << 0
CapWriteIsolated uint64 = 1 << 1
CapCompileCode uint64 = 1 << 2
CapExecuteBinary uint64 = 1 << 3
CapNetworkEgress uint64 = 1 << 4
CapInspectAST uint64 = 1 << 5
CapReadMemory uint64 = 1 << 6
CapWriteMemory uint64 = 1 << 7
CapAdmin uint64 = 1 << 8
)
Authorization evaluates in constant $O(1)$ CPU register time with zero heap allocations: $$\text{Authorized} \iff (\text{Token.Bitmask} \ \& \ \text{Required}) == \text{Required}$$
Non-Transitive Delegation Ceiling
When an orchestrator delegates tasks to subagents, the child's capability mask cannot exceed the parent's assigned scope. Any attempt to elevate privileges trips an immediate ErrPrivilegeEscalation trap: $$(\text{ChildMask} \ \& \ \sim\!\text{ParentMask}) \neq 0 \implies \text{ABORT}$$
2. Sandboxing, Process Isolation, and Network Egress Policies
Interpreted language wrappers (e.g. Python monkey-patching or string inspection filters) provide zero meaningful containment against malicious payloads or prompt injections. The operating system kernel is the only authoritative security boundary.
Physical Isolation Primitives
1. Go Process Silos: Independent sub-processes executed under dedicated, unprivileged operating system accounts. 2. Cgroups V2: Hard ceilings on RAM consumption (e.g. 512 MB per worker) and CPU quotas (e.g. 50% of 1 core) to neutralize Economic Denial of Service (EDoS) attacks and compiler fork-bombs. 3. Ephemeral Ramdisks (tmpfs): Mutations occur in volatile memory mounted strictly for the task duration, unmounted and scrubbed upon completion with zero residual storage footprint. 4. SafePath Path Containment: Strict canonical path resolution prohibiting ../ traversal or symbolic link escapes targeting host filesystems: go cleanRel := filepath.Clean(relPath) targetPath := filepath.Join(s.rootDir, cleanRel) rel, err := filepath.Rel(s.rootDir, targetPath) if err != nil || strings.HasPrefix(rel, "..") { return "", ErrPathEscape }
Zero-Trust Network Egress Interceptor
Outbound network access is disabled by default. When egress is explicitly granted (CapNetworkEgress):
- Direct IP literals (e.g.
10.0.0.1,169.254.169.254) are blocked unconditionally to defeat SSRF attacks. - Only explicit, corporate-whitelisted FQDNs (e.g.
pkg.go.dev) over TLS port 443 are routed.
3. Memory as a Trust Boundary and Memory Poisoning
Persistent memory (vector databases, semantic stores, key-value caches) is an independent attack and failure surface. In HEFESTO, persistence alone does not establish credibility:
PERSISTED ≠ TRUSTED
RETRIEVED ≠ SAFE
Memory Threat Taxonomy
- Stale Benign Memory: Outdated configurations or revoked credentials inducing operational failures.
- Untrusted Provenance: Records lacking cryptographic author signatures or timestamp proofs.
- Cross-Tenant Leakage: Semantic vector searches surfacing sensitive data across client boundaries (
ErrTenantMismatch). - Agent-to-Agent Contamination: Unverified speculative claims from one agent persisted and ingested by subsequent agents as canonical truth.
Memory Poisoning: Accidental vs Adversarial
1. Accidental Contamination: LLM hallucinations or flawed intermediate deductions persisted without oracle verification. 2. Adversarial Poisoning: Deliberate trojan payloads ingested via external repositories, documentation, or tool outputs designed to hijack future agent planning.
Pre-Admission Firewalls & Non-Contradiction Oracles
Every persistent candidate must pass pre-admission filtering (pkg/memory):
func (e *MemoryEnclave) Admit(record *MemoryRecord) error {
if !record.Verified { return ErrUnverifiedMemory }
if activeVal, exists := e.activeInvariants[record.InvariantKey]; exists {
if record.InvariantValue != activeVal {
return fmt.Errorf("%w: key=%s active=%s candidate=%s", ErrMemoryContradiction, ...)
}
}
e.records[record.ID] = record
return nil
}
4. Prompt Injection, Tool Poisoning, and Agent-to-Agent Cascades
While direct prompt injections from end-users are relatively easy to sanitize, indirect prompt injection is the primary threat in autonomous software workflows. It occurs when an agent ingests untrusted environmental data: source code comments, README files, bug reports, or tool outputs containing adversarial override directives.
The Von Neumann Conflation in Transformers
Because LLMs concatenate instructions and data into a single token stream, relying on prompt instructions to ignore injected text is fundamentally flawed. HEFESTO replaces stochastic LLM judges with Deterministic Defense-in-Depth Oracles:
┌────────────────────────────────────────────────────────────────────────┐
│ UNTRUSTED EXTERNAL DATA │
│ GitHub Issues · Scraped Web Docs · Command Stdout / Stderr │
└───────────────────────────────────┬────────────────────────────────────┘
│ Raw Byte Stream
▼
┌────────────────────────────────────────────────────────────────────────┐
│ LAYER 1: PAYLOAD SANITIZER │
│ ANSI Control Code Stripping · Override Regex Pattern Filtering │
└───────────────────────────────────┬────────────────────────────────────┘
│ Sanitized Structured Envelope
▼
┌────────────────────────────────────────────────────────────────────────┐
│ LAYER 2: CRYPTOGRAPHIC CANARY │
│ HMAC-SHA256 Token Ingestion · Instant Prompt Leakage Detection │
└───────────────────────────────────┬────────────────────────────────────┘
│ Clean AST Generation
▼
┌────────────────────────────────────────────────────────────────────────┐
│ LAYER 3: GO AST IMPORT ORACLE │
│ go/parser Static Validation · Unauthorized Package Blocking │
└────────────────────────────────────────────────────────────────────────┘
1. HMAC Canary Engine: Secret cryptographic canary tokens seeded inside system prompts; any response or tool echo containing the canary triggers an immediate ErrCanaryLeaked alarm. 2. Observation Sanitizer: Strips terminal ANSI control escape sequences and blocks override patterns (ignore previous instructions, developer mode enabled). 3. Go AST Import Oracle: Inspects agent-generated code using standard library go/parser, verifying all imports against an approved package whitelist before compilation.
5. Autonomous Git, Supply Chain Security, and Secret Brokerage
Allowing an autonomous agent to execute direct git commit or git push operations against central branches represents an existential compliance risk.
Ephemeral Git Worktrees
The central upstream repository is mounted strictly read-only. All agent mutations occur in isolated Git Worktrees (git worktree add -b sandbox/task-101 /tmp/wt-101 HEAD). If deterministic verification tests fail, the worktree is scrubbed with zero residue on upstream branches.
Neutralizing Generative Package Typosquatting
Foundation models frequently hallucinate plausible third-party package names. Adversaries monitor LLM outputs and register these squatting packages on npm or PyPI with embedded malware.
- HEFESTO mandates immutable lockfiles (
go.sum,go.mod). - The Go compiler is invoked with
-mod=readonlyand network access disabled. - The AST Import Oracle drops any patch referencing undeclared dependencies.
Credential Brokerage Pattern
Agents must never view raw production secrets. The runtime implements a local mediation proxy:
- Agent targets
http://127.0.0.1:9090with ephemeral, single-use capability tokens. - The Security Kernel validates
CapNetworkEgressand injects productionAuthorization: Bearer <SECRET>headers directly into the outbound TLS socket. - The agent reasons and executes tasks without ever holding secret keys in memory.
6. Human Approval Gates and Cryptographic Governance Receipts
Human-in-the-loop (HITL) is an audit and escalation control, never a replacement for least privilege or sandboxing. Designing naive confirmation dialogs leads to Approval Fatigue, causing human operators to rubber-stamp complex actions without adequate review.
Deterministic Escalation Triggers
Human intervention is invoked exclusively upon deterministic, high-impact thresholds: 1. Irreversible Mutations: Production schema migrations, direct cloud deployments, or CI/CD workflow alterations. 2. Budget Ceilings: Token consumption or compute expenditures exceeding assigned task allowances. 3. Contradictory Oracles: Independent verification oracles yielding conflicting verdicts (e.g. unit tests PASS, but AST security linters FAIL).
Cryptographic Governance Receipts
Every security-relevant decision emits an immutable, non-repudiable receipt sealed via HMAC-SHA256 in pure Go:
type Receipt struct {
ReceiptID string `json:"receipt_id"`
TaskID string `json:"task_id"`
Role string `json:"role"`
Action string `json:"action"`
CapabilityMask uint64 `json:"capability_mask"`
Verdict string `json:"verdict"` // APPROVED | BLOCKED
Reason string `json:"reason"`
IssuedAt time.Time `json:"issued_at"`
HMACSignature string `json:"hmac_signature"`
}
Receipt signatures are verified in constant time via hmac.Equal, providing mathematical proof for SOC 2 Type II, HIPAA, and ISO 27001 audits that no agent action bypassed corporate access control policies.