← Lab L12 — Trajectory Tracing & Decision Lineage All modules Lab L14 — Failure Injection →

Lab L13 — Equal-Compute Benchmark & Real-Time Token Budget Allocation (Pure Go)

Phase 9 · COST & TOKEN ENGINEERING — Lab L13
Status: Authored & Empirically Verified.
Student Lab Package: Authenticated Direct Download from VTAlgo Platform (hefesto-lab13-equal-compute-benchmark.zip)
Branches: main (starter template) · solution (reference architecture).
Canonical Path: Student repo only (hefesto-lab13-equal-compute-benchmark) — not E:\bridle, not the live product.

1. Laboratory Objective

Construct from scratch in pure Go standard library (zero external dependencies in go.mod) an enterprise-grade Equal-Compute Benchmark & Token Budget Engine, solving the fundamental architecture controversy of autonomous systems:

Under an identical financial compute constraint (e.g. exactly $0.0500 per task), does a single frontier model with deep reasoning (SAS) beat an ensemble of specialized lightweight models (MAS)?

The student implements:


2. System Architecture

┌──────────────────────────────────────────────────────────────────────────┐
│      HEFESTO EQUAL-COMPUTE BENCHMARK & BUDGET ENGINE (PURE GO)           │
│                                                                          │
│  ┌─────────────────────────┐           ┌──────────────────────────────┐  │
│  │ 1. Atomic Budget Kernel │           │ 2. Prompt Cache Simulator    │  │
│  │    RateCards (3 Tiers)  ├──────────►│    KV-Cache Block Matching   │  │
│  │    Reserve() / Settle() │           │    Canonical vs Naïve Layout │  │
│  │    Hard Circuit Breaker │           │    33%+ Cost Reduction       │  │
│  └───────────┬─────────────┘           └──────────────┬───────────────┘  │
│              │                                        │                  │
│              ▼                                        ▼                  │
│  ┌─────────────────────────┐           ┌──────────────────────────────┐  │
│  │ 3. SAS Harness (Frontier│           │ 4. MAS Harness (Multi-Tier)  │  │
│  │    Deep Internal CoT    ├──────────►│    Router + Worker + Verifier│  │
│  │    Self-Refine Loop     │           │    Coordination Tax Tracking │  │
│  │    CTax = 0.0%          │           │    CTax = 20% to 45%         │  │
│  └───────────┬─────────────┘           └──────────────┬───────────────┘  │
│              │                                        │                  │
│              ▼                                        ▼                  │
│  ┌────────────────────────────────────────────────────────────────────┐  │
│  │ 5. Equal-Compute Runner & CLI: $0.05 Hard Cap · Pareto Summary     │  │
│  └────────────────────────────────────────────────────────────────────┘  │
└──────────────────────────────────────────────────────────────────────────┘

3. The Five Core Engine Subsystems

1. Atomic Budget Engine & Hard Circuit Breaker (pkg/budget)

2. Prefix Cache Simulator (pkg/caching)

3. Single-Agent System (SAS) Harness (pkg/sas)

4. Multi-Agent System (MAS) Harness (pkg/mas)

5. Benchmark Runner & Reporting Engine (pkg/benchmark, pkg/metrics)


4. Benchmark CLI Execution Modes

Compile and run the unified CLI (cmd/benchmark-cli):

Mode 1: Nominal Equal-Compute Benchmark

go run cmd/benchmark-cli/main.go -mode nominal

Runs the full 3-task suite under budget parity ($0.05 per task), printing the comparative ASCII matrix:

Mode 2: Hard Circuit Breaker Audit

go run cmd/benchmark-cli/main.go -mode budget-breaker

Injects an adversarial oscillating error trap and proves that the Go mutex governor trips at turn 3 ($0.0360 spent of $0.0500 max), preventing financial exhaustion with 0% budget overrun.

Mode 3: Prefix Caching Optimization Audit

go run cmd/benchmark-cli/main.go -mode cache-opt

Executes a live 4-turn sequence under naive vs canonical context ordering, proving an increase in hit ratio from 29.9% to 79.2% and a 33.4% net cost reduction.


5. Certification Protocol

Run the automated test suite with race detection:

go test -v ./...

Verified Test Results

=== RUN   TestSASLinearReasoningNominal
--- PASS: TestSASLinearReasoningNominal (0.00s)
=== RUN   TestMASParallelSearchNominal
--- PASS: TestMASParallelSearchNominal (0.00s)
=== RUN   TestEqualComputeLinearReasoningComparison
--- PASS: TestEqualComputeLinearReasoningComparison (0.00s)
=== RUN   TestEqualComputeParallelSearchComparison
--- PASS: TestEqualComputeParallelSearchComparison (0.00s)
=== RUN   TestAdversarialRunawayBreaker
--- PASS: TestAdversarialRunawayBreaker (0.00s)
=== RUN   TestBenchmarkReporting
--- PASS: TestBenchmarkReporting (0.00s)
=== RUN   TestWinnerDeterminationBothFail
--- PASS: TestWinnerDeterminationBothFail (0.00s)
=== RUN   TestDeterministicTestSuiteIntegrity
--- PASS: TestDeterministicTestSuiteIntegrity (0.00s)
=== RUN   TestParetoEfficiencyComparison
--- PASS: TestParetoEfficiencyComparison (0.00s)
=== RUN   TestBudgetComputeCost
--- PASS: TestBudgetComputeCost (0.00s)
=== RUN   TestBudgetReservationAndSettlement
--- PASS: TestBudgetReservationAndSettlement (0.00s)
=== RUN   TestHardCircuitBreakerOnReservation
--- PASS: TestHardCircuitBreakerOnReservation (0.00s)
=== RUN   TestDirectDeductAndOverrunProtection
--- PASS: TestDirectDeductAndOverrunProtection (0.00s)
=== RUN   TestConcurrentBudgetReservations
--- PASS: TestConcurrentBudgetReservations (0.00s)
=== RUN   TestBudgetCustomRateCards
--- PASS: TestBudgetCustomRateCards (0.00s)
=== RUN   TestBudgetInvalidReservationID
--- PASS: TestBudgetInvalidReservationID (0.00s)
=== RUN   TestManualTripBreaker
--- PASS: TestManualTripBreaker (0.00s)
=== RUN   TestZeroBudgetImmediateTrip
--- PASS: TestZeroBudgetImmediateTrip (0.00s)
=== RUN   TestPrefixCacheSimulation
--- PASS: TestPrefixCacheSimulation (0.00s)
=== RUN   TestNaiveVsCanonicalCacheInvalidation
--- PASS: TestNaiveVsCanonicalCacheInvalidation (0.00s)
PASS
ok  	hefesto-lab13-equal-compute-benchmark/tests	0.243s

Passing Criteria: 20/20 tests PASS in under 250 milliseconds with zero race conditions.