Lab L13 — Equal-Compute Benchmark & Real-Time Token Budget Allocation (Pure Go)
Phase 9 · COST & TOKEN ENGINEERING — Lab L13
Status: Authored & Empirically Verified.
Student Lab Package: Authenticated Direct Download from VTAlgo Platform (hefesto-lab13-equal-compute-benchmark.zip)
Branches:main(starter template) ·solution(reference architecture).
Canonical Path: Student repo only (hefesto-lab13-equal-compute-benchmark) — notE:\bridle, not the live product.
1. Laboratory Objective
Construct from scratch in pure Go standard library (zero external dependencies in go.mod) an enterprise-grade Equal-Compute Benchmark & Token Budget Engine, solving the fundamental architecture controversy of autonomous systems:
Under an identical financial compute constraint (e.g. exactly $0.0500 per task), does a single frontier model with deep reasoning (SAS) beat an ensemble of specialized lightweight models (MAS)?
The student implements:
- An atomic, thread-safe budget tracking kernel (
pkg/budget) with micro-dollar precision, parameterized rate cards for three compute tiers (Frontier, Standard, Fast), two-phase reservation (Reserve/Settle), and hard circuit breakers (ErrBudgetExceeded) guaranteeing zero financial overdraft. - A deterministic prefix cache simulator (
pkg/caching) evaluating KV-cache hit ratios and proving how canonical prompt ordering (static system rules and sorted tools first, dynamic timestamps at the tail) reduces input costs by more than 33%. - A Single-Agent System (SAS) harness (
pkg/sas) executing multi-step reasoning and self-refinement under a Frontier model with zero coordination tax. - A Multi-Agent System (MAS) harness (
pkg/mas) coordinating a Router, specialized Workers, and a Verifier, explicitly tracking and quantifying the Multi-Agent Coordination Tax. - An Equal-Compute Benchmark runner and workload suite (
pkg/benchmark) testing both architectures on tightly coupled linear proofs, decoupled parallel search, and adversarial runaway loops. - An aggregate reporting engine (
pkg/metrics) generating Pareto efficiency frontiers, Cost-Per-Success metrics, and formatted ASCII/JSON tables. - A multi-mode CLI utility (
cmd/benchmark-cli) supportingnominal,budget-breaker, andcache-optexecution modes. - A 20/20 deterministic test suite passing in $<250$ ms with zero race conditions.
2. System Architecture
┌──────────────────────────────────────────────────────────────────────────┐
│ HEFESTO EQUAL-COMPUTE BENCHMARK & BUDGET ENGINE (PURE GO) │
│ │
│ ┌─────────────────────────┐ ┌──────────────────────────────┐ │
│ │ 1. Atomic Budget Kernel │ │ 2. Prompt Cache Simulator │ │
│ │ RateCards (3 Tiers) ├──────────►│ KV-Cache Block Matching │ │
│ │ Reserve() / Settle() │ │ Canonical vs Naïve Layout │ │
│ │ Hard Circuit Breaker │ │ 33%+ Cost Reduction │ │
│ └───────────┬─────────────┘ └──────────────┬───────────────┘ │
│ │ │ │
│ ▼ ▼ │
│ ┌─────────────────────────┐ ┌──────────────────────────────┐ │
│ │ 3. SAS Harness (Frontier│ │ 4. MAS Harness (Multi-Tier) │ │
│ │ Deep Internal CoT ├──────────►│ Router + Worker + Verifier│ │
│ │ Self-Refine Loop │ │ Coordination Tax Tracking │ │
│ │ CTax = 0.0% │ │ CTax = 20% to 45% │ │
│ └───────────┬─────────────┘ └──────────────┬───────────────┘ │
│ │ │ │
│ ▼ ▼ │
│ ┌────────────────────────────────────────────────────────────────────┐ │
│ │ 5. Equal-Compute Runner & CLI: $0.05 Hard Cap · Pareto Summary │ │
│ └────────────────────────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────────────────────────┘
3. The Five Core Engine Subsystems
1. Atomic Budget Engine & Hard Circuit Breaker (pkg/budget)
RateCardSpecification: Parametrizes input, output, and cached read costs per million tokens forTierFrontier($3.00 / $15.00 / $0.75),TierStandard($0.50 / $1.50 / $0.10), andTierFast($0.15 / $0.60 / $0.03).- Two-Phase Reservation Protocol:
Reserve(tier, worstIn, worstOut)locks worst-case funds under async.Mutexand returns a reservation ID.Settle(id, actualIn, actualOut, actualCached)charges actual usage and immediately releases the unused hold. - Hard Circuit Breaker: When
spentUSD + worstCost > maxBudgetUSD, the breaker trips instantly (isTripped = true), recording the exact trip timestamp and reason, and returningErrBudgetExceeded.
2. Prefix Cache Simulator (pkg/caching)
- Block Segmentation: Chunks serialized prompt strings into discrete blocks, computes SHA-256 digests, and matches consecutive prefix blocks turn-over-turn.
- Layout Order Verification:
AssembleNaive: Injects timestamps at the top of the context, triggering a cache miss on block 1 and destroying downstream hit ratios (29.9% hit ratio).AssembleCanonical: Keeps system instructions, role contracts, and sorted tools static at the top; pushes timestamps and queries strictly to the tail (79.2% hit ratio, 33.4% net cost savings).
3. Single-Agent System (SAS) Harness (pkg/sas)
- Concentrates the entire task budget into invoking a
TierFrontierreasoning model. - Executes iterative self-refinement passes.
- Computes with zero coordination tax (
CTax = 0.0%), achieving 96% accuracy on formal cryptographic proofs.
4. Multi-Agent System (MAS) Harness (pkg/mas)
- Subdivides the problem across three specialized agents:
Router(TierFast),Worker(TierStandard), andVerifier(TierFast). - Explicitly measures worker execution tokens vs inter-agent message passing: $$\text{CoordinationTax} = \frac{\text{RouterTokens} + \text{VerifierTokens}}{\text{TotalTokens}}$$
- Shines on decoupled parallel extraction (98% accuracy at $0.0022 spend), but suffers significant degradation on coupled formal proofs (42% accuracy with 43.7% CTax).
5. Benchmark Runner & Reporting Engine (pkg/benchmark, pkg/metrics)
- Executes both harnesses against a standardized test battery under identical budget constraints ($0.0500 per task).
- Evaluates the Cost-Per-Success (CPS) metric: $$\text{CPS} = \frac{\text{Spend}}{\text{Accuracy}}$$
- Outputs structured JSON receipts and formatted ASCII summary tables.
4. Benchmark CLI Execution Modes
Compile and run the unified CLI (cmd/benchmark-cli):
Mode 1: Nominal Equal-Compute Benchmark
go run cmd/benchmark-cli/main.go -mode nominal
Runs the full 3-task suite under budget parity ($0.05 per task), printing the comparative ASCII matrix:
- SAS dominates on linear formal proofs (96% accuracy vs 42% for MAS).
- MAS dominates on parallel search shards (98% accuracy vs 84% for SAS, at 13x lower cost).
- Both trip the circuit breaker on adversarial runaway loops.
Mode 2: Hard Circuit Breaker Audit
go run cmd/benchmark-cli/main.go -mode budget-breaker
Injects an adversarial oscillating error trap and proves that the Go mutex governor trips at turn 3 ($0.0360 spent of $0.0500 max), preventing financial exhaustion with 0% budget overrun.
Mode 3: Prefix Caching Optimization Audit
go run cmd/benchmark-cli/main.go -mode cache-opt
Executes a live 4-turn sequence under naive vs canonical context ordering, proving an increase in hit ratio from 29.9% to 79.2% and a 33.4% net cost reduction.
5. Certification Protocol
Run the automated test suite with race detection:
go test -v ./...
Verified Test Results
=== RUN TestSASLinearReasoningNominal
--- PASS: TestSASLinearReasoningNominal (0.00s)
=== RUN TestMASParallelSearchNominal
--- PASS: TestMASParallelSearchNominal (0.00s)
=== RUN TestEqualComputeLinearReasoningComparison
--- PASS: TestEqualComputeLinearReasoningComparison (0.00s)
=== RUN TestEqualComputeParallelSearchComparison
--- PASS: TestEqualComputeParallelSearchComparison (0.00s)
=== RUN TestAdversarialRunawayBreaker
--- PASS: TestAdversarialRunawayBreaker (0.00s)
=== RUN TestBenchmarkReporting
--- PASS: TestBenchmarkReporting (0.00s)
=== RUN TestWinnerDeterminationBothFail
--- PASS: TestWinnerDeterminationBothFail (0.00s)
=== RUN TestDeterministicTestSuiteIntegrity
--- PASS: TestDeterministicTestSuiteIntegrity (0.00s)
=== RUN TestParetoEfficiencyComparison
--- PASS: TestParetoEfficiencyComparison (0.00s)
=== RUN TestBudgetComputeCost
--- PASS: TestBudgetComputeCost (0.00s)
=== RUN TestBudgetReservationAndSettlement
--- PASS: TestBudgetReservationAndSettlement (0.00s)
=== RUN TestHardCircuitBreakerOnReservation
--- PASS: TestHardCircuitBreakerOnReservation (0.00s)
=== RUN TestDirectDeductAndOverrunProtection
--- PASS: TestDirectDeductAndOverrunProtection (0.00s)
=== RUN TestConcurrentBudgetReservations
--- PASS: TestConcurrentBudgetReservations (0.00s)
=== RUN TestBudgetCustomRateCards
--- PASS: TestBudgetCustomRateCards (0.00s)
=== RUN TestBudgetInvalidReservationID
--- PASS: TestBudgetInvalidReservationID (0.00s)
=== RUN TestManualTripBreaker
--- PASS: TestManualTripBreaker (0.00s)
=== RUN TestZeroBudgetImmediateTrip
--- PASS: TestZeroBudgetImmediateTrip (0.00s)
=== RUN TestPrefixCacheSimulation
--- PASS: TestPrefixCacheSimulation (0.00s)
=== RUN TestNaiveVsCanonicalCacheInvalidation
--- PASS: TestNaiveVsCanonicalCacheInvalidation (0.00s)
PASS
ok hefesto-lab13-equal-compute-benchmark/tests 0.243s
Passing Criteria: 20/20 tests PASS in under 250 milliseconds with zero race conditions.