adventureonthewave

case-studies / turbo-rig

Case Study: One Week Inside a Governed Agentic Coding Rig

The claim behind everything on this site is that AI agents can do the building while humans keep the authority. This study is that claim, measured.

One audited production week — September 21–28, 2026 — across my own GitHub account, with every number drawn from the append-only gate and merge logs.

Problem

Turbo Flow — the predecessor — had proven that one developer could orchestrate serious agentic horsepower: swarms of agents, thousands of lines of orchestration, impressive output. It had also produced the opposite of reliability: more agents meant more surface area, more failure modes, and more time spent herding the machinery instead of shipping.

Observation

More agents does not mean better engineering.— the observation that redesigned the workflow

The bottleneck was never generation capacity — models generate plenty. The bottleneck was trust: nobody could prove, for any given change, that an independent reviewer had challenged it and a human had decided on it.

Redesign

Before (Turbo Flow v4-era)After (Turbo Rig)
3,181 lines of orchestration~160 lines, five scripts
Many coordinating agentsOne cross-model review gate
Trust by conventionParseable verdicts, fail-closed
Memory in tooling stateGit-versioned agent memory
Automation with broad accessA written constitution, enforced
Merges by momentumHuman-only merges

Operating principle

The builder never approves its own work.

Everything else is implementation. The reviewer is always a different model family, the verdict always machine-parsed, a missing verdict always a rejection, and the merge button always human.

Result — the audited week

Figures below come from one audited production week (September 21–28, 2026) across my own GitHub account — merged-PR counts, gate logs, and token accounting, published weekly on the Turbo Rig stats page. They are measurements of my own workflow, not industry averages.

MetricWeek of Sept 21–28, 2026
Merged PRs (account-wide)180 across 16 repositories
Cross-model gate reviews967
Tokens consumed6.57 billion
Prompt-cache hit rate96.5%
Commits (account-wide)595
PRs merged by an agent0 — by construction

The ratio is the story: 967 gate reviews for 180 merges means every change that landed was challenged — repeatedly — before a human decided. Some PRs cleared in one round; the hard ones took several, and the rejections are logged with reasons. That log is the difference between "we use AI carefully" and proof.

The cost line matters too: a 96.5% prompt-cache hit rate is what keeps a multi-billion-token week affordable. Agent cost is an engineering variable, not a fact of nature — you manage it with caching, scoping, and ceilings. See agent observability.

What this means for your team

The rig is one person's workflow, and the numbers are one person's — they are evidence of the mechanism, not a promise of your throughput. The transferable part is the architecture: one gate, one constitution, human merges, and an append-only log. Installing that shape in a team's repositories is the team onboarding engagement.

Talk it through with someone who runs this stack on his own systems every day.