case-studies / turbo-rig
Case Study: One Week Inside a Governed Agentic Coding Rig
The claim behind everything on this site is that AI agents can do the building while humans keep the authority. This study is that claim, measured.
One audited production week — September 21–28, 2026 — across my own GitHub account, with every number drawn from the append-only gate and merge logs.
Problem
Turbo Flow — the predecessor — had proven that one developer could orchestrate serious agentic horsepower: swarms of agents, thousands of lines of orchestration, impressive output. It had also produced the opposite of reliability: more agents meant more surface area, more failure modes, and more time spent herding the machinery instead of shipping.
Observation
More agents does not mean better engineering.— the observation that redesigned the workflow
The bottleneck was never generation capacity — models generate plenty. The bottleneck was trust: nobody could prove, for any given change, that an independent reviewer had challenged it and a human had decided on it.
Redesign
| Before (Turbo Flow v4-era) | After (Turbo Rig) |
|---|---|
| 3,181 lines of orchestration | ~160 lines, five scripts |
| Many coordinating agents | One cross-model review gate |
| Trust by convention | Parseable verdicts, fail-closed |
| Memory in tooling state | Git-versioned agent memory |
| Automation with broad access | A written constitution, enforced |
| Merges by momentum | Human-only merges |
Operating principle
The builder never approves its own work.
Everything else is implementation. The reviewer is always a different model family, the verdict always machine-parsed, a missing verdict always a rejection, and the merge button always human.
Result — the audited week
Figures below come from one audited production week (September 21–28, 2026) across my own GitHub account — merged-PR counts, gate logs, and token accounting, published weekly on the Turbo Rig stats page. They are measurements of my own workflow, not industry averages.
| Metric | Week of Sept 21–28, 2026 |
|---|---|
| Merged PRs (account-wide) | 180 across 16 repositories |
| Cross-model gate reviews | 967 |
| Tokens consumed | 6.57 billion |
| Prompt-cache hit rate | 96.5% |
| Commits (account-wide) | 595 |
| PRs merged by an agent | 0 — by construction |
The ratio is the story: 967 gate reviews for 180 merges means every change that landed was challenged — repeatedly — before a human decided. Some PRs cleared in one round; the hard ones took several, and the rejections are logged with reasons. That log is the difference between "we use AI carefully" and proof.
The cost line matters too: a 96.5% prompt-cache hit rate is what keeps a multi-billion-token week affordable. Agent cost is an engineering variable, not a fact of nature — you manage it with caching, scoping, and ceilings. See agent observability.
What this means for your team
The rig is one person's workflow, and the numbers are one person's — they are evidence of the mechanism, not a promise of your throughput. The transferable part is the architecture: one gate, one constitution, human merges, and an append-only log. Installing that shape in a team's repositories is the team onboarding engagement.
Talk it through with someone who runs this stack on his own systems every day.