research
Research
Original experiments and benchmarks, run on working systems, with methodology and raw numbers published.
The test for everything here: could someone cite it? If not, it does not get written.
What belongs here
- Coding-agent benchmarks — the major coding agents measured against standardized engineering tasks: completion rate, review rejection, tokens, cost, time, human intervention.
- Workflow experiments — what happens when the builder cannot approve its own work; what a week at multi-billion-token scale actually looks like; where multi-agent swarms make things worse.
- Incident write-ups — production failures in the ecosystem's own systems, dissected with the logs.
The bar is the one this site holds itself to elsewhere: every metric traceable to its writer, methodology published alongside results, and no industry-percent stat that cannot name its source.
First data already public
The research program effectively started with the rig itself: the Turbo Rig case study publishes one full audited week of a governed agentic workflow — 180 merged PRs, 967 gate reviews, 6.57B tokens, 96.5% cache — with the logging architecture that produced the numbers. Future studies follow the same pattern: mechanism first, measurement second, narrative last.
Cadence
Approximately two genuinely original pieces per month, rather than fifteen commodity AI posts. Depth compounds; listicles evaporate. When a piece lands it will be linked from writing and this page.
Talk it through with someone who runs this stack on his own systems every day.