Claude now writes 80% of Anthropic's own production code — and its test infrastructure nearly buckled under the load
Anthropic's engineering team now ships roughly eight times more code per quarter than in 2021–25, with Claude authoring about 80% of it, which drove a 25-fold increase in CI jobs and a 10-fold increase in test volume over six months — a real-world case study in what verification and testing infrastructure has to become once AI-assisted development scales past the pilot stage.
16 September 2026
Anthropic published an engineering post this month with numbers worth sitting with: its engineers now ship about eight times as much code per quarter as they did between 2021 and 2025, and Claude authors roughly 80% of that output. That volume increase drove a 25-fold rise in continuous-integration jobs and a 10-fold increase in the number of tests across the codebase, over just six months. This is the company that builds the AI writing the code, running into the same scaling wall that any team adopting agentic coding at pace eventually hits — just at a scale most organisations will never approach.
The part that matters more than the 80% headline
The more instructive detail is what broke, and what it took to fix it. Anthropic’s test-impact-analysis service — the system that decides which tests actually need to run for a given change — got patched three times as load grew, with each fix buying less time than the last: seventy days, then twenty-nine, then under a day, before the team replaced the whole architecture with a distributed, in-memory design built to scale independently of any single service. In other words: more AI-generated code doesn’t just mean more code to review, it means the infrastructure around verification — CI capacity, test selection, build pipelines — has to be engineered for a different order of magnitude, or it becomes the actual bottleneck long before code quality does.
Why this is the opposite of a cautionary tale about AI writing code
The headline “80% of code is AI-written” invites a lazy read: that oversight has been abandoned. The opposite is true here — Anthropic’s test volume grew faster than its code volume, and it rebuilt core infrastructure specifically to keep every one of those tests running reliably at the new scale. The lesson for anyone commissioning AI-assisted development isn’t “trust the AI more,” it’s “the discipline has to scale with the throughput.” A team that lets an AI coding tool multiply output without multiplying its testing and review capacity in step is heading for exactly the kind of strain Anthropic hit — just without the engineering budget to fix it in a weekend.
So what
Ask any development partner using agentic coding tools at pace a direct question: how has your testing and CI setup changed as your AI-assisted output has grown? “It hasn’t needed to” is the wrong answer. Anthropic’s own numbers show that AI-assisted development done properly means your verification infrastructure scales in lockstep with your code volume, not behind it. That’s the standard we hold builds to under our AI-assisted development approach — high throughput from the tools, matched testing discipline underneath it.