88% of AI agent pilots never reach production, Forrester and Anaconda find — the gap between 'we tried Claude Code' and 'it runs our business'
New Forrester and Anaconda research puts the AI agent pilot-to-production failure rate at 88%, with evaluation gaps, governance friction and model reliability cited as the top blockers — a sharp contrast to adoption figures that show most developers already using AI coding tools daily.
10 August 2026
Adoption and deployment have quietly become two different numbers. Stack Overflow’s 2026 developer survey puts daily AI coding agent use at 71% of professional developers, and GitHub Copilot alone claims 40% adoption in large enterprises. But new research from Forrester and Anaconda finds that 88% of AI agent pilots never reach production — the largest gap between “we’re using this” and “this is load-bearing” that either firm has recorded.
The blockers aren’t mysterious. Leaders cite evaluation gaps — no reliable way to test whether an agent’s output is actually correct before it ships — as the top obstacle (64%), followed by governance friction (57%) and model reliability (51%). None of these are solved by a better model. They’re solved by the unglamorous work of building test harnesses, review gates, and rollback paths around the model, which is exactly the part a Cursor or Claude Code subscription doesn’t include.
The industry split makes the pattern clearer still. Banking and insurance run 47% of their AI agent pilots through to production; software and internet companies aren’t far behind at 44%. Healthcare and government trail badly, at 18% and 14%. That’s not a capability gap — it’s a compliance and procurement gap. The sectors with the strictest audit and governance requirements are the ones where “it worked in the demo” is furthest from “it’s safe to ship,” because they’re the only ones forced to prove it before they can.
So what
If your organisation has AI coding tools in every developer’s hands but nothing running in production, you’re not behind — you’re at the median. The 88% figure is a useful gut-check against any vendor pitch that treats agentic coding as already-solved: the model is rarely the constraint anymore, the surrounding engineering discipline is. That’s also the case for treating “vibe coded” prototypes as production-ready without the evaluation, governance, and reliability work layered on top. If you’re weighing whether to build a pilot in-house or bring in a team that’s already solved the production gap, our approach to AI-assisted development is built around exactly that gap, or get in touch to talk through where your pilot actually stands.