A startup just raised $10m to fix the reason 88% of AI agent pilots never reach production — by cloning your actual software first
Arga Labs raised a $10m seed round led by General Catalyst, reported 26 August 2026, to build high-fidelity sandbox 'digital twins' of enterprise software like Salesforce and Workday so AI agents can be trained and tested against realistic conditions before going live — a direct market response to the well-documented pilot-to-production gap in AI agent deployment.
28 August 2026
Arga Labs, a Y Combinator-backed startup founded in 2025, raised a $10m seed round led by General Catalyst — with Box Group, Emergence, Gradient and SV Angel also participating — reported by TechCrunch on 26 August 2026. What it builds is narrow and specific: full-scale sandbox environments that clone real enterprise software like Salesforce, Workday and email clients, permission systems and webhooks intact, so an AI agent can be trained and stress-tested against something that behaves like the real system before it ever touches production data.
That’s a direct answer to a problem this site flagged in detail earlier this month: Forrester and Anaconda research puts the AI agent pilot-to-production failure rate at 88%, with the top blocker being evaluation gaps — no reliable way to test whether an agent’s output is actually correct before it ships. Arga’s founders, CEO Phillip Li (previously at Amazon) and CTO Akira Tong (previously at Stripe and Goldman Sachs), are betting that most AI agent failures aren’t a model-quality problem at all — they’re a testing-infrastructure problem, because most teams are evaluating agents against stateless API mocks that don’t reproduce the permission edge cases, rate limits, and workflow quirks of the real software the agent will eventually run inside.
Why a funding round for test infrastructure is a signal, not just a data point
Venture money chasing “how do we safely test an AI agent before it’s live” is a tell about where the market actually is. A year ago, most investment in this space went to the agents themselves — better models, faster completions, broader tool access. Money now flowing into the unglamorous layer underneath — sandboxes, evaluation harnesses, digital twins — suggests the industry has broadly accepted that the model was never the bottleneck. The bottleneck is proving an agent won’t break something real before you let it near something real.
So what
If you’re a founder or CTO weighing whether an AI agent feature — customer-facing or internal — is ready to leave the pilot stage, “how was this actually tested against production-like conditions” is now a fair, specific question to ask any team building it, in-house or agency. It’s also exactly the kind of evaluation and governance layer that separates a working AI product from a demo that never ships. Our AI-assisted development work is built around that gap between pilot and production — get in touch if you’re trying to work out where your own agent project actually stands.