Terminal-Bench 4.0 arrives with harder tasks — and most AI coding agents still fail 70% of them
Vals AI's Terminal-Bench 4.0, a fresh set of 66 expert-reviewed coding tasks released this month, shows GPT-6 Astra narrowly leading Claude Fable 5.1 at the top (58.2% vs 57.9%), but most evaluated AI coding agents still fail at least 70% of its tasks — a reminder that frontier benchmark races and reliable AI-assisted development are two different questions.
20 September 2026
Vals AI published Terminal-Bench 4.0 this month — a full rebuild of the benchmark used to grade AI coding agents on real, end-to-end terminal work, with 66 new community-contributed, expert-reviewed tasks that share nothing with the previous version. The tasks aren’t toy problems: the median one is estimated at four hours of expert human work, with some running past 60 hours, and each is graded strictly on whether the final deliverable actually works.
At the top of the new leaderboard, GPT-6 Astra edges out Claude Fable 5.1 by a hair — 58.2% to 57.9% — continuing a pattern that’s held for months now, where the gap between leading models has narrowed to fractions of a percentage point. Zoom out from that headline race, though, and the more decision-relevant number is this: most AI coding agents evaluated against Terminal-Bench 4.0 still fail at least 70% of its tasks.
Two numbers, two different stories
The 58% figure is a benchmark story — useful for engineers picking a model, but not the number that should drive a commissioning decision. The 70%-plus failure rate is the operational story, and it’s the one that matters if you’re a founder or CTO weighing how much of a build to hand to an AI agent versus a reviewed engineering process. These aren’t autocomplete-style tasks; they’re the kind of sustained, multi-step work — refactors, environment setup, debugging across a real codebase — that a development team is actually paid to get right. A benchmark built around that work still shows the frontier losing most of the time on the harder end of the task distribution.
Why the race at the top keeps tightening
Claude Fable 5.1 still leads on broader indices — Artificial Analysis’s Intelligence Index (66 vs Astra’s 61) and its Coding Agent Index (70 vs 67) — while GPT-6 Astra takes Terminal-Bench 4.0 specifically, particularly on tasks involving computer use and long-context retrieval. That split matters more than either single number: it says the “best” AI coding model increasingly depends on the shape of the task, not a single leaderboard position.
So what
Terminal-Bench 4.0 is a useful gut-check against the idea that AI coding tools have solved sustained, autonomous software delivery. They haven’t — not even the frontier models, not on tasks explicitly designed to look like real engineering work. What’s changed is capability at the margins, not the need for review discipline, testing, and architectural judgement around the tools. If you’re evaluating what an AI-assisted build can realistically take off your plate versus what still needs an engineer checking the output, that’s exactly the conversation our AI-assisted development approach is built around.