Signal

Terminal-Bench 4.0 arrives with harder tasks — and most AI coding agents still fail 70% of them

Vals AI's Terminal-Bench 4.0, a fresh set of 66 expert-reviewed coding tasks released this month, shows GPT-6 Astra narrowly leading Claude Fable 5.1 at the top (58.2% vs 57.9%), but most evaluated AI coding agents still fail at least 70% of its tasks — a reminder that frontier benchmark races and reliable AI-assisted development are two different questions.