Google shipped Gemini 3.6 Flash instead of the Pro everyone was waiting for
Google DeepMind released Gemini 3.6 Flash and 3.5 Flash-Lite on 21 July — a cheaper, faster tier built specifically for agentic coding and tool-calling workloads — while Gemini 3.5 Pro, originally due in June, is still stuck in limited preview with no confirmed date.
22 July 2026
On 21 July, Google DeepMind released three new models — Gemini 3.6 Flash, 3.5 Flash-Lite, and a specialised 3.5 Flash Cyber build for vulnerability research — and teased Gemini 4. What it didn’t ship was Gemini 3.5 Pro, the flagship model originally slated for June that’s now missed at least two rumoured dates and remains in limited enterprise preview with no confirmed launch window.
The detail that matters more than the miss is what Google chose to ship in its place. Gemini 3.6 Flash is explicitly positioned for “agents that reason, call tools, inspect visual inputs, edit code, and continue across multiple steps” — this is a model built for the agentic coding workflows that Cursor, Claude Code, Windsurf and their peers run on, not for chat. It’s also cheaper and more token-efficient than its predecessor, cutting token usage by up to 17% at a lower price point.
This is the second time in a month a major lab has led with a cheaper, faster tier while its flagship model slips (Anthropic did something similar with Fable 5’s export-control saga). Read together, it suggests the commercial pressure in AI right now is on the cost and speed of models doing agentic, tool-calling work — the exact category that AI coding assistants and AI-native product features sit in — rather than on chasing frontier reasoning benchmarks that most production use cases don’t actually need.
So what
If you’re specifying which model powers an AI feature or an internal coding agent, the flagship isn’t automatically the right default. A cheaper, faster model tuned for tool-calling and multi-step agent work will often out-perform a more expensive general-purpose flagship on the tasks that actually make up an agentic workflow — and the cost difference compounds fast at production volume. This is a live decision on every AI product build we scope, not a one-off pick made at kickoff. If you’re weighing model choice for an AI feature or agent, see our AI products work or get in touch to talk through the tradeoffs for your build.