Google spent July 21 doing something its rivals in the AI race haven't been doing much lately: shipping the mid-tier models people actually use for production work. The company rolled out three additions to its Gemini lineup — 3.6 Flash, 3.5 Flash-Lite, and a specialized security model called 3.5 Flash Cyber — while staying quiet on the flagship update everyone had been waiting for.
Gemini 3.6 Flash, priced at $1.50 per million input tokens and $7.50 per million output tokens, trims token usage by roughly 17% compared to its predecessor and pushes the model's knowledge cutoff forward 14 months, from January 2025 to March 2026. On SWE-Bench Pro, a coding benchmark built around real-world software tasks, it scores 58.7% against the old model's 55.1%.
Gemini 3.5 Flash-Lite undercuts it on price at $0.30 per million input tokens while still outrunning some pricier models on coding tests, delivering up to 350 tokens per second. And 3.5 Flash Cyber — trained specifically to hunt and patch security vulnerabilities — is restricted for now to governments and select partners through Google's CodeMender agent.
What's missing from the announcement says as much as what's in it. Gemini 3.5 Pro, the flagship model Google had promised for June, still hasn't shipped. Bloomberg has reported the model has been falling short of internal quality targets. At the same time, Google confirmed it has begun what it calls its most ambitious pretraining run yet — for Gemini 4, the next full generation of the family.
Read together, the two facts tell a coherent story. Google is patching the gap in its mid-tier lineup while its top model stalls, and simultaneously starting the next generation from scratch. That timing matters because OpenAI has already rolled out GPT-5.6 and Anthropic has pushed Claude Opus 4.8 and Sonnet 5 into the market. A stalled Gemini 3.5 Pro isn't just a missed date — it's ground Google needs to make up before Gemini 4 is ready to ship.



