Google DeepMind launches cost-focused Gemini models, delays flagship Pro release
TechCrunch reports on Google’s latest AI rollout, highlighting a strategic pivot toward lower-cost, high-efficiency models while the anticipated Gemini 3.5 Pro faces internal delays.

Google DeepMind has released three new iterations of its Gemini artificial intelligence family, marking a distinct shift toward cost efficiency and specialised utility. The new suite includes Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. According to the company, the primary objective of this release is to enhance efficiency, latency, and reliability for developers building AI agents at scale.
Gemini 3.6 Flash is positioned as the primary workhorse model, offering improved capabilities in coding, knowledge work, and multimodal performance. It consumes up to 17% fewer output tokens than its predecessor, Gemini 3.5 Flash, according to the Artificial Analysis Index. Pricing for the new model is set at $1.50 per million input tokens and $7.50 per million output tokens, representing a cost reduction for users.
The second variant, Gemini 3.5 Flash-Lite, is described as the most cost-effective model in the class, targeting applications where lower computational expense is critical. The third model, Gemini 3.5 Flash Cyber, is a specialised version fine-tuned to detect and fix cybersecurity vulnerabilities. It is available exclusively to governments and trusted partners through a limited-access pilot program via CodeMender.
The release notably omits the anticipated Gemini 3.5 Pro, which was last updated in February. This gap has drawn attention as competitors have maintained a rapid release pace, with OpenAI launching GPT-5.5 and beginning to roll out GPT-5.6, while Anthropic has released Claude Opus 4.8, Claude Sonnet 5, and expanded access to its Fable 5 model.
Google had previously teased the Gemini 3.5 Pro in May, stating it was already being used internally and expected to roll out the following month. However, a Bloomberg report last week indicated that the company faced internal delays due to struggles in meeting performance goals. Google DeepMind product lead Logan Kilpatrick confirmed on Tuesday that the model is currently being tested with partners, with the company hoping to launch it soon.
In addition to the current releases, Kilpatrick noted that the team has commenced its most ambitious pre-training run yet for the future Gemini 4 model. The omission of the Pro tier suggests a continued focus on Flash models, which prioritise lower cost and faster response times for production applications, rather than the highest-capability offerings typically reserved for complex reasoning tasks.
