Google DeepMind is rolling out three new Gemini models focused on cheaper and faster AI for large-scale use: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. Google says the lineup improves efficiency for “agentic” systems that run over long periods, targeting reduced token use to lower total costs. VentureBeat reports API pricing for Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens, while Gemini 3.5 Flash-Lite is priced at $0.30 per million input tokens and $2.50 per million output tokens. Gemini 3.5 Flash is described as the prior generation option, and Google’s older Gemini 3.1 Flash-Lite remains the company’s most cost-efficient model, though it is slower than the newer Flash-Lite.

Across third-party benchmarks cited by VentureBeat, Google says Gemini 3.6 Flash reduces output token usage by 17% versus Gemini 3.5 Flash, with savings up to 65% on long-horizon engineering tasks such as DeepSWE. It also reports gains on additional evaluations and highlights safety guardrails.

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available through Google’s Gemini API and related products, while Gemini 3.5 Flash Cyber is planned for limited access via Google’s CodeMender for governments and trusted partners. Google also says Gemini 3.5 Pro remains under partner testing and will launch broadly once ready, and that Gemini 4 pre-training has started.