Google Launches Three New Gemini Models and Teases Gemini 4 While Its Flagship Sits in Limbo

Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber are live - and Google confirmed its Gemini 4 pre-training has already begun.

Saganote
Saganote ·
3 Min Read

Google launched three new Gemini models on Tuesday - Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber - keeping developers fed while its frontier model continues to miss deadlines. Gemini 3.5 Pro has now been delayed past June and July, with Google saying only that partners are testing it and it will ship "when it is ready." As a brief consolation, the company also confirmed Gemini 4 pre-training has started.

Gemini 3.6 Flash: The New Mainstream Model

Gemini 3.6 Flash steps into the role of Google's everyday workhorse, covering coding, knowledge work, multimodal processing, and computer-use tasks. Against Gemini 3.5 Flash on the Artificial Analysis Index, it uses 17% fewer output tokens while delivering more accurate results. Benchmark improvements are real: DeepSWE jumped from 37% to 49%, MLE Bench went from 49.7% to 63.9%, and OSWorld-Verified climbed from 78.4% to 83%.

Pricing lands at $1.50 per million input tokens and $7.50 per million output tokens - identical positioning to where 3.5 Flash sat. Gemini 3.6 Flash is available now via the Gemini API, Google AI Studio, Android Studio, Gemini Enterprise, and the Gemini app. Developers comparing these costs against Claude Fable 5, GPT-5.5, and other frontier models will find it priced as a capable mid-tier option.

Gemini 3.5 Flash-Lite: 350 Tokens per Second, Built for Speed

Flash-Lite targets latency-sensitive jobs: agentic search pipelines, document processing, receipt analysis, and translation. Speed is the headline - up to 350 output tokens per second. Cost drops to $0.30 per million input tokens and $2.50 per million output tokens, making it one of Google's cheapest production options. Benchmark gains over the prior Flash-Lite are substantial: Terminal-Bench 2.1 rose from 31% to 54%, and the GDM-MRCR v2 long-context benchmark improved from 60.1% to 72.2%.

Flash-Lite is available in the same channels as 3.6 Flash, plus it is rolling out to Google Search directly. That Search integration matters - it means Flash-Lite will handle AI Overviews at scale, so its speed ceiling directly affects how fast Google's core product responds to queries.

Gemini 3.5 Flash Cyber: Vulnerability Research, Governments Only - For Now

Gemini 3.5 Flash Cyber is positioned as a low-cost alternative to GPT-5.5 Cyber and Claude Mythos 5. Google built it to identify, validate, and patch software vulnerabilities, and is running it through its CodeMender security agent. Availability is restricted: initially governments and trusted partners only, under a limited pilot. No general API access date was announced.

Gemini 4 Is Pre-Training - No Timeline Given

Google's closest thing to a headline was the Gemini 4 mention. The company said its team has begun "the most ambitious pre-training run yet" for the next generation. No dates, no benchmarks, no architecture details. For context, Gemini 3.5 Pro was announced months before its still-incomplete release - so a Gemini 4 pre-training announcement today likely means availability is well over a year out.

Three Flash-class models and a tease is a reasonable holding pattern for a company managing both a compute crunch and a delayed flagship. Google is already building dedicated inference silicon to serve Gemini more efficiently - adding faster, cheaper Flash models now keeps API adoption moving while the infrastructure and frontier model work catches up.


Share this
Saganote

About Author

Saganote

Saganote is an independent technology publication covering artificial intelligence, startups, cybersecurity, consumer technology, science, and innovation. Our editorial team reports on the companies, products, and ideas shaping the future.