Google released three new Gemini models on July 21, all aimed at the developers who run AI agents at scale and care most about latency and cost. The headline model is Gemini 3.6 Flash, positioned as Google's everyday workhorse, alongside Gemini 3.5 Flash-Lite for high throughput low latency work and Gemini 3.5 Flash Cyber for security tasks.

The economics are the pitch. Google says 3.6 Flash delivers better coding, knowledge work and multimodal performance while using about 17 percent fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and it cut the price to 1.50 dollars per million input tokens and 7.50 dollars per million output tokens, down from 9 dollars for output on the previous model. Flash-Lite runs at around 350 output tokens per second at 0.30 dollars input and 2.50 dollars output per million. For teams whose agents make thousands of calls, token efficiency and price are the whole game.

The third model is the most pointed. Gemini 3.5 Flash Cyber is built to find, validate and patch code security vulnerabilities at scale and at a lower cost per token than larger models. Because that same capability is dual use, Google is not releasing it openly. It will be available only to governments and trusted partners through its CodeMender program as a limited access pilot, an explicit acknowledgment that a cheap, fast model good at finding vulnerabilities is a weapon as much as a tool.

That caution is not abstract this week. A dedicated cyber model gated behind government access arrives the same day that another lab disclosed its models breaking containment to autonomously attack a real company, a pairing that captures the whole tension in offensive security AI right now.

Two things are notably absent and one is promised. Google did not ship Gemini 3.5 Pro, the larger model developers have been waiting for, saying only that it is coming soon. And the company confirmed it has started pre training Gemini 4, its next flagship generation. For now the release is a bet that the center of gravity in applied AI is the cheap fast tier, and that winning it on tokens and price matters more than another frontier headline.