Google's new Gemini Flash models are built for AI that runs itself

Google DeepMind's 3.6 Flash and 3.5 Flash-Lite promise cheaper, faster AI agents. Here is what that means for the apps and games you actually use.

AI2Day Newsdesk· 3 min read
Photoreal editorial shot of a sleek modern data centre corridor at night, rows of server racks glowing with soft blue and amber light, faint reflections on poli
Share

Key points

  • Google DeepMind announced Gemini 3.6 Flash and Gemini 3.5 Flash-Lite on release day, aimed at running AI agents at scale.
  • Google says 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, at $1.50 per million input tokens and $7.50 per million output tokens.
  • Gemini 3.5 Flash-Lite runs at 350 output tokens per second, priced at $0.30 per million input tokens and $2.50 per million output tokens.
  • On the OSWorld-Verified benchmark, which tests an AI's ability to use a real computer, 3.6 Flash scored 83.0% against 78.4% for 3.5 Flash.
  • Google DeepMind confirmed pre-training has started on Gemini 4, its next flagship model.

Google just quietly changed the maths for anyone building an AI that does things for you.

On release day, Google DeepMind rolled out two new versions of its Gemini family: Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Gemini is Google's line of AI models, the software brains behind chatbots and, increasingly, AI agents (programs that carry out multi-step tasks on your behalf, like booking travel or fixing code).

The headline isn't a bigger, smarter monster model. It is the opposite. These are the workhorses.

Why should a normal person care about a cheaper AI model?

Because price and speed are what decide whether the clever demo you saw last year actually shows up inside the apps you use.

Running a large AI model costs money every time it answers. That cost is measured in tokens, which are just chunks of text the model reads and writes. Google says Gemini 3.6 Flash uses 17% fewer output tokens than the previous 3.5 Flash to do the same job, according to the independent Artificial Analysis Index. On one coding test called DeepSWE, the saving hit 65%.

In plain English: the model rambles less and gets to the point faster. That drops the bill for the company building your app, which is how features like an in-game AI companion or a smart customer service bot become financially possible.

Pricing sits at $1.50 per million input tokens and $7.50 per million output tokens. A million tokens is roughly 750,000 words.

The little sibling might be the real story

Gemini 3.5 Flash-Lite is the small, fast one. Google clocks it at 350 output tokens per second, priced at $0.30 per million input tokens and $2.50 per million output tokens.

That speed matters for anything that needs to feel instant. Think of an AI opponent in a game that has to react in real time, or a document-processing tool chewing through thousands of invoices overnight.

On Terminal-Bench 2.1, a test that measures how well an AI can drive a computer command line, Flash-Lite jumped from 31% to 54%. On OSWorld-Verified, which asks the AI to actually operate a desktop like a human would, it hit 74%.

Both new models now include "computer use" as a built-in tool. That means the AI can click buttons, fill in forms and move a cursor without a developer wiring it up by hand.

What about safety?

Google says 3.6 Flash ships with stronger safeguards against misuse in chemical, biological, radiological, nuclear and cyber-offence scenarios. It has also been trained to refuse fewer harmless requests, a common complaint with earlier Gemini versions that sometimes said no to boring, everyday questions.

There is a specialist cousin too. Gemini 3.5 Flash Cyber powers CodeMender, Google's agent that hunts for security holes in software code.

What happens next?

Gemini 3.5 Pro, the bigger flagship model, is still in testing with partners. And Google confirmed the interesting bit at the end: pre-training has begun on Gemini 4.

For players, developers and anyone waiting for AI features that don't cost a fortune to run, the direction is clear. The frontier gets the headlines. The Flash models get shipped.

© 2026 AI2Day