Google's Gemini 3.8 Flash Thinks Harder, But That Extra Effort Will Cost You More

Same price per token as its predecessor, but Google warns the new model uses more tokens to get better results. Here is what developers and everyday subscribers need to know.

AI2Day Newsdesk4 min read
Abstract lattice of light filaments suspended on near-black
Share

Key points

  • Google launched Gemini 3.8 Flash in May 2025, just weeks after Gemini 3.7 Flash.
  • Per-token pricing stays the same: $0.75 per million words fed in, $3.75 per million words generated.
  • Early testing found total task costs running roughly 40% higher than 3.7 Flash because the model produces more output.
  • The model topped two industry benchmarks for software coding and legal AI agents.
  • A separate, restricted version called Gemini 3.8 Flash Cyber is available only to governments and vetted security partners.

Google's newest AI model is faster and smarter than its predecessor. It is also likely to run up a bigger bill.

Gemini 3.8 Flash arrived just weeks after Gemini 3.7 Flash, and Google says the upgrade is simple to summarise: the model "works harder." On complex tasks it carries out more reasoning steps and calls outside tools repeatedly until it finds a good answer. Think of it like the difference between an employee who gives you a first draft and one who reads it back, spots the gaps, and sends three revised versions before handing it over.

What does it actually cost?

The sticker price is unchanged. Google charges $0.75 per million input tokens and $3.75 per million output tokens. A token, for context, is roughly three-quarters of a word.

But Google itself flags the catch: "the model might use more tokens to maximize performance, especially at higher effort levels." Early numbers from AI benchmarking firm Artificial Analysis, first highlighted by The Verge AI, put the real-world bill about 40% higher than 3.7 Flash. The firm traced that to a 30% increase in the words the model generates per task, plus more back-and-forth on agentic tasks. An AI agent, here, means software that can carry out multi-step jobs on its own, like filing a support ticket or running a web search and summarising the result.

Developers who want to hold the line on spending can stay on Gemini 3.7 Flash. Google confirmed both models remain available.

Model Input (per 1M tokens) Output (per 1M tokens) Est. cost vs 3.7 Flash
Gemini 3.7 Flash $0.75 $3.75 baseline
Gemini 3.8 Flash $0.75 $3.75 ~40% higher in practice

Is it actually better at real work?

Yes, on the tasks Google chose to test. The model beat competitors on three external benchmarks: DeepSWE v1.1 for software engineering, Vals Finance Agent V2 for financial AI tasks, and Harvey's Legal Agent benchmark.

For comparison, it outscored Anthropic Claude's flagship model on the coding benchmark. Aigora.ai CEO John Ennis described the result as "Opus 5 coding quality but at a fraction of the cost and super fast."

That quote deserves a small asterisk. CEOs announcing that a cheap model matches a premium one are rarely a neutral source. Real-world results vary by task.

What about the security version?

Alongside the main release, Google quietly launched Gemini 3.8 Flash Cyber, a version tuned for offensive security research. It is not open to the public.

Access goes only through Google's new Fairwind Program, a vetted group of 650 members including cybersecurity firm CrowdStrike and the Center for Internet Security. Participants can also use CodeMender, an AI agent designed to find and fix software vulnerabilities automatically. Google says the standard 3.8 Flash ships with safeguards against misuse in chemical, biological, radiological and nuclear domains, as well as cyber attack assistance.

What should you do?

If you are a Google AI Pro or Ultra subscriber, 3.8 Flash is available now in your existing plan at no extra charge per session, though heavier use may push you toward a higher tier faster.

If you are a developer paying per token, run a small test batch before switching. Check your actual token counts against your 3.7 Flash baseline. A 40% cost increase on a modest workload might be worth the quality jump. On a large one, the maths deserve a closer look before you commit.

© 2026 AI2Day