LMTimeline ← Back to the timeline

Google releases Gemini 3.6 Flash and 3.5 Flash-Lite

Google refreshed its mid-tier Flash line with two proprietary models. Gemini 3.6 Flash cuts cost and verbosity, roughly 17% fewer output tokens than 3.5 Flash, while improving coding (DeepSWE 49% vs 37%), computer use (OSWorld-Verified 83.0%) and ML research (MLE-Bench 63.9%), priced at $1.50/$7.50 per million input/output tokens. Gemini 3.5 Flash-Lite targets high-throughput agentic work at about 350 tokens/second with a built-in computer-use tool, priced at $0.30/$2.50. Both ship across Google AI Studio, Gemini Enterprise and the Gemini app.

Google

Google shipped Gemini 3.6 Flash and Gemini 3.5 Flash-Lite on July 21, 2026. Both are proprietary, both were available the same day, and the announcement is built around cost per finished task rather than raw intelligence. That framing matters more than it sounds, because the economics of an AI agent are set by how many words it burns getting to an answer, not by how clever it is on a single question.

Gemini 3.6 Flash

The efficiency claim leads. Google says 3.6 Flash produces about 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and on some benchmarks such as DeepSWE the reduction reaches 65%. Output tokens are the expensive half of any AI bill, so a model that reaches the same place in fewer words costs less to run even at an unchanged price.

Capability moved as well, concentrated in coding and agent work:

  • DeepSWE, a software engineering test: 49%, up from 37%
  • MLE-Bench, which measures machine learning research tasks: 63.9%, up from 49.7%
  • OSWorld-Verified, where the model operates a computer to finish a job: 83.0%, up from 78.4%
  • GDPval-AA v2, covering general knowledge work: 1421, up from 1349

Pricing is $1.50 per million input tokens and $7.50 per million output tokens. You can use it through Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform and the Gemini app. Google notes extra Frontier Safety safeguards on this release covering biological and chemical weapons misuse along with cyber offense.

Gemini 3.5 Flash-Lite

Flash-Lite is built for volume. Google quotes roughly 350 output tokens per second, which is fast enough that text appears quicker than most people read, and it ships with a built in computer use tool. That last detail is unusual at this price. It means the cheapest model in the lineup can drive a browser or an application directly, which is normally a premium feature.

The jump over the previous Flash-Lite is steep:

  • Terminal-Bench 2.1, working in a command line: 54%, up from 31%
  • GDM-MRCR v2, recalling detail from very long documents: 72.2%, up from 60.1%
  • GDPval-AA v2: 1140, up from 642
  • OSWorld-Verified: 74.0%, up from 65.1%

It also beats the older Gemini 3 Flash on SWE-Bench Pro, 54.2% against 49.6%. A budget model outscoring the previous generation's standard model is the result worth pausing on, because it means the floor moved rather than the ceiling.

Pricing is $0.30 per million input tokens and $2.50 per million output tokens, and it is rolling out through Google AI Studio, Android Studio, Gemini Enterprise and Google Search.

What Google left out

Neither model got a stated context window, the amount of material it can hold in mind at once. That is an odd gap in a release that advertises a long context benchmark result. There is also no independent testing yet, so every figure above is Google's own measurement.

Two other things surfaced in the same post. Gemini 3.5 Pro is being tested with partners ahead of a wider release, and pre-training has begun on Gemini 4. The announcement was written by Tulsee Doshi, a senior director of product management at Google.

Sources

LMTimeline writes its own account of each event. Primary sources are linked above so you can read them directly.