The AI arena is free today

Open Superagent
Back to blog
Model Release·Technical Deep Dive

Gemini 3.7 Flash Release, Benchmarks And Price

Gemini 3.7 Flash scores 65.3% on DeepSWE v1.1 and 43.6% on FrontierCode 1.1 Main, at $0.75/$3.75 through 2026. Same 1M context as 3.6 Flash. List price doubles on January 1, 2027.

Sebastian Crossa
Sebastian Crossa
Co-Founder @ LLM Stats
·8 min read
Gemini 3.7 Flash Release, Benchmarks And Price

Key Numbers

Gemini 3.7 Flash · Aug 13, 2026

0.0%
DeepSWE v1.1
0.0%
FrontierCode 1.1 Main
0
WebDev Arena Elo
0.0%
GDP.pdf
0.0%
AutomationBench
0M
Context

Price Clock

intro expires Dec 31, 2026

$0.75 / $3.75
Intro through Dec 31, 2026
per 1M tokens
$1.50 / $7.50
From Jan 1, 2027
per 1M tokens

Intro is half the original 3.6 Flash list. Google's footnote: the introductory rate expires December 31, 2026.

Google released Gemini 3.7 Flash on August 13, 2026, three weeks after 3.6 Flash. Same 1,048,576 / 65,536 context class. Same multimodal inputs (text, image, video, audio, PDF). The story is a workhorse Flash that moves on coding and real workflows, at an introductory $0.75 / $3.75 per 1M tokens through December 31, 2026.

The headline vendor numbers vs 3.6 Flash: 65.3% on DeepSWE v1.1 (from 49.0%) and 43.6% on FrontierCode 1.1 Main(from 34.4%). Around them sit WebDev Arena Elo 1588 vs 1538, GDP.pdf 34.0% vs 22.0%, and AutomationBench 30.4% vs 17.0%. The benches move. The more interesting change is the price clock: half-off 3.6 list until New Year's, then $1.50 / $7.50 unless Google extends it.


At a Glance

  • Release: August 13, 2026. Generally available.
  • Model ID: gemini-3.7-flash (stable) on the Gemini API.
  • Intro price through Dec 31, 2026: $0.75 / $3.75 per 1M.
  • From Jan 1, 2027: $1.50 / $7.50.
  • Context: 1,048,576 in / 65,536 out (same as 3.6 Flash).
  • Modalities: text, image, video, audio, PDF in. Text out. No native image or audio generation on this ID. Live API not supported.
  • Thinking: low / medium / high. minimal is not supported and returns an error.
  • Also: caching, code execution, function calling, search grounding, Maps grounding, structured outputs, URL context, file search, computer use (preview). Batch, Flex, Priority inference.
  • Where: Gemini API (AI Studio, Android Studio), Google Antigravity, Gemini Enterprise Agent Platform and app. Spark for Google AI Pro and Ultra (160+ countries) runs on 3.7 Flash starting today.

What's New in 3.7

None of this is a new context class. Together it is a faster hop on the Flash workhorse: better first-pass code, denser document/workflow scores, and a temporary list-price cut.

A three-week follow-up, not a new family

Google frames 3.7 as the next Gemini 3 iteration from developer feedback and algorithmic innovations it says will land in later models. Public surface is the same Flash contract: 1M in, 64K out, natively multimodal, thinking on.

First-pass code and fewer retries

Vendor coding claim: higher first-pass accuracy. Numbers: FrontierCode 1.1 Main 43.6% vs 34.4%, DeepSWE v1.1 65.3% vs 49.0%. Web: more functional layouts in fewer prompts, design adherence from a screenshot/image/design system. WebDev Arena 1588 vs 1538.

Developer experience (claim, not a bench)

Google's developer-experience pitch is not a bench: thinks more diligently on multi-step planning and tool calls, adapts to roadblocks, asks for clarification. Treat as a claim to test on your own traffic.

Spark, Antigravity, Enterprise

Spark (24/7 personal agent for Pro/Ultra) switches to 3.7 today. Examples are Workspace-shaped. Developers: Antigravity and Gemini API. Enterprises: Agent Platform and Gemini Enterprise app.


Benchmarks

All scores are self-reported by Google in the Aug 13 launch post. Not LLM Stats verified. The blog has a Detailed benchmarks block that did not render as a usable table; we do not invent extra rows.

3.7 Flash vs 3.6 Flash

3.7 Flash3.6 Flash
DeepSWE v1.1
65.349.0+16.3
FrontierCode 1.1 Main
43.634.4+9.2
AutomationBench
30.417.0+13.4
GDP.pdf
34.022.0+12.0
Self-reported by Google, Aug 13 launch post. Not LLM Stats verified. DeepSWE is the movement. GDP.pdf and AutomationBench are still low in absolute terms. Scores on a 0-100 scale. WebDev Arena Elo is omitted (not a percent bench).
Benchmark3.7 Flash3.6 FlashDelta
FrontierCode 1.1 Main43.6%34.4%+9.2
DeepSWE v1.165.3%49.0%+16.3
GDP.pdf34.0%22.0%+12.0
AutomationBench30.4%17.0%+13.4
WebDev Arena Elo15881538+50

DeepSWE is the movement(+16.3). FrontierCode Main is a cleaner +9.2. WebDev Arena is Elo, +50, real but not a regime change. GDP.pdf and AutomationBench jump +12.0 and +13.4 and are still low in absolute terms (34% and 30%). Relative win vs 3.6 Flash, not “document QA is solved.”


Pricing and the 2027 Clock

The 2027 Clock

input / output · per 1M

Half off until New Year's. Then the clock hits.

Intro
$0.75
$3.75

through Dec 31, 2026

List
$1.50
$7.50

from Jan 1, 2027

Intro is half the original 3.6 Flash list ($1.50 / $7.50). Google's footnote: introductory pricing expires December 31, 2026. Budget 2027 at the doubled rate unless a new notice lands.
WindowInput / 1MOutput / 1M
Intro through Dec 31, 2026$0.75$3.75
From Jan 1, 2027$1.50$7.50

Intro is half the original 3.6 Flash list ($1.50 / $7.50). Google's footnote: intro expires December 31, 2026. Budget 2027 at the doubled rate unless a new notice lands.

Caching is supported on the model page. Cloud/Agent Platform tables may publish cached and long-context (>200k) rows that differ by SKU. Prefer the Gemini API pricing page for the ID you actually call. Batch / Flex / Priority are levers on the same ID, not separate Flash models.


When to Use It

Good fit: high-volume coding/agent loops where 3.6 Flash was the default, UI-from-reference, document-heavy workflows, Spark on Pro/Ultra, shops that can burn intro pricing before January.

Watch the calendar: 2x list on Jan 1 rewrites cost-per-task. Re-measure token-per-task. Google says 3.7 puts in more effort on planning and tool calls, which can raise output tokens.

Not a new window: if you needed more than 1M / 64K, this does not give it to you.


Migrating from 3.6 Flash

Drop-in swap to gemini-3.7-flash.

  • 1. Flip the model string. Caps match.
  • 2. Thinking: low / medium / high only. minimal errors.
  • 3. Re-run your own coding and document evals. DeepSWE and AutomationBench are the ones worth reproducing.
  • 4. Put a 2027-01-01 price check on the calendar.
  • 5.Spark users do not pick the ID. The app switch is on Google's side.

Outlook

Incremental Flash release with a real coding delta and a promotional price. Benches vs 3.6 move most on DeepSWE, FrontierCode, GDP.pdf, and AutomationBench. Platform surface is the same 1M Flash contract, plus Spark and Antigravity. The thing that will age this post is January 1, 2027.

For the official announcement and model docs, see Google's launch post, the Gemini 3.7 Flash model page, and Google AI Studio.

Questions

Frequently Asked Questions

  • Google released Gemini 3.7 Flash on August 13, 2026. It is available on the Gemini API (AI Studio, Android Studio), Antigravity, Gemini Enterprise, and Spark for Google AI Pro and Ultra.
  • Introductory pricing is $0.75 / $3.75 per 1M input / output tokens through December 31, 2026. From January 1, 2027, list price is $1.50 / $7.50. That intro rate is half the original 3.6 Flash list until then.
  • Gemini 3.7 Flash supports 1,048,576 input tokens and 65,536 output tokens, matching 3.6 Flash.
  • Vendor scores vs 3.6 Flash: FrontierCode 1.1 Main 43.6% vs 34.4%, DeepSWE v1.1 65.3% vs 49.0%, WebDev Arena 1588 vs 1538 Elo, GDP.pdf 34.0% vs 22.0%, and AutomationBench 30.4% vs 17.0%. Same context window. Half price until the end of 2026.
  • Yes: low, medium, and high. The minimal thinking level is not supported and returns an error.
  • Try it in Google AI Studio, the Gemini API, Antigravity, Gemini Enterprise, and Spark in the Gemini app (Pro / Ultra, supported countries).

Continue Reading