- Organizations
- Gemini 3.5 Flash
Gemini 3.5 Flash: Benchmarks, Pricing & Context Window
Gemini 3.5 Flash is a language model from Google, released in May 2026, with multimodal input, a 1.0M-token context window, and pricing from $1.50/M input and $9.00/M output.
Gemini 3.5 Flash is Google's strongest agentic and coding model in the Flash series, delivering frontier-level performance at up to 4x the speed of comparable frontier models and often at less than half the cost. Built to execute complex,
Gemini 3.5 Flash benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Gemini 3.5 Flash across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Gemini 3.5 Flash holds up as conversations get longer.
Quality Tracker
Gemini 3.5 Flash Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Gemini 3.5 Flash pricing
Providers
Gemini 3.5 Flash starts at $1.50 per million input tokens and $9.00 per million output tokens via DeepInfra. See all 2 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $1.50 | — | $9.00 | 1.0M/1.0M | — | — | / | |
| $1.50 | — | $9.00 | 1.0M/65.5K | 2.70 | 564 | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Gemini 3.5 Flash context window
Input and output token limits for Gemini 3.5 Flash, plus how it ranks on long-context understanding.
Try now
Make it with
Gemini 3.5 Flash.
Gemini 3.5 Flash latency
Gemini 3.5 Flash time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.
Provider operational metrics
Time to first token, output throughput, and failed-request rate from live model usage
Gemini 3.5 Flash examples
Recent arena outputs from Gemini 3.5 Flash, picked from the highest-ranked matchups.
Gemini 3.5 Flash license
Gemini 3.5 Flash is a proprietary model available under its provider's product and API terms, has a knowledge cutoff of January 2026.
- License
- Proprietary
- Hosted access
- Knowledge cutoff
- January 2026
Proprietary license - usage restrictions apply
Gemini 3.5 Flash resources
Official sources for Gemini 3.5 Flash: provider documentation, official playground, official launch post.
Gemini 3.5 Flash vs other models
The most-compared alternatives to Gemini 3.5 Flash are GPT-5.2, Qwen3.7-Plus, DeepSeek-V4-Pro-Max. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Gemini 3.5 Flash
Models ranked just above and below Gemini 3.5 Flash by LLM Stats score.
FAQ
Common questions about Gemini 3.5 Flash.