The AI arena is free today

Open Superagent
GoogleReleased on Jun 17, 2025

Gemini 2.5 Flash-Lite: Benchmarks, Pricing & Context Window

Gemini 2.5 Flash-Lite is a language model from Google, released in June 2025, with multimodal input.

Gemini 2.5 Flash-Lite is a model developed by Google DeepMind, designed to handle various tasks including reasoning, science, mathematics, code generation, and more. It features advanced capabilities in multilingual performance and long

Input
TextImage
Output
Text

Gemini 2.5 Flash-Lite benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for Gemini 2.5 Flash-Lite across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How Gemini 2.5 Flash-Lite holds up as conversations get longer.

Quality Tracker

Gemini 2.5 Flash-Lite Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Thu Oct 08 2026
Notice missing or incorrect data?

Gemini 2.5 Flash-Lite pricing

Providers

Gemini 2.5 Flash-Lite starts at $0.100 per million input tokens and $0.400 per million output tokens via Google.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
Google logoGoogle
$0.100—$0.4001.0M/65.5K
0.44
—
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Gemini 2.5 Flash-Lite context window

Input and output token limits for Gemini 2.5 Flash-Lite, plus how it ranks on long-context understanding.

InputOutput
1.0Mtokens
66Ktokens
≈ 1.6k pages of text
1.0M
8K128K1M

Try now

huggle
Gemini 2.5 Flash-Litein Huggle

Make it with
Gemini 2.5 Flash-Lite.

Gemini 2.5 Flash-Lite

Gemini 2.5 Flash-Lite latency

Gemini 2.5 Flash-Lite time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.

Gemini 2.5 Flash-Lite examples

Recent arena outputs from Gemini 2.5 Flash-Lite, picked from the highest-ranked matchups.

Gemini 2.5 Flash-Lite license

Gemini 2.5 Flash-Lite is a proprietary model available under its provider's product and API terms, has a knowledge cutoff of January 2025.

License
Creative Commons Attribution 4.0 License
Hosted access
Knowledge cutoff
January 2025

Gemini 2.5 Flash-Lite resources

Official sources for Gemini 2.5 Flash-Lite: provider documentation, official playground, paper or system card.

Gemini 2.5 Flash-Lite vs other models

The most-compared alternatives to Gemini 2.5 Flash-Lite are DeepSeek R1 Distill Llama 70B, Kimi K2 Instruct, Qwen3 VL 4B Thinking. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Gemini 2.5 Flash-Lite

Models ranked just above and below Gemini 2.5 Flash-Lite by LLM Stats score.

 

DeepSeek R1 Distill Llama 70B

Score pending
 

Kimi K2 Instruct

Score pending
 

Qwen3 VL 4B Thinking

Score pending
 

Sarvam-30B

Score pending
 

Kimi K2-Instruct-0905

Score pending
 

IBM Granite 4.2 8B

Score pending

FAQ

Common questions about Gemini 2.5 Flash-Lite.

When was Gemini 2.5 Flash-Lite released?

Gemini 2.5 Flash-Lite was released on June 17, 2025 by Google. This is the official Gemini 2.5 Flash-Lite release date tracked on LLM Stats.

How much does Gemini 2.5 Flash-Lite cost?

Gemini 2.5 Flash-Lite pricing starts at $0.10 per million input tokens and $0.40 per million output tokens via Google, the lowest price among tracked providers.

Who created Gemini 2.5 Flash-Lite?

Gemini 2.5 Flash-Lite was created by Google.

What is the license for Gemini 2.5 Flash-Lite?

Gemini 2.5 Flash-Lite is released under the Creative Commons Attribution 4.0 License license.

What is the knowledge cutoff date for Gemini 2.5 Flash-Lite?

Gemini 2.5 Flash-Lite has a knowledge cutoff of January 2025, meaning it was trained on data up to that point and may not know about events after it.

Is Gemini 2.5 Flash-Lite multimodal?

Yes, Gemini 2.5 Flash-Lite is multimodal and can accept both text and images as input.

What is Gemini 2.5 Flash-Lite latency?

Gemini 2.5 Flash-Lite p95 time to first token is 0.44 seconds via Google over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and model workloads.

Where can I use Gemini 2.5 Flash-Lite?

Gemini 2.5 Flash-Lite is available through 1 provider including Google.

Where is the Gemini 2.5 Flash-Lite paper or technical report?

Gemini 2.5 Flash-Lite has a paper or technical report available at https://arxiv.org/abs/2503.16534. Use that source for architecture, training, release and evaluation details.

What models should I compare Gemini 2.5 Flash-Lite against?

Common Gemini 2.5 Flash-Lite comparisons include Gemini 2.5 Flash-Lite vs DeepSeek R1 Distill Llama 70B, Gemini 2.5 Flash-Lite vs Kimi K2 Instruct, Gemini 2.5 Flash-Lite vs Qwen3 VL 4B Thinking. Compare them side by side for benchmark scores, pricing, context window, latency and provider availability.