GoogleReleased on Feb 5, 2025

Gemini 2.0 Flash-Lite: API Pricing, Context Window & Benchmarks

Gemini 2.0 Flash-Lite is a language model from Google, released in February 2025, with multimodal input.

A Gemini 2.0 Flash model optimized for cost efficiency and low latency

Input
TextImage
Output
Text

Gemini 2.0 Flash-Lite benchmarks

Rankings

Quality Tracker

Gemini 2.0 Flash-Lite Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Wed Jul 29 2026
Notice missing or incorrect data?

Gemini 2.0 Flash-Lite pricing

Providers

Gemini 2.0 Flash-Lite starts at $0.0700 per million input tokens and $0.300 per million output tokens via Google.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
Google logoGoogle
$0.0700$0.3001.0M/8.2K
/0.70
85/
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Gemini 2.0 Flash-Lite context window

Input and output token limits for Gemini 2.0 Flash-Lite, plus how it ranks on long-context understanding.

InputOutput
1.0Mtokens
8Ktokens
1.6k pages of text
1.0M
8K128K1M

Gemini 2.0 Flash-Lite API

Available from the model provider

Gemini 2.0 Flash-Lite has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Gemini 2.0 Flash-Lite latency

Gemini 2.0 Flash-Lite time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Gemini 2.0 Flash-Lite examples

Recent arena outputs from Gemini 2.0 Flash-Lite, picked from the highest-ranked matchups.

Gemini 2.0 Flash-Lite license

Gemini 2.0 Flash-Lite is a proprietary model available under its provider's product and API terms, has a knowledge cutoff of June 2024.

License
Proprietary
Hosted access
Knowledge cutoff
June 2024

Proprietary license - usage restrictions apply

Gemini 2.0 Flash-Lite resources

Official sources for Gemini 2.0 Flash-Lite: api documentation, official playground, official launch post.

Gemini 2.0 Flash-Lite vs other models

The most-compared alternatives to Gemini 2.0 Flash-Lite are MiMo-V2.5-Pro, GPT-4o, Grok-2. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Gemini 2.0 Flash-Lite

Models ranked just above and below Gemini 2.0 Flash-Lite by LLM Stats score.

 

MiMo-V2.5-Pro

Score pending
 

GPT-4o

Score pending
 

Grok-2

Score pending
 

Gemini 1.5 Pro

Score pending
 

Grok-2 mini

Score pending
 

Claude 3.5 Sonnet

Score pending

FAQ

Common questions about Gemini 2.0 Flash-Lite.

When was Gemini 2.0 Flash-Lite released?

Gemini 2.0 Flash-Lite was released on February 5, 2025 by Google. This is the official Gemini 2.0 Flash-Lite release date tracked on LLM Stats.

How much does Gemini 2.0 Flash-Lite cost?

Gemini 2.0 Flash-Lite pricing starts at $0.07 per million input tokens and $0.30 per million output tokens via Google, the lowest price among tracked providers.

Is Gemini 2.0 Flash-Lite available via API?

Yes, Gemini 2.0 Flash-Lite is available via API. See the official documentation for authentication and endpoint details. It is served by 1 provider tracked on LLM Stats.

Who created Gemini 2.0 Flash-Lite?

Gemini 2.0 Flash-Lite was created by Google.

What is the license for Gemini 2.0 Flash-Lite?

Gemini 2.0 Flash-Lite is released under the Proprietary license.

What is the knowledge cutoff date for Gemini 2.0 Flash-Lite?

Gemini 2.0 Flash-Lite has a knowledge cutoff of June 2024, meaning it was trained on data up to that point and may not know about events after it.

Is Gemini 2.0 Flash-Lite multimodal?

Yes, Gemini 2.0 Flash-Lite is multimodal and can accept both text and images as input.

What is Gemini 2.0 Flash-Lite latency?

Gemini 2.0 Flash-Lite p95 time to first token is 0.70 seconds via Google over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Gemini 2.0 Flash-Lite?

Gemini 2.0 Flash-Lite is available through 1 provider including Google.

Where is the Gemini 2.0 Flash-Lite paper or technical report?

Gemini 2.0 Flash-Lite has a paper or technical report available at https://developers.googleblog.com/en/gemini-2-family-expands. Use that source for architecture, training, release and evaluation details.

What models should I compare Gemini 2.0 Flash-Lite against?

Common Gemini 2.0 Flash-Lite comparisons include Gemini 2.0 Flash-Lite vs MiMo-V2.5-Pro, Gemini 2.0 Flash-Lite vs GPT-4o, Gemini 2.0 Flash-Lite vs Grok-2. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.