The AI arena is free today

Open Superagent
GoogleReleased on May 1, 2024

Gemini 1.5 Flash: API Pricing, Context Window & Benchmarks

Gemini 1.5 Flash is a language model from Google, released in May 2024, with multimodal input.

Gemini 1.5 Flash is a fast and versatile multimodal model for scaling across diverse tasks. It supports audio, images, video, and text input, and produces text output. The model is optimized for generating code, extracting data, editing

Input
TextImage
Output
Text

Gemini 1.5 Flash benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

How Gemini 1.5 Flash performs across real-world prompt categories.

Performance by conversation depth

How Gemini 1.5 Flash holds up as conversations get longer.

Quality Tracker

Gemini 1.5 Flash Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Mon Aug 24 2026
Notice missing or incorrect data?

Gemini 1.5 Flash pricing

Providers

Gemini 1.5 Flash starts at $0.150 per million input tokens and $0.600 per million output tokens via Google.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
Google logoGoogle
$0.150$0.6001.0M/8.2K
0.30
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Gemini 1.5 Flash context window

Input and output token limits for Gemini 1.5 Flash, plus how it ranks on long-context understanding.

InputOutput
1.0Mtokens
8Ktokens
1.6k pages of text
1.0M
8K128K1M

Gemini 1.5 Flash API

Available from the model provider

Gemini 1.5 Flash has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Gemini 1.5 Flash latency

Gemini 1.5 Flash time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Gemini 1.5 Flash examples

Recent arena outputs from Gemini 1.5 Flash, picked from the highest-ranked matchups.

Gemini 1.5 Flash license

Gemini 1.5 Flash is a proprietary model available under its provider's product and API terms, has a knowledge cutoff of November 2023.

License
Proprietary
Hosted access
Knowledge cutoff
November 2023

Proprietary license - usage restrictions apply

Gemini 1.5 Flash resources

Official sources for Gemini 1.5 Flash: api documentation, official playground, paper or system card, official launch post.

Gemini 1.5 Flash vs other models

The most-compared alternatives to Gemini 1.5 Flash are Llama 3.3 70B Instruct, Llama 3.1 405B Instruct, Grok-2 mini. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Gemini 1.5 Flash

Models ranked just above and below Gemini 1.5 Flash by LLM Stats score.

 

Llama 3.3 70B Instruct

Score pending
 

Llama 3.1 405B Instruct

Score pending
 

Grok-2 mini

Score pending
 

Claude 3 Sonnet

Score pending
 

Claude 3.5 Sonnet

Score pending
 

Nova Pro

Score pending

FAQ

Common questions about Gemini 1.5 Flash.

When was Gemini 1.5 Flash released?

Gemini 1.5 Flash was released on May 1, 2024 by Google. This is the official Gemini 1.5 Flash release date tracked on LLM Stats.

How much does Gemini 1.5 Flash cost?

Gemini 1.5 Flash pricing starts at $0.15 per million input tokens and $0.60 per million output tokens via Google, the lowest price among tracked providers.

Is Gemini 1.5 Flash available via API?

Yes, Gemini 1.5 Flash is available via API. See the official documentation for authentication and endpoint details. It is served by 1 provider tracked on LLM Stats.

Who created Gemini 1.5 Flash?

Gemini 1.5 Flash was created by Google.

What is the license for Gemini 1.5 Flash?

Gemini 1.5 Flash is released under the Proprietary license.

What is the knowledge cutoff date for Gemini 1.5 Flash?

Gemini 1.5 Flash has a knowledge cutoff of November 2023, meaning it was trained on data up to that point and may not know about events after it.

Is Gemini 1.5 Flash multimodal?

Yes, Gemini 1.5 Flash is multimodal and can accept both text and images as input.

What is Gemini 1.5 Flash latency?

Gemini 1.5 Flash p95 time to first token is 0.30 seconds via Google over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Gemini 1.5 Flash?

Gemini 1.5 Flash is available through 1 provider including Google.

Where is the Gemini 1.5 Flash paper or technical report?

Gemini 1.5 Flash has a paper or technical report available at https://arxiv.org/pdf/2403.05530. Use that source for architecture, training, release and evaluation details.

What models should I compare Gemini 1.5 Flash against?

Common Gemini 1.5 Flash comparisons include Gemini 1.5 Flash vs Llama 3.3 70B Instruct, Gemini 1.5 Flash vs Llama 3.1 405B Instruct, Gemini 1.5 Flash vs Grok-2 mini. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.