GoogleReleased on Mar 15, 2024

Gemini 1.5 Flash 8B: API Pricing, Context Window & Benchmarks

Gemini 1.5 Flash 8B is a language model from Google, released in March 2024, with multimodal input.

A multimodal model capable of processing audio, images, video, and text with high efficiency. Features JSON mode, function calling, code execution, and system instructions support. Optimized for fast inference with 8B parameters.

Input
TextImage
Output
Text

Gemini 1.5 Flash 8B benchmarks

Rankings

Quality Tracker

Gemini 1.5 Flash 8B Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Sun Aug 02 2026
Notice missing or incorrect data?

Gemini 1.5 Flash 8B pricing

Providers

Gemini 1.5 Flash 8B starts at $0.0700 per million input tokens and $0.300 per million output tokens via Google.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
Google logoGoogle
$0.0700$0.3001.0M/8.2K
/0.30
150/
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Gemini 1.5 Flash 8B context window

Input and output token limits for Gemini 1.5 Flash 8B, plus how it ranks on long-context understanding.

InputOutput
1.0Mtokens
8Ktokens
1.6k pages of text
1.0M
8K128K1M

Gemini 1.5 Flash 8B API

Available from the model provider

Gemini 1.5 Flash 8B has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Gemini 1.5 Flash 8B latency

Gemini 1.5 Flash 8B time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Gemini 1.5 Flash 8B examples

Recent arena outputs from Gemini 1.5 Flash 8B, picked from the highest-ranked matchups.

Gemini 1.5 Flash 8B license

Gemini 1.5 Flash 8B is a proprietary model available under its provider's product and API terms, has 8.0B parameters, has a knowledge cutoff of October 2024.

License
Proprietary
Hosted access
Parameters
8.0B
Knowledge cutoff
October 2024

Proprietary license - usage restrictions apply

Gemini 1.5 Flash 8B resources

Official sources for Gemini 1.5 Flash 8B: api documentation, official playground, source repository.

Gemini 1.5 Flash 8B vs other models

The most-compared alternatives to Gemini 1.5 Flash 8B are Claude 3 Sonnet, Phi-4-multimodal-instruct, Grok-1.5V. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Gemini 1.5 Flash 8B

Models ranked just above and below Gemini 1.5 Flash 8B by LLM Stats score.

 

Claude 3 Sonnet

Score pending
 

Phi-4-multimodal-instruct

Score pending
 

Grok-1.5V

Score pending
 

Nova Micro

Score pending
 

Jamba 1.5 Large

Score pending
 

Qwen2 72B Instruct

Score pending

FAQ

Common questions about Gemini 1.5 Flash 8B.

When was Gemini 1.5 Flash 8B released?

Gemini 1.5 Flash 8B was released on March 15, 2024 by Google. This is the official Gemini 1.5 Flash 8B release date tracked on LLM Stats.

How much does Gemini 1.5 Flash 8B cost?

Gemini 1.5 Flash 8B pricing starts at $0.07 per million input tokens and $0.30 per million output tokens via Google, the lowest price among tracked providers.

Is Gemini 1.5 Flash 8B available via API?

Yes, Gemini 1.5 Flash 8B is available via API. See the official documentation for authentication and endpoint details. It is served by 1 provider tracked on LLM Stats.

How big is Gemini 1.5 Flash 8B?

Gemini 1.5 Flash 8B has 8 billion parameters.

Who created Gemini 1.5 Flash 8B?

Gemini 1.5 Flash 8B was created by Google.

What is the license for Gemini 1.5 Flash 8B?

Gemini 1.5 Flash 8B is released under the Proprietary license.

What is the knowledge cutoff date for Gemini 1.5 Flash 8B?

Gemini 1.5 Flash 8B has a knowledge cutoff of October 2024, meaning it was trained on data up to that point and may not know about events after it.

Is Gemini 1.5 Flash 8B multimodal?

Yes, Gemini 1.5 Flash 8B is multimodal and can accept both text and images as input.

What is Gemini 1.5 Flash 8B latency?

Gemini 1.5 Flash 8B p95 time to first token is 0.30 seconds via Google over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Gemini 1.5 Flash 8B?

Gemini 1.5 Flash 8B is available through 1 provider including Google.

What models should I compare Gemini 1.5 Flash 8B against?

Common Gemini 1.5 Flash 8B comparisons include Gemini 1.5 Flash 8B vs Claude 3 Sonnet, Gemini 1.5 Flash 8B vs Phi-4-multimodal-instruct, Gemini 1.5 Flash 8B vs Grok-1.5V. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.