GoogleReleased on Mar 12, 2025

Gemma 3 4B: API Pricing, Context Window & Benchmarks

Gemma 3 4B is a language model from Google, released in March 2025, with multimodal input.

Gemma 3 4B is a 4-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for question answering,

Input
TextImage
Output
Text

Gemma 3 4B benchmarks

Rankings

Quality Tracker

Gemma 3 4B Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Mon Jul 20 2026
Notice missing or incorrect data?

Gemma 3 4B pricing

Providers

Gemma 3 4B starts at $0.0200 per million input tokens and $0.0400 per million output tokens via DeepInfra.

ProviderInput $/MOutput $/MWorkload 1M + 100KContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
DeepInfra logoDeepInfra
$0.0200$0.0400$0.0240131.1K/131.1K
/0.20
33/
/

Workload cost uses 1M input tokens plus 100K output tokens at published list prices without assuming a cache hit. Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Gemma 3 4B model size

Gemma 3 4B has 4 billion parameters and was trained on 4 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
4B
4Ttokens
1000× tokens-to-params ratio
Small (3–10B)
4B
1B7B70B405B

Gemma 3 4B context window

Input and output token limits for Gemma 3 4B, plus how it ranks on long-context understanding.

InputOutput
131Ktokens
131Ktokens
197 pages of text
131K
8K128K1M

Gemma 3 4B latency

Gemma 3 4B time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Gemma 3 4B examples

Recent arena outputs from Gemma 3 4B, picked from the highest-ranked matchups.

Gemma 3 4B license

Gemma 3 4B is released under the Gemma license, which permits commercial use, has 4.0B parameters, has a knowledge cutoff of August 2024.

License
Gemma
Commercial use allowed
Parameters
4.0B
Knowledge cutoff
August 2024

Google Gemma Terms of Use

Gemma 3 4B resources

Official sources for Gemma 3 4B: paper or system card, official launch post, model weights.

Gemma 3 4B vs other models

The most-compared alternatives to Gemma 3 4B are GPT-4o, Grok-2, DeepSeek-V2.5. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Gemma 3 4B

Models ranked just above and below Gemma 3 4B by LLM Stats score.

 

GPT-4o

Score pending
 

Grok-2

Score pending
 

DeepSeek-V2.5

Score pending
 

Llama 3.2 11B Instruct

Score pending
 

Phi-3.5-MoE-instruct

Score pending
 

Qwen2.5-Omni-7B

Score pending

FAQ

Common questions about Gemma 3 4B.

When was Gemma 3 4B released?

Gemma 3 4B was released on March 12, 2025 by Google. This is the official Gemma 3 4B release date tracked on LLM Stats.

How much does Gemma 3 4B cost?

Gemma 3 4B pricing starts at $0.02 per million input tokens and $0.04 per million output tokens via DeepInfra, the lowest price among tracked providers.

How big is Gemma 3 4B?

Gemma 3 4B has 4 billion parameters. It was trained on 4.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Gemma 3 4B?

Gemma 3 4B was created by Google.

What is the license for Gemma 3 4B?

Gemma 3 4B is released under the Gemma license. This is an open-source / open-weight license that permits self-hosting.

What is the knowledge cutoff date for Gemma 3 4B?

Gemma 3 4B has a knowledge cutoff of August 2024, meaning it was trained on data up to that point and may not know about events after it.

Is Gemma 3 4B multimodal?

Yes, Gemma 3 4B is multimodal and can accept both text and images as input.

What is Gemma 3 4B latency?

Gemma 3 4B p95 time to first token is 0.20 seconds via DeepInfra over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Gemma 3 4B?

Gemma 3 4B is available through 1 provider including DeepInfra.

Where is the Gemma 3 4B paper or technical report?

Gemma 3 4B has a paper or technical report available at https://storage.googleapis.com/deepmind-media/gemma/Gemma3Report.pdf. Use that source for architecture, training, release and evaluation details.

What models should I compare Gemma 3 4B against?

Common Gemma 3 4B comparisons include Gemma 3 4B vs GPT-4o, Gemma 3 4B vs Grok-2, Gemma 3 4B vs DeepSeek-V2.5. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.