GoogleReleased on Jun 26, 2025

Gemma 3n E4B Instructed: API Pricing, Context Window & Benchmarks

Gemma 3n E4B Instructed is a language model from Google, released in June 2025, with multimodal input.

Gemma 3n is a multimodal model designed to run locally on hardware, supporting image, text, audio, and video inputs. It features a language decoder, audio encoder, and vision encoder, and is available in two sizes: E2B and E4B. The model

Input
TextImage
Output
Text

Gemma 3n E4B Instructed benchmarks

Rankings

Quality Tracker

Gemma 3n E4B Instructed Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Tue Jul 21 2026
Notice missing or incorrect data?

Gemma 3n E4B Instructed pricing

Providers

Gemma 3n E4B Instructed starts at $20.00 per million input tokens and $40.00 per million output tokens via Together.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
Together logoTogether
$20.00$40.0032.0K/32.0K
/0.43
42/
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Gemma 3n E4B Instructed context window

Input and output token limits for Gemma 3n E4B Instructed, plus how it ranks on long-context understanding.

InputOutput
32Ktokens
32Ktokens
48 pages of text
32K
8K128K1M

Gemma 3n E4B Instructed API

Available from the model provider

Gemma 3n E4B Instructed has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Gemma 3n E4B Instructed latency

Gemma 3n E4B Instructed time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Gemma 3n E4B Instructed examples

Recent arena outputs from Gemma 3n E4B Instructed, picked from the highest-ranked matchups.

Gemma 3n E4B Instructed license

Gemma 3n E4B Instructed is a proprietary model available under its provider's product and API terms, has 8.0B parameters, has a knowledge cutoff of June 2024.

License
Proprietary
Hosted access
Parameters
8.0B
Knowledge cutoff
June 2024

Proprietary license - usage restrictions apply

Gemma 3n E4B Instructed resources

Official sources for Gemma 3n E4B Instructed: api documentation, official playground, official launch post, model weights.

Gemma 3n E4B Instructed vs other models

The most-compared alternatives to Gemma 3n E4B Instructed are Granite 3.3 8B Instruct, Phi 4 Mini, Granite 3.3 8B Base. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Gemma 3n E4B Instructed

Models ranked just above and below Gemma 3n E4B Instructed by LLM Stats score.

 

Granite 3.3 8B Instruct

Score pending
 

Phi 4 Mini

Score pending
 

Granite 3.3 8B Base

Score pending
 

Grok-1.5

Score pending
 

Qwen2.5-Coder 32B Instruct

Score pending
 

Ministral 8B Instruct

Score pending

FAQ

Common questions about Gemma 3n E4B Instructed.

When was Gemma 3n E4B Instructed released?

Gemma 3n E4B Instructed was released on June 26, 2025 by Google. This is the official Gemma 3n E4B Instructed release date tracked on LLM Stats.

How much does Gemma 3n E4B Instructed cost?

Gemma 3n E4B Instructed pricing starts at $20.00 per million input tokens and $40.00 per million output tokens via Together, the lowest price among tracked providers.

Is Gemma 3n E4B Instructed available via API?

Yes, Gemma 3n E4B Instructed is available via API. See the official documentation for authentication and endpoint details. It is served by 1 provider tracked on LLM Stats.

How big is Gemma 3n E4B Instructed?

Gemma 3n E4B Instructed has 8 billion parameters. It was trained on 11.0 trillion tokens.

Who created Gemma 3n E4B Instructed?

Gemma 3n E4B Instructed was created by Google.

What is the license for Gemma 3n E4B Instructed?

Gemma 3n E4B Instructed is released under the Proprietary license.

What is the knowledge cutoff date for Gemma 3n E4B Instructed?

Gemma 3n E4B Instructed has a knowledge cutoff of June 2024, meaning it was trained on data up to that point and may not know about events after it.

Is Gemma 3n E4B Instructed multimodal?

Yes, Gemma 3n E4B Instructed is multimodal and can accept both text and images as input.

What is Gemma 3n E4B Instructed latency?

Gemma 3n E4B Instructed p95 time to first token is 0.43 seconds via Together over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Gemma 3n E4B Instructed?

Gemma 3n E4B Instructed is available through 1 provider including Together.

Where is the Gemma 3n E4B Instructed paper or technical report?

Gemma 3n E4B Instructed has a paper or technical report available at https://ai.google.dev/gemma/docs/gemma-3n. Use that source for architecture, training, release and evaluation details.

What models should I compare Gemma 3n E4B Instructed against?

Common Gemma 3n E4B Instructed comparisons include Gemma 3n E4B Instructed vs Granite 3.3 8B Instruct, Gemma 3n E4B Instructed vs Phi 4 Mini, Gemma 3n E4B Instructed vs Granite 3.3 8B Base. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.