IBMReleased on Apr 16, 2025

Granite 3.3 8B Instruct: API Pricing, Context Window & Benchmarks

Granite 3.3 8B Instruct is a language model from IBM, released in April 2025, with multimodal input.

Granite 3.3 models feature enhanced reasoning capabilities and support for Fill-in-the-Middle (FIM) code completion. They are built on a foundation of open-source instruction datasets with permissive licenses, alongside internally curated

Input
Text
Output
Text

Granite 3.3 8B Instruct benchmarks

Rankings

Quality Tracker

Granite 3.3 8B Instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Tue Jul 21 2026
Notice missing or incorrect data?

Granite 3.3 8B Instruct pricing

Providers

Granite 3.3 8B Instruct starts at $0.500 per million input tokens and $0.500 per million output tokens via Replicate.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
Replicate logoReplicate
$0.500$0.500128.0K/8.2K
/0.30
50/
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Granite 3.3 8B Instruct model size

Granite 3.3 8B Instruct has 8 billion parameters. See how it compares to other models in the same parameter range.

Parameters
8B
Small (3–10B)
8B
1B7B70B405B

Granite 3.3 8B Instruct context window

Input and output token limits for Granite 3.3 8B Instruct, plus how it ranks on long-context understanding.

InputOutput
128Ktokens
8Ktokens
192 pages of text
128K
8K128K1M

Granite 3.3 8B Instruct API

Available from the model provider

Granite 3.3 8B Instruct has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Granite 3.3 8B Instruct latency

Granite 3.3 8B Instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Granite 3.3 8B Instruct examples

Recent arena outputs from Granite 3.3 8B Instruct, picked from the highest-ranked matchups.

Granite 3.3 8B Instruct license

Granite 3.3 8B Instruct is released under the Apache 2.0 license, which permits commercial use, has 8.0B parameters, has a knowledge cutoff of April 2024.

License
Apache 2.0
Commercial use allowed
Parameters
8.0B
Knowledge cutoff
April 2024

Apache License 2.0 - allows commercial use

Granite 3.3 8B Instruct resources

Official sources for Granite 3.3 8B Instruct: api documentation, official playground, official launch post, source repository.

Granite 3.3 8B Instruct vs other models

The most-compared alternatives to Granite 3.3 8B Instruct are Phi 4 Reasoning Plus, GPT-4o, Qwen3 30B A3B. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Granite 3.3 8B Instruct

Models ranked just above and below Granite 3.3 8B Instruct by LLM Stats score.

 

Phi 4 Reasoning Plus

Score pending
 

GPT-4o

Score pending
 

Qwen3 30B A3B

Score pending
 

Nova Pro

Score pending
 

Phi 4 Mini

Score pending
 

Qwen3 32B

Score pending

FAQ

Common questions about Granite 3.3 8B Instruct.

When was Granite 3.3 8B Instruct released?

Granite 3.3 8B Instruct was released on April 16, 2025 by IBM. This is the official Granite 3.3 8B Instruct release date tracked on LLM Stats.

How much does Granite 3.3 8B Instruct cost?

Granite 3.3 8B Instruct pricing starts at $0.50 per million input tokens and $0.50 per million output tokens via Replicate, the lowest price among tracked providers.

Is Granite 3.3 8B Instruct available via API?

Yes, Granite 3.3 8B Instruct is available via API. See the official documentation for authentication and endpoint details. It is served by 1 provider tracked on LLM Stats.

How big is Granite 3.3 8B Instruct?

Granite 3.3 8B Instruct has 8 billion parameters. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Granite 3.3 8B Instruct?

Granite 3.3 8B Instruct was created by IBM.

What is the license for Granite 3.3 8B Instruct?

Granite 3.3 8B Instruct is released under the Apache 2.0 license. This is an open-source / open-weight license that permits self-hosting.

What is the knowledge cutoff date for Granite 3.3 8B Instruct?

Granite 3.3 8B Instruct has a knowledge cutoff of April 2024, meaning it was trained on data up to that point and may not know about events after it.

Is Granite 3.3 8B Instruct multimodal?

Yes, Granite 3.3 8B Instruct is multimodal and can accept both text and images as input.

What is Granite 3.3 8B Instruct latency?

Granite 3.3 8B Instruct p95 time to first token is 0.30 seconds via Replicate over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Granite 3.3 8B Instruct?

Granite 3.3 8B Instruct is available through 1 provider including Replicate.

Where is the Granite 3.3 8B Instruct paper or technical report?

Granite 3.3 8B Instruct has a paper or technical report available at https://huggingface.co/ibm-granite/granite-3.3-8b-instruct. Use that source for architecture, training, release and evaluation details.

What models should I compare Granite 3.3 8B Instruct against?

Common Granite 3.3 8B Instruct comparisons include Granite 3.3 8B Instruct vs Phi 4 Reasoning Plus, Granite 3.3 8B Instruct vs GPT-4o, Granite 3.3 8B Instruct vs Qwen3 30B A3B. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.