IBMReleased on Apr 16, 2025

Granite 3.3 8B Base: API Pricing, Context Window & Benchmarks

Granite 3.3 8B Base is a language model from IBM, released in April 2025, with multimodal input.

Granite-3.3-8B-Base is a decoder-only language model with a 128K token context window. It improves upon Granite-3.1-8B-Base by adding support for Fill-in-the-Middle (FIM) using specialized tokens, enabling the model to generate content

Granite 3.3 8B Base benchmarks

Rankings

Quality Tracker

Granite 3.3 8B Base Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Sun Jul 26 2026
Notice missing or incorrect data?

Granite 3.3 8B Base model size

Granite 3.3 8B Base has 8.2 billion parameters. See how it compares to other models in the same parameter range.

Parameters
8.2B
Small (3–10B)
8.2B
1B7B70B405B

Granite 3.3 8B Base API

Available from the model provider

Granite 3.3 8B Base has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Granite 3.3 8B Base latency

Granite 3.3 8B Base time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Granite 3.3 8B Base examples

Recent arena outputs from Granite 3.3 8B Base, picked from the highest-ranked matchups.

Granite 3.3 8B Base license

Granite 3.3 8B Base is released under the Apache 2.0 license, which permits commercial use, has 8.2B parameters, has a knowledge cutoff of April 2024.

License
Apache 2.0
Commercial use allowed
Parameters
8.2B
Knowledge cutoff
April 2024

Apache License 2.0 - allows commercial use

Granite 3.3 8B Base resources

Official sources for Granite 3.3 8B Base: api documentation, official playground, official launch post, source repository.

Granite 3.3 8B Base vs other models

The most-compared alternatives to Granite 3.3 8B Base are Phi 4 Reasoning Plus, GPT-4o, Qwen3 30B A3B. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Granite 3.3 8B Base

Models ranked just above and below Granite 3.3 8B Base by LLM Stats score.

 

Phi 4 Reasoning Plus

Score pending
 

GPT-4o

Score pending
 

Qwen3 30B A3B

Score pending
 

DeepSeek R1 Distill Qwen 14B

Score pending
 

Granite 3.3 8B Instruct

Score pending
 

Qwen2.5 VL 32B Instruct

Score pending

FAQ

Common questions about Granite 3.3 8B Base.

When was Granite 3.3 8B Base released?

Granite 3.3 8B Base was released on April 16, 2025 by IBM. This is the official Granite 3.3 8B Base release date tracked on LLM Stats.

Is Granite 3.3 8B Base available via API?

Yes, Granite 3.3 8B Base is available via API. See the official documentation for authentication and endpoint details.

How big is Granite 3.3 8B Base?

Granite 3.3 8B Base has 8.2 billion parameters. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Granite 3.3 8B Base?

Granite 3.3 8B Base was created by IBM.

What is the license for Granite 3.3 8B Base?

Granite 3.3 8B Base is released under the Apache 2.0 license. This is an open-source / open-weight license that permits self-hosting.

What is the knowledge cutoff date for Granite 3.3 8B Base?

Granite 3.3 8B Base has a knowledge cutoff of April 2024, meaning it was trained on data up to that point and may not know about events after it.

Is Granite 3.3 8B Base multimodal?

Yes, Granite 3.3 8B Base is multimodal and can accept both text and images as input.

Where is the Granite 3.3 8B Base paper or technical report?

Granite 3.3 8B Base has a paper or technical report available at https://huggingface.co/ibm-granite/granite-3.3-8b-base. Use that source for architecture, training, release and evaluation details.

What models should I compare Granite 3.3 8B Base against?

Common Granite 3.3 8B Base comparisons include Granite 3.3 8B Base vs Phi 4 Reasoning Plus, Granite 3.3 8B Base vs GPT-4o, Granite 3.3 8B Base vs Qwen3 30B A3B. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.