IBMReleased on May 2, 2025

IBM Granite 4.0 Tiny Preview: API Pricing, Context Window & Benchmarks

IBM Granite 4.0 Tiny Preview is a language model from IBM, released in May 2025.

A preliminary version of the smallest model in the upcoming Granite 4.0 family, released May 2025. It utilizes a novel hybrid Mamba-2/Transformer, fine-grained mixture of experts (MoE) architecture (7B total parameters, 1B active at

IBM Granite 4.0 Tiny Preview benchmarks

Rankings

Quality Tracker

IBM Granite 4.0 Tiny Preview Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Mon Jul 20 2026
Notice missing or incorrect data?

IBM Granite 4.0 Tiny Preview model size

IBM Granite 4.0 Tiny Preview has 7 billion parameters and was trained on 2.5 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
7B
2.5Ttokens
357× tokens-to-params ratio
Small (3–10B)
7B
1B7B70B405B

IBM Granite 4.0 Tiny Preview API

Available from the model provider

IBM Granite 4.0 Tiny Preview has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

IBM Granite 4.0 Tiny Preview latency

IBM Granite 4.0 Tiny Preview time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

IBM Granite 4.0 Tiny Preview examples

Recent arena outputs from IBM Granite 4.0 Tiny Preview, picked from the highest-ranked matchups.

IBM Granite 4.0 Tiny Preview license

IBM Granite 4.0 Tiny Preview is released under the Apache 2.0 license, which permits commercial use, has 7.0B parameters.

License
Apache 2.0
Commercial use allowed
Parameters
7.0B

Apache License 2.0 - allows commercial use

IBM Granite 4.0 Tiny Preview resources

Official sources for IBM Granite 4.0 Tiny Preview: api documentation, official playground, official launch post, model weights.

IBM Granite 4.0 Tiny Preview vs other models

The most-compared alternatives to IBM Granite 4.0 Tiny Preview are Command R+, Nova Micro, Phi 4. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like IBM Granite 4.0 Tiny Preview

Models ranked just above and below IBM Granite 4.0 Tiny Preview by LLM Stats score.

 

Command R+

Score pending
 

Nova Micro

Score pending
 

Phi 4

Score pending
 

Gemma 2 27B

Score pending
 

Qwen2.5 14B Instruct

Score pending
 

Granite 3.3 8B Base

Score pending

FAQ

Common questions about IBM Granite 4.0 Tiny Preview.

When was IBM Granite 4.0 Tiny Preview released?

IBM Granite 4.0 Tiny Preview was released on May 2, 2025 by IBM. This is the official IBM Granite 4.0 Tiny Preview release date tracked on LLM Stats.

Is IBM Granite 4.0 Tiny Preview available via API?

Yes, IBM Granite 4.0 Tiny Preview is available via API. See the official documentation for authentication and endpoint details.

How big is IBM Granite 4.0 Tiny Preview?

IBM Granite 4.0 Tiny Preview has 7 billion parameters. It was trained on 2.5 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created IBM Granite 4.0 Tiny Preview?

IBM Granite 4.0 Tiny Preview was created by IBM.

What is the license for IBM Granite 4.0 Tiny Preview?

IBM Granite 4.0 Tiny Preview is released under the Apache 2.0 license. This is an open-source / open-weight license that permits self-hosting.

Where is the IBM Granite 4.0 Tiny Preview paper or technical report?

IBM Granite 4.0 Tiny Preview has a paper or technical report available at https://www.ibm.com/new/announcements/ibm-granite-4-0-tiny-preview-sneak-peek. Use that source for architecture, training, release and evaluation details.

What models should I compare IBM Granite 4.0 Tiny Preview against?

Common IBM Granite 4.0 Tiny Preview comparisons include IBM Granite 4.0 Tiny Preview vs Command R+, IBM Granite 4.0 Tiny Preview vs Nova Micro, IBM Granite 4.0 Tiny Preview vs Phi 4. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.