The AI arena is free today

Open Playground
MetaReleased on Sep 25, 2024

Llama 3.2 3B Instruct: API Pricing, Context Window & Benchmarks

Llama 3.2 3B Instruct is a language model from Meta, released in September 2024.

Llama 3.2 3B Instruct is a large language model that supports a context length of 128K tokens and are state-of-the-art in their class for on-device use cases like summarization, instruction following, and rewriting tasks running locally at

Input
Text
Output
Text

Llama 3.2 3B Instruct benchmarks

Rankings

Quality Tracker

Llama 3.2 3B Instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Fri Aug 07 2026
Notice missing or incorrect data?

Llama 3.2 3B Instruct pricing

Providers

Llama 3.2 3B Instruct starts at $0.0100 per million input tokens and $0.0200 per million output tokens via DeepInfra.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
DeepInfra logoDeepInfra
$0.0100$0.0200128.0K/128.0K
0.24
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Llama 3.2 3B Instruct model size

Llama 3.2 3B Instruct has 3.2 billion parameters and was trained on 9 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
3.2B
9Ttokens
2804× tokens-to-params ratio
Small (3–10B)
3.2B
1B7B70B405B

Llama 3.2 3B Instruct context window

Input and output token limits for Llama 3.2 3B Instruct, plus how it ranks on long-context understanding.

InputOutput
128Ktokens
128Ktokens
192 pages of text
128K
8K128K1M

Llama 3.2 3B Instruct API

Available from the model provider

Llama 3.2 3B Instruct has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Llama 3.2 3B Instruct latency

Llama 3.2 3B Instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Llama 3.2 3B Instruct examples

Recent arena outputs from Llama 3.2 3B Instruct, picked from the highest-ranked matchups.

Llama 3.2 3B Instruct license

Llama 3.2 3B Instruct is released under the Llama 3.2 Community License license, which restricts commercial use, has 3.2B parameters.

License
Llama 3.2 Community License
Non-commercial
Parameters
3.2B

Llama 3.2 3B Instruct resources

Official sources for Llama 3.2 3B Instruct: api documentation, official playground, official launch post, model weights.

Llama 3.2 3B Instruct vs other models

The most-compared alternatives to Llama 3.2 3B Instruct are Claude 3 Haiku, Granite 3.3 8B Base, Llama 3.2 11B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Llama 3.2 3B Instruct

Models ranked just above and below Llama 3.2 3B Instruct by LLM Stats score.

 

Claude 3 Haiku

Score pending
 

Granite 3.3 8B Base

Score pending
 

Llama 3.2 11B Instruct

Score pending
 

IBM Granite 4.0 Tiny Preview

Score pending
 

MiniCPM-SALA

Score pending
 

Phi-3.5-mini-instruct

Score pending

FAQ

Common questions about Llama 3.2 3B Instruct.

When was Llama 3.2 3B Instruct released?

Llama 3.2 3B Instruct was released on September 25, 2024 by Meta. This is the official Llama 3.2 3B Instruct release date tracked on LLM Stats.

How much does Llama 3.2 3B Instruct cost?

Llama 3.2 3B Instruct pricing starts at $0.01 per million input tokens and $0.02 per million output tokens via DeepInfra, the lowest price among tracked providers.

Is Llama 3.2 3B Instruct available via API?

Yes, Llama 3.2 3B Instruct is available via API. See the official documentation for authentication and endpoint details. It is served by 1 provider tracked on LLM Stats.

How big is Llama 3.2 3B Instruct?

Llama 3.2 3B Instruct has 3.2 billion parameters. It was trained on 9.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Llama 3.2 3B Instruct?

Llama 3.2 3B Instruct was created by Meta.

What is the license for Llama 3.2 3B Instruct?

Llama 3.2 3B Instruct is released under the Llama 3.2 Community License license. This is an open-source / open-weight license that permits self-hosting.

What is Llama 3.2 3B Instruct latency?

Llama 3.2 3B Instruct p95 time to first token is 0.24 seconds via DeepInfra over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Llama 3.2 3B Instruct?

Llama 3.2 3B Instruct is available through 1 provider including DeepInfra.

Where is the Llama 3.2 3B Instruct paper or technical report?

Llama 3.2 3B Instruct has a paper or technical report available at https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/. Use that source for architecture, training, release and evaluation details.

What models should I compare Llama 3.2 3B Instruct against?

Common Llama 3.2 3B Instruct comparisons include Llama 3.2 3B Instruct vs Claude 3 Haiku, Llama 3.2 3B Instruct vs Granite 3.3 8B Base, Llama 3.2 3B Instruct vs Llama 3.2 11B Instruct. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.