The AI arena is free today

Open Superagent
MetaReleased on Sep 25, 2024

Llama 3.2 90B Instruct: Benchmarks, Pricing & Context Window

Llama 3.2 90B Instruct is a language model from Meta, released in September 2024, with multimodal input.

Llama 3.2 90B is a large multimodal language model optimized for visual recognition, image reasoning, and captioning tasks. It supports a context length of 128,000 tokens and is designed for deployment on edge and mobile devices, offering

Input
TextImage
Output
Text

Llama 3.2 90B Instruct benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for Llama 3.2 90B Instruct across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How Llama 3.2 90B Instruct holds up as conversations get longer.

Quality Tracker

Llama 3.2 90B Instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Thu Oct 08 2026
Notice missing or incorrect data?

Llama 3.2 90B Instruct pricing

Providers

Llama 3.2 90B Instruct starts at $0.350 per million input tokens and $0.400 per million output tokens via DeepInfra. See all 5 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
DeepInfra logoDeepInfra
$0.350—$0.400128.0K/128.0K
0.50
—
/
Bedrock logoBedrock
$0.720—$0.720128.0K/128.0K
0.50
—
/
Fireworks logoFireworks
$0.890—$0.890128.0K/128.0K
0.50
—
/
Together logoTogether
$1.20—$1.20128.0K/128.0K
0.50
—
/
Hyperbolic logoHyperbolic
$2.00—$2.00128.0K/128.0K
0.50
—
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Loading chart...
Loading chart...
Loading chart...

Llama 3.2 90B Instruct model size

Llama 3.2 90B Instruct has 90 billion parameters. See how it compares to other models in the same parameter range.

Parameters
90B
Very large (80–200B)
90B
1B7B70B405B

Llama 3.2 90B Instruct context window

Input and output token limits for Llama 3.2 90B Instruct, plus how it ranks on long-context understanding.

InputOutput
128Ktokens
128Ktokens
≈ 192 pages of text
128K
8K128K1M

Try now

huggle
Llama 3.2 90B Instructin Huggle

Make it with
Llama 3.2 90B Instruct.

Llama 3.2 90B Instruct

Llama 3.2 90B Instruct latency

Llama 3.2 90B Instruct time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.

Llama 3.2 90B Instruct examples

Recent arena outputs from Llama 3.2 90B Instruct, picked from the highest-ranked matchups.

Llama 3.2 90B Instruct license

Llama 3.2 90B Instruct is released under the Llama 3.2 license, which permits commercial use, has 90.0B parameters.

License
Llama 3.2
Commercial use allowed
Parameters
90.0B

Meta Llama 3.2 Community License

Llama 3.2 90B Instruct resources

Official sources for Llama 3.2 90B Instruct: provider documentation, official launch post.

Llama 3.2 90B Instruct vs other models

The most-compared alternatives to Llama 3.2 90B Instruct are Llama 3.3 70B Instruct, Qwen2-VL-72B-Instruct, Grok-2 mini. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Llama 3.2 90B Instruct

Models ranked just above and below Llama 3.2 90B Instruct by LLM Stats score.

 

Llama 3.3 70B Instruct

Score pending
 

Qwen2-VL-72B-Instruct

Score pending
 

Grok-2 mini

Score pending
 

Nova Pro

Score pending
 

Phi-4-multimodal-instruct

Score pending
 

Qwen3 VL 32B Instruct

Score pending

FAQ

Common questions about Llama 3.2 90B Instruct.

When was Llama 3.2 90B Instruct released?

Llama 3.2 90B Instruct was released on September 25, 2024 by Meta. This is the official Llama 3.2 90B Instruct release date tracked on LLM Stats.

How much does Llama 3.2 90B Instruct cost?

Llama 3.2 90B Instruct pricing starts at $0.35 per million input tokens and $0.40 per million output tokens via DeepInfra, the lowest price among tracked providers.

How big is Llama 3.2 90B Instruct?

Llama 3.2 90B Instruct has 90 billion parameters. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Llama 3.2 90B Instruct?

Llama 3.2 90B Instruct was created by Meta.

What is the license for Llama 3.2 90B Instruct?

Llama 3.2 90B Instruct is released under the Llama 3.2 license. This is an open-source / open-weight license that permits self-hosting.

Is Llama 3.2 90B Instruct multimodal?

Yes, Llama 3.2 90B Instruct is multimodal and can accept both text and images as input.

What is Llama 3.2 90B Instruct latency?

Llama 3.2 90B Instruct p95 time to first token is 0.50 seconds via DeepInfra over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and model workloads.

Where can I use Llama 3.2 90B Instruct?

Llama 3.2 90B Instruct is available through 5 providers including DeepInfra, Bedrock, Fireworks, and 2 more.

Where is the Llama 3.2 90B Instruct paper or technical report?

Llama 3.2 90B Instruct has a paper or technical report available at https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/. Use that source for architecture, training, release and evaluation details.

What models should I compare Llama 3.2 90B Instruct against?

Common Llama 3.2 90B Instruct comparisons include Llama 3.2 90B Instruct vs Llama 3.3 70B Instruct, Llama 3.2 90B Instruct vs Qwen2-VL-72B-Instruct, Llama 3.2 90B Instruct vs Grok-2 mini. Compare them side by side for benchmark scores, pricing, context window, latency and provider availability.