- Organizations
- Meta
- Llama 3.2 90B Instruct
Llama 3.2 90B Instruct: Benchmarks, Pricing & Context Window
Llama 3.2 90B Instruct is a language model from Meta, released in September 2024, with multimodal input.
Llama 3.2 90B is a large multimodal language model optimized for visual recognition, image reasoning, and captioning tasks. It supports a context length of 128,000 tokens and is designed for deployment on edge and mobile devices, offering
Llama 3.2 90B Instruct benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Llama 3.2 90B Instruct across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Llama 3.2 90B Instruct holds up as conversations get longer.
Quality Tracker
Llama 3.2 90B Instruct Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Llama 3.2 90B Instruct pricing
Providers
Llama 3.2 90B Instruct starts at $0.350 per million input tokens and $0.400 per million output tokens via DeepInfra. See all 5 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.350 | — | $0.400 | 128.0K/128.0K | 0.50 | — | / | |
| $0.720 | — | $0.720 | 128.0K/128.0K | 0.50 | — | / | |
| $0.890 | — | $0.890 | 128.0K/128.0K | 0.50 | — | / | |
| $1.20 | — | $1.20 | 128.0K/128.0K | 0.50 | — | / | |
| $2.00 | — | $2.00 | 128.0K/128.0K | 0.50 | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Llama 3.2 90B Instruct model size
Llama 3.2 90B Instruct has 90 billion parameters. See how it compares to other models in the same parameter range.
Llama 3.2 90B Instruct context window
Input and output token limits for Llama 3.2 90B Instruct, plus how it ranks on long-context understanding.
Try now
Make it with
Llama 3.2 90B Instruct.
Llama 3.2 90B Instruct latency
Llama 3.2 90B Instruct time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.
Llama 3.2 90B Instruct examples
Recent arena outputs from Llama 3.2 90B Instruct, picked from the highest-ranked matchups.
Llama 3.2 90B Instruct license
Llama 3.2 90B Instruct is released under the Llama 3.2 license, which permits commercial use, has 90.0B parameters.
- License
- Llama 3.2
- Commercial use allowed
- Parameters
- 90.0B
Meta Llama 3.2 Community License
Llama 3.2 90B Instruct resources
Official sources for Llama 3.2 90B Instruct: provider documentation, official launch post.
Llama 3.2 90B Instruct vs other models
The most-compared alternatives to Llama 3.2 90B Instruct are Llama 3.3 70B Instruct, Qwen2-VL-72B-Instruct, Grok-2 mini. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Llama 3.2 90B Instruct
Models ranked just above and below Llama 3.2 90B Instruct by LLM Stats score.
FAQ
Common questions about Llama 3.2 90B Instruct.