- Organizations
- MiniMax
- MiniMax M3
MiniMax M3: API Pricing, Context Window & Benchmarks
MiniMax M3 is a language model from MiniMax, released in June 2026, with multimodal input, a 1M-token context window, and pricing from $0.300/M input and $1.20/M output.
MiniMax M3 is the first open-weight model to combine three frontier capabilities: top-tier coding and agentic performance, a 1M-token context window, and native multimodality. It is powered by MiniMax Sparse Attention (MSA), a new sparse
MiniMax M3 benchmarks
Rankings
Quality Tracker
MiniMax M3 Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
MiniMax M3 pricing
Providers
MiniMax M3 starts at $0.300 per million input tokens and $1.20 per million output tokens via Fireworks. See all 4 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Output $/M | Context in / out | TTFT p50 / p95 s | Output avg / p5 c/s | Success 7d | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.300 | $1.20 | 512.0K/131.1K | 0.91/2.37 | 300/10 | 62.07%(29) | / | |
| $0.300 | $1.20 | 1.0M/131.1K | 4.43/7.72 | 150/7 | 81.82%(11) | / | |
| $0.300 | $1.20 | 1.0M/131.1K | 1.14/1.45 | 44/44 | 50.00%(4) | / | |
| $0.600 | $2.40 | 1.0M/1.0M | 1.35/3.59 | 200/9 | 100.00%(30) | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.
MiniMax M3 context window
Input and output token limits for MiniMax M3, plus how it ranks on long-context understanding.
MiniMax M3 API
Available from the model provider
MiniMax M3 is available from Fireworks, Novita, Together and 1 more. It is not currently routed through the LLM Stats gateway.
Read the official API documentationMiniMax M3 latency
MiniMax M3 time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.
Provider operational metrics
Time to first token, output throughput, and failed-request rate from live API traffic
MiniMax M3 examples
Recent arena outputs from MiniMax M3, picked from the highest-ranked matchups.
MiniMax M3 license
MiniMax M3 is released under the MIT license, which permits commercial use.
- License
- MIT
- Commercial use allowed
MIT License - allows commercial use
MiniMax M3 resources
Official sources for MiniMax M3: api documentation, official playground, official launch post.
MiniMax M3 vs other models
The most-compared alternatives to MiniMax M3 are Claude Opus 4.6, Qwen3.7 Max, Muse Spark 1.1. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like MiniMax M3
Models ranked just above and below MiniMax M3 by LLM Stats score.
FAQ
Common questions about MiniMax M3.