MiniMaxReleased on Jun 1, 2026

MiniMax M3: API Pricing, Context Window & Benchmarks

MiniMax M3 is a language model from MiniMax, released in June 2026, with multimodal input, a 1M-token context window, and pricing from $0.300/M input and $1.20/M output.

MiniMax M3 is the first open-weight model to combine three frontier capabilities: top-tier coding and agentic performance, a 1M-token context window, and native multimodality. It is powered by MiniMax Sparse Attention (MSA), a new sparse

Input
TextImageVideo
Output
Text

MiniMax M3 benchmarks

Rankings

Quality Tracker

MiniMax M3 Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Tue Jul 21 2026
Notice missing or incorrect data?

MiniMax M3 pricing

Providers

MiniMax M3 starts at $0.300 per million input tokens and $1.20 per million output tokens via Fireworks. See all 4 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
Fireworks logoFireworks
$0.300$1.20512.0K/131.1K
0.91/2.37
300/10
62.07%(29)
/
Novita logoNovita
$0.300$1.201.0M/131.1K
4.43/7.72
150/7
81.82%(11)
/
Together logoTogether
$0.300$1.201.0M/131.1K
1.14/1.45
44/44
50.00%(4)
/
MiniMax logoMiniMax
$0.600$2.401.0M/1.0M
1.35/3.59
200/9
100.00%(30)
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Loading chart...
Loading chart...
Loading chart...

MiniMax M3 context window

Input and output token limits for MiniMax M3, plus how it ranks on long-context understanding.

InputOutput
1Mtokens
1Mtokens
1.5k pages of text
1M
8K128K1M

MiniMax M3 API

Available from the model provider

MiniMax M3 is available from Fireworks, Novita, Together and 1 more. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

MiniMax M3 latency

MiniMax M3 time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Provider operational metrics

Time to first token, output throughput, and failed-request rate from live API traffic

Loading chart...
Loading chart...
Loading chart...

MiniMax M3 examples

Recent arena outputs from MiniMax M3, picked from the highest-ranked matchups.

MiniMax M3 license

MiniMax M3 is released under the MIT license, which permits commercial use.

License
MIT
Commercial use allowed

MIT License - allows commercial use

MiniMax M3 resources

Official sources for MiniMax M3: api documentation, official playground, official launch post.

MiniMax M3 vs other models

The most-compared alternatives to MiniMax M3 are Claude Opus 4.6, Qwen3.7 Max, Muse Spark 1.1. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like MiniMax M3

Models ranked just above and below MiniMax M3 by LLM Stats score.

 

Claude Opus 4.6

Score pending
 

Qwen3.7 Max

Score pending
 

Muse Spark 1.1

Score pending
 

Kimi K2.6

Score pending
 

GPT-5

Score pending
 

MiMo-V2.5

Score pending

FAQ

Common questions about MiniMax M3.

When was MiniMax M3 released?

MiniMax M3 was released on June 1, 2026 by MiniMax. This is the official MiniMax M3 release date tracked on LLM Stats.

How much does MiniMax M3 cost?

MiniMax M3 pricing starts at $0.30 per million input tokens and $1.20 per million output tokens via Fireworks, the lowest price among tracked providers.

Is MiniMax M3 available via API?

Yes, MiniMax M3 is available via API. See the official documentation for authentication and endpoint details. It is served by 4 providers tracked on LLM Stats.

Who created MiniMax M3?

MiniMax M3 was created by MiniMax.

What is the license for MiniMax M3?

MiniMax M3 is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

Is MiniMax M3 multimodal?

Yes, MiniMax M3 is multimodal and can accept both text and images as input.

What is MiniMax M3 latency?

MiniMax M3 p95 time to first token is 1.45 seconds via Together over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use MiniMax M3?

MiniMax M3 is available through 4 providers including Fireworks, Novita, Together, and 1 more.

Where is the MiniMax M3 paper or technical report?

MiniMax M3 has a paper or technical report available at https://www.minimax.io/blog/minimax-m3. Use that source for architecture, training, release and evaluation details.

What models should I compare MiniMax M3 against?

Common MiniMax M3 comparisons include MiniMax M3 vs Claude Opus 4.6, MiniMax M3 vs Qwen3.7 Max, MiniMax M3 vs Muse Spark 1.1. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.