The AI arena is free today

Open Superagent
Back to blog
Model Release·Technical Deep Dive

Mistral Large 4: Europe-Trained MoE Preview at Half List

Mistral Large 4 (also ML4 / le Chonk; mistral-large-4) Studio + API public preview: Europe-trained open-weight-bound MoE. Docs 1.05T/52B active/1.6B vision/1M ctx. Self-reported DeepSWE 61.7%, Coding Agent Index 49.8%, CyberGym-E2E 82%. Current sale $0.68/$2.09 (list $1.36/$4.18); no official end date published.

Sebastian Crossa
Sebastian Crossa
Co-Founder @ LLM Stats
·9 min read
Mistral Large 4: Europe-Trained MoE Preview at Half List

Key Numbers

Mistral Large 4 · Oct 6, 2026

0.0%
DeepSWE 1.1
0.0%
Coding Agent Index
0%
CyberGym-E2E
highest of any model per Mistral
0.0%
Dense200 bbox
vs Astra 41%
$0.00 / $2.09
Sale input / output
per 1M · current sale · no official end date
0M
Context
docs: 1.05T / 52B active MoE

Spec Rail

Europe-trained MoE · open-weight-bound

1.05T / 52B
Total / active (docs)
1.6B
Vision encoder
mistral-large-4
API id
Studio + API
Preview · weights end Oct
Self-reported Mistral announce and docs. Prefer docs specs (1.05T / 52B active / 1.6B vision / 1M context) over announce rounding (~1T / 49B). Current sale $0.68 / $2.09 vs list $1.36 / $4.18; no official end date published. Studio + API preview. Not LLM Stats verified. Do not read as beating closed frontier overall.

Mistral Large 4 is a Europe-trained MoE in Studio + API public preview that Mistral positions as open-weight-bound by end of October. It posts strong self-reported coding, agentic, and cyber numbers at a current half-of-list sale on the pricing page. Prefer docs specs: 1.05T total / 52B active / 1.6B vision / 1M context over announce rounding (~1T / 49B).

The honest frame is open-model strength under preview pricing, not a closed-frontier crown. Headline self-reported numbers: 61.7% DeepSWE 1.1, 49.8% Coding Agent Index, 82% CyberGym-E2E, 42.0% Dense200 bbox (vs Astra 41%). Current sale is $0.68 / $2.09 per 1M (list $1.36 / $4.18); no official end date is published. API id: mistral-large-4. Cyber wins partly vs models that refuse. None of these scores are LLM Stats verified yet.


At a Glance

  • Catalog / API id: mistral-large-4 (alias mistral-large-4-0 exists; primary is without -0)
  • Organization: Mistral AI
  • Availability: Studio + API public preview (mistral-large-4). October 6, 2026. Weights planned end of October 2026.
  • Architecture: Granular MoE, natively multimodal. Docs: 1.05T total / 52B active + 1.6B vision encoder. Announce rounds to ~1T / 49B.
  • Pricing (current sale): $0.68 input / $0.07 cached / $2.09 output per 1M. List: $1.36 / $0.14 / $4.18. No official sale end date published.
  • Context: 1,000,000 tokens. Max output not published.
  • Modalities: Text + image in, text out
  • Training:From scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's European datacenters; preview served on the same infra. Training data multilingual across more than 160 languages, including every official EU language.
  • License: Proprietary until the weight drop. Weight license TBD.

What's New

Large 4 is Mistral's frontier open-weight-bound MoE in Studio + API public preview, trained and served in Europe under European law, with weights planned for end of October. That availability story (Studio + API now, weights soon, EU deployment) matters as much as any single bench row.

On capability, Mistral's announce shows a strong open-model showing on coding and agentic harnesses, Dense200 grounding slightly ahead of GPT-6-Astra (42% vs 41%), and cyber numbers where closed models often refuse. SciCode-Verified is claimed as open-weight SOTA at 91.8% pass@1 (n=6).

On price, the current preview sale on the pricing page is half of list: $0.68 / $2.09 versus list $1.36 / $4.18, with 1M context. No official sale end date is published. RL is still in flight per Mistral, so treat scores as a moving preview snapshot.


Benchmarks

All scores below are self-reportedfrom Mistral's announce prose and charts. Not LLM Stats verified. Interpret by job. Do not invent peer scores where Mistral only says "ahead of." Do not claim overall closed-frontier superiority.

Coding / agents · absolute

ML4Astra (Dense200 only)
DeepSWE 1.1
61.7
SWE-Atlas-QnA
59.4
AutomationBenchahead of Kimi K3, MiMo-V2.6-Pro, DeepSeek V4 Pro (no peer %)
59.9
Coding Agent Indexahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max (no peer %)
49.8
Terminal-Bench 4.0
28.3
Dense200 bboxvs GPT-6-Astra
42.041.01.0
Self-reported by Mistral (announce prose and charts). Not LLM Stats verified. Peer bars only where Mistral published an exact peer number (Dense200 42.0% vs GPT-6-Astra 41%). Coding Agent Index and AutomationBench peer leads are qualitative only.

Coding and agents

DeepSWE 1.1 at 61.7% and SWE-Atlas-QnA at 59.4% are the software-engineering headlines. Coding Agent Index at 49.8% is ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max per Mistral (no peer percentages published). AutomationBench at 59.9% covers 657 business workflows and is ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro. Terminal-Bench 4.0 at 28.3% is the quieter terminal row. AA-Briefcase at 1393 Elo is the long-horizon knowledge-work score, ahead of DeepSeek V4 Pro per Mistral.

Vision, science, and professional

Dense200 bbox at 42.0%edges GPT-6-Astra's 41% on grounding. ChartQA Pro is 63.1%; GDP.pdf is 18.6%. SciCode-Verified pass@1 (n=6) is 91.8% (open-weight SOTA per Mistral). Finance Agent v2 54.7% and Harvey LAB 15.8% are the professional rows; Mistral says finance/legal exceeds Astra and Harvey outperforms all open-source (qualitative peer claims). Finch / FinWorkBench is 67.4%.

Cyber capability

Cybench at 93% (40 CTF-style challenges) and CyberGym-E2E at 82% are the cyber headlines. CyberGym-E2E is highest of any model per Mistral; closed models land near zero partly because they refuse. Read cyber as capability under that harness, not as universal superiority without refusal notes.

Catalog domains by job

Fifteen catalog capability benches, grouped by job. Absolute bars. Safety and refusal rows are skipped. Self-reported; not LLM Stats verified.

Catalog domains · by job

absolute · 15 capability benches

Coding / agents

Self-reported coding and agent harnesses. Peer leads are qualitative except where noted.

DeepSWE 1.161.7%
SWE-Atlas-QnA59.4%
AutomationBench59.9%

657 business workflows · ahead of Kimi K3, MiMo-V2.6-Pro, DeepSeek V4 Pro

Coding Agent Index49.8%

ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max

Terminal-Bench 4.028.3%
AA-Briefcase1,393 Elo

long-horizon knowledge work · ahead of DeepSeek V4 Pro

Vision / docs

Grounding and document suites from announce charts.

Dense200 bbox42%

vs GPT-6-Astra 41%

ChartQA Pro63.1%
GDP.pdf18.6%

Science / professional

SciCode plus finance and legal rows. Finance/legal peer claims are qualitative unless noted.

SciCode-Verified91.8%

pass@1 (n=6) · open-weight SOTA per Mistral

Finch / FinWorkBench67.4%

spreadsheet create / edit finance work

Finance Agent v254.7%

exceeds Astra per Mistral (vals.ai)

Harvey LAB15.8%

outperforms all open-source per Mistral

Cyber

Capability rows only. Closed models near zero on CyberGym-E2E partly due to refusals.

Cybench93%

40 CTF-style challenges

CyberGym-E2E82%

highest of any model per Mistral · refusal effect on closed peers

Self-reported Mistral announce prose and charts for mistral-large-4. Absolute bars, not invented peer deltas. AA-Briefcase is Elo (rail scaled for display only). Safety and refusal rows are omitted. Cyber wins partly vs models that refuse. Not LLM Stats verified.

How to read the table

Coding agents: DeepSWE, SWE-Atlas, Coding Agent Index, AutomationBench, Terminal-Bench, plus AA-Briefcase Elo for long-horizon work. Vision / docs: Dense200, ChartQA Pro, GDP.pdf. Science / professional: SciCode, Finance Agent, Harvey LAB, Finch. Cyber: Cybench and CyberGym-E2E, with the refusal caveat attached. Prefer Large 4 when Europe-served open-weight-bound pricing and those jobs dominate; prefer closed flagships when you need peak verified score on the hardest suites.


Pricing

First-party Mistral docs, USD per million tokens. Current preview sale on the pricing page is half of list. No official sale end date is published.

Pricing · sale vs list

first-party Mistral docs

$0.68
list $1.36
Input
per 1M · current sale
$0.07
list $0.14
Cached input
per 1M · sale vs list
$2.09
list $4.18
Output
per 1M · current sale
half list
Current sale
Pricing page shows sale vs list; no official end date published
Studio + API
Availability
Public preview on Mistral Studio and API only
end of Oct
Weights planned
Open weights planned end of October 2026

Current preview sale on the pricing page is half of list: $0.68 / $2.09 versus list $1.36 / $4.18 per 1M, with cached input at $0.07 (list $0.14). No official sale end date is published.

USD per million tokens from Mistral docs pricing. Public API id mistral-large-4 (alias mistral-large-4-0 exists). Regional / Batch / Priority multipliers not extracted for this post.

Public callers use mistral-large-4. An alias mistral-large-4-0 exists on the docs model page; primary id is without -0.


When to Use / Migrate

  • Good fit: Coding and agent loops, business workflow automation, document grounding, finance/legal agent prototypes, and cyber research harnesses where you want a strong open-weight-bound Europe-served option under sale pricing.
  • Try now if the current half-of-list sale and 1M context matter and you can accept self-reported Studio + API preview scores while RL is still running.
  • Wait for weights if on-prem / self-host license terms matter more than Studio + API preview access. Weights are planned for end of October; license TBD.
  • Prefer closed flagships when you need peak verified score on the hardest agentic or science jobs, or when refusal behavior on cyber harnesses is a product requirement rather than a scoreboard artifact.
  • Do not migrate on cyber alone. CyberGym-E2E wins partly vs models that refuse. Validate on your own harness before treating 82% as production superiority.

Caveats

  • Public preview, not GA. Mistral says RL is still in flight; scores can move.
  • All benches above are self-reported by Mistral from announce prose and charts. Not LLM Stats verified. Harness versions and agent scaffolds matter, especially SWE, Terminal, Automation, and cyber.
  • Spec conflict: announce rounds to ~1T / 49B active; docs say 1.05T / 52B active + 1.6B vision. Prefer docs for specs.
  • Current sale on the pricing page is $0.68 / $0.07 / $2.09 vs list $1.36 / $0.14 / $4.18. No official sale end date is published.
  • Cyber wins partly vs closed models that refuse. Do not overread CyberGym-E2E 82% as universal cyber superiority.
  • Do not claim Large 4 beats closed frontier overall. Peer leads without published peer percentages stay qualitative.
  • Max output and knowledge cutoff are not published on the official model card. Weights / license TBD until the end-of-October drop.
  • Safety, refusal, and jailbreak scoreboards are intentionally omitted from this post.

Outlook

Near-term watch items: whether sale pricing changes on the docs page, the end-of-October weight drop and license, and whether independent harnesses confirm the self-reported coding and cyber rows. This post ages when LLM Stats verifies DeepSWE, Coding Agent Index, or CyberGym-E2E, and when weights ship.

Through-line for now: Europe-trained open-weight-bound MoE in Studio + API public preview, strong self-reported open-model showing on coding/agents and cyber at current half-of-list sale, with 1M context and docs-size 1.05T / 52B. For primary sources, see Mistral's Large 4 announcement, the docs model card, and pricing.

Questions

Frequently Asked Questions

  • Mistral announced Mistral Large 4 (also called ML4 / le Chonk) on October 6, 2026 as a Studio + API public preview. API id is mistral-large-4 (alias mistral-large-4-0 exists). Open weights are planned for end of October 2026.
  • Current preview sale on Mistral's pricing page is $0.68 / $2.09 per 1M input / output, with cached input at $0.07. List is $1.36 / $0.14 cached / $4.18. No official sale end date is published.
  • Docs list a 1,000,000 token context window. Max output is not published on the official model card. Modalities are text + image in, text out. Prefer docs size: 1.05T total / 52B active MoE plus a 1.6B vision encoder (announce rounds to ~1T / 49B).
  • Self-reported announce prose and charts (not LLM Stats verified): DeepSWE 1.1 61.7%, SWE-Atlas-QnA 59.4%, Terminal-Bench 4.0 28.3%, Coding Agent Index 49.8%, AutomationBench 59.9%, AA-Briefcase 1393 Elo, Dense200 bbox 42.0%, ChartQA Pro 63.1%, GDP.pdf 18.6%, SciCode-Verified pass@1 (n=6) 91.8%, Finance Agent v2 54.7%, Harvey LAB 15.8%, Finch / FinWorkBench 67.4%, Cybench 93%, CyberGym-E2E 82%.
  • No overall claim. Mistral positions Large 4 as a strong open-weight-bound showing on coding, agents, cyber, and Dense200 grounding (42% vs Astra 41%). CyberGym-E2E at 82% is highest of any model per Mistral, but closed peers land near zero partly because they refuse. Prefer Large 4 for Europe-served open-weight roadmap and current sale pricing; prefer closed flagships when peak verified score on hardest jobs matters.

  • Mistral says open weights are planned for end of October 2026. Until then the model is proprietary SaaS / self-deploy under Mistral terms. License for the weight drop is not published yet.

Continue Reading