PostTrainBench
Progress Over Time
Interactive timeline showing model performance evolution on PostTrainBench
PostTrainBench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | MiniMax | — | 1.0M | $0.30 / $1.20 | ||
| 2 | Kimi K3New Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 3 | Zhipu AI | 753B | 1.0M | $0.95 / $3.00 | ||
| 4 | ByteDance | — | — | — | ||
| 5 | ByteDance | — | — | — |
What is PostTrainBench?
PostTrainBench evaluates a model's ability to autonomously post-train base models. Given pretrain-only base models, the agent must complete the full pipeline of data synthesis, training, evaluation, and iteration within a time budget, scored across downstream benchmarks such as AIME2025, BFCL, GPQA Main, GSM8K, and HumanEval.
PostTrainBench is a text benchmark evaluating models on reasoning, agents, code, and systems tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.4.
Compare leaders on the best AI for reasoning, best AI for agents, best AI for code and best AI for systems leaderboards.
Current leaders
MiniMax M3 from MiniMax currently leads the PostTrainBench leaderboard with a score of 0.371 across 5 evaluated AI models.
FAQ
Common questions about the PostTrainBench benchmark and leaderboard.