PostTrainBench
Progress Over Time
Interactive timeline showing model performance evolution on PostTrainBench
PostTrainBench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Zhipu AI | 753B | 1.0M | $1.40 / $4.40 | ||
| 2 | MiniMax | 428B | 1.0M | $0.30 / $1.20 | ||
| 3 | Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 4 | Tencent | 770B | — | — | ||
| 5 | Zhipu AI | 753B | 1.0M | $0.95 / $3.00 | ||
| 6 | ByteDance | — | — | — | ||
| 7 | ByteDance | — | — | — |
What is PostTrainBench?
PostTrainBench evaluates a model's ability to autonomously post-train base models. Given pretrain-only base models, the agent must complete the full pipeline of data synthesis, training, evaluation, and iteration within a time budget, scored across downstream benchmarks such as AIME2025, BFCL, GPQA Main, GSM8K, and HumanEval.
PostTrainBench is a text benchmark evaluating models on reasoning, agents, code, and systems tasks. LLM Stats tracks 7 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.4.
Compare leaders on the best AI for reasoning, best AI for agents, best AI for code and best AI for systems leaderboards.
Current leaders
GLM-5.3 from Zhipu AI currently leads the PostTrainBench leaderboard with a score of 0.398 across 7 evaluated AI models.
FAQ
Common questions about the PostTrainBench benchmark and leaderboard.