The AI arena is free today

Open Superagent

PostTrainBench

Progress Over Time

Interactive timeline showing model performance evolution on PostTrainBench

State-of-the-art frontier
Open
Proprietary

PostTrainBench Leaderboard

7 models
ContextCostLicense
1
Zhipu AI
Zhipu AI
753B1.0M$1.40 / $4.40
2
MiniMax
MiniMax
428B1.0M$0.30 / $1.20
3
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
4770B
5
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
6
7
ByteDance
ByteDance
Notice missing or incorrect data?
About this benchmark

What is PostTrainBench?

PostTrainBench evaluates a model's ability to autonomously post-train base models. Given pretrain-only base models, the agent must complete the full pipeline of data synthesis, training, evaluation, and iteration within a time budget, scored across downstream benchmarks such as AIME2025, BFCL, GPQA Main, GSM8K, and HumanEval.

PostTrainBench is a text benchmark evaluating models on reasoning, agents, code, and systems tasks. LLM Stats tracks 7 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.4.

Compare leaders on the best AI for reasoning, best AI for agents, best AI for code and best AI for systems leaderboards.

Current leaders

GLM-5.3 from Zhipu AI currently leads the PostTrainBench leaderboard with a score of 0.398 across 7 evaluated AI models.

1GLM-5.3Zhipu AI39.8%
2MiniMax M3MiniMax37.1%
3Kimi K3Moonshot AI36.6%

FAQ

Common questions about the PostTrainBench benchmark and leaderboard.

What is the PostTrainBench benchmark?

PostTrainBench evaluates a model's ability to autonomously post-train base models. Given pretrain-only base models, the agent must complete the full pipeline of data synthesis, training, evaluation, and iteration within a time budget, scored across downstream benchmarks such as AIME2025, BFCL, GPQA Main, GSM8K, and HumanEval.

What is the PostTrainBench leaderboard?

The PostTrainBench leaderboard ranks 7 AI models based on their performance on this benchmark. Currently, GLM-5.3 by Zhipu AI leads with a score of 0.398. The average score across all models is 0.312.

What is the highest PostTrainBench score?

The highest PostTrainBench score is 0.398, achieved by GLM-5.3 from Zhipu AI.

How many models are evaluated on PostTrainBench?

7 models have been evaluated on the PostTrainBench benchmark, with 0 verified results and 7 self-reported results.

What categories does PostTrainBench cover?

PostTrainBench is categorized under reasoning, agents, code, and systems. The benchmark evaluates text models.

What is the best open-source model on PostTrainBench?

MiniMax M3 by MiniMax is the top-ranked open-source model on PostTrainBench, with a score of 0.371 (rank #2).

Which model offers the best value on PostTrainBench?

Among models scoring within 10% of the leader, MiniMax M3 from MiniMax is the cheapest, at $0.30 per million input tokens with a score of 0.371.

How recent are the PostTrainBench leaderboard results?

The PostTrainBench leaderboard was last updated in September 2026 and currently includes 7 evaluated models.