KernelBench Hard
Progress Over Time
Interactive timeline showing model performance evolution on KernelBench Hard
KernelBench Hard Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | MiniMax | — | 1.0M | $0.30 / $1.20 |
What is KernelBench Hard?
KernelBench Hard evaluates agentic GPU kernel optimization on the hardest problem set. Each question is scored by the agent's submitted operator TFLOPs relative to the theoretical peak of the current hardware, with the benchmark score being the average across all questions.
KernelBench Hard is a text benchmark evaluating models on agents, code, and systems tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.3.
Compare leaders on the best AI for agents, best AI for code and best AI for systems leaderboards.
Current leaders
MiniMax M3 from MiniMax currently leads the KernelBench Hard leaderboard with a score of 0.288 across 1 evaluated AI models.
FAQ
Common questions about the KernelBench Hard benchmark and leaderboard.