KernelBench Hard

Progress Over Time

Interactive timeline showing model performance evolution on KernelBench Hard

State-of-the-art frontier
Open
Proprietary

KernelBench Hard Leaderboard

1 models
ContextCostLicense
1
MiniMax
MiniMax
1.0M$0.30 / $1.20
Notice missing or incorrect data?
About this benchmark

What is KernelBench Hard?

KernelBench Hard evaluates agentic GPU kernel optimization on the hardest problem set. Each question is scored by the agent's submitted operator TFLOPs relative to the theoretical peak of the current hardware, with the benchmark score being the average across all questions.

KernelBench Hard is a text benchmark evaluating models on agents, code, and systems tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.3.

Compare leaders on the best AI for agents, best AI for code and best AI for systems leaderboards.

Current leaders

MiniMax M3 from MiniMax currently leads the KernelBench Hard leaderboard with a score of 0.288 across 1 evaluated AI models.

1MiniMax M3MiniMax28.8%

FAQ

Common questions about the KernelBench Hard benchmark and leaderboard.

What is the KernelBench Hard benchmark?

KernelBench Hard evaluates agentic GPU kernel optimization on the hardest problem set. Each question is scored by the agent's submitted operator TFLOPs relative to the theoretical peak of the current hardware, with the benchmark score being the average across all questions.

What is the KernelBench Hard leaderboard?

The KernelBench Hard leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, MiniMax M3 by MiniMax leads with a score of 0.288. The average score across all models is 0.288.

What is the highest KernelBench Hard score?

The highest KernelBench Hard score is 0.288, achieved by MiniMax M3 from MiniMax.

How many models are evaluated on KernelBench Hard?

1 models have been evaluated on the KernelBench Hard benchmark, with 0 verified results and 1 self-reported results.

What categories does KernelBench Hard cover?

KernelBench Hard is categorized under agents, code, and systems. The benchmark evaluates text models.

What is the best open-source model on KernelBench Hard?

MiniMax M3 by MiniMax is the top-ranked open-source model on KernelBench Hard, with a score of 0.288 (rank #1).

Which model offers the best value on KernelBench Hard?

Among models scoring within 10% of the leader, MiniMax M3 from MiniMax is the cheapest, at $0.30 per million input tokens with a score of 0.288.

How recent are the KernelBench Hard leaderboard results?

The KernelBench Hard leaderboard was last updated in July 2026 and currently includes 1 evaluated models.