Artifacts Bench
Progress Over Time
Interactive timeline showing model performance evolution on Artifacts Bench
Artifacts Bench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | ByteDance | — | — | — | ||
| 2 | ByteDance | — | — | — | ||
| 3 | Microsoft | — | — | — |
What is Artifacts Bench?
Artifacts Bench evaluates a model's ability to generate visual code artifacts, measuring the quality of generated interactive and visual front-end outputs from natural-language requests.
Artifacts Bench is a text benchmark evaluating models on frontend development and code tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.5.
Compare leaders on the best AI for frontend development and best AI for code leaderboards.
Current leaders
Seed 2.1 Pro from ByteDance currently leads the Artifacts Bench leaderboard with a score of 0.510 across 3 evaluated AI models.
FAQ
Common questions about the Artifacts Bench benchmark and leaderboard.