AI Arena
Explore historical blind comparisons across code, image, video, and chat. The rankings and results remain available as a public archive.
Arenas are moving to Playground
We're sunsetting new Arena submissions in favor of Playground—a more powerful way to compare models across multiple turns. Existing rankings and results will remain available here.
Image
Archived image-generation matchups ranked by human preference across quality, style, and prompt accuracy.
Video
Historical text-to-video and image-to-video matchups, with results preserved in the Arena leaderboard.
Agent
Continue in Playground to compare frontier models and agents across multiple turns, tools, files, and longer tasks.
How the rankings were built
Prompt
People described what they wanted—a website, an image, a video clip, or an answer—and multiple models responded in parallel.
Compare blind
Outputs appeared side by side with model identities hidden. No logos or brand names influenced the comparison.
Vote & rank
Each preference updated the live rating system, producing rankings built from direct human judgment.






