The AI arena is free today

Open Superagent
Jonathan Chavez

Jonathan Chavez

Co-Founder at LLM Stats

Jonathan co-founded LLM Stats to build independent, reproducible measurement infrastructure for AI models. He leads the platform's benchmark evaluation methodology, arena design, and data pipeline architecture. His work focuses on eliminating bias from AI model evaluation — designing blind voting systems, standardizing benchmark collection across providers, and publishing transparent ranking methodologies that frontier labs and Fortune 500 teams rely on for model selection decisions.

Expertise

AI Model EvaluationBenchmark DesignLLM InfrastructureData EngineeringMachine Learning Systems

Articles (26)

How to use DeepSeek V4 Flash 0731 today

Aug 4, 2026

Claude Opus 5 vs GPT-5.6 Sol: Frontier Analysis

Jul 24, 2026

Claude Opus 5 vs Claude Fable 5: Ultimate Comparison

Jul 24, 2026

Claude Code vs Cursor: The Ultimate Comparison

Jul 17, 2026

Gemini vs ChatGPT: A Complete Comparison in 2026

Jul 17, 2026

Claude vs ChatGPT: The Ultimate Comparison 2026

Jul 17, 2026

Kimi K3 vs Claude Fable 5: Complete Analysis

Jul 16, 2026

Claude Sonnet 5 vs Claude Opus 4.8: The Complete Comparison

Jun 30, 2026

GLM-5.2 vs Claude Opus 4.8: Full Comparison

Jun 16, 2026

Claude Fable 5 vs Claude Opus 4.8: Complete Comparison

Jun 9, 2026

Claude Fable 5: Review, Benchmarks and Pricing

Jun 9, 2026

Claude Opus 4.8 Release, Benchmarks And More

May 28, 2026

The Position of Your Context Matters for LLMs

May 27, 2026

Gemini 3.5 Flash: Benchmarks, Pricing, and Complete Specs

May 19, 2026

What Is a CUDA Kernel? A Visual Explainer

Apr 30, 2026

What Is a Contaminated LLM? Detection, Famous Cases, 2026 Guide

Apr 30, 2026

Is Fine-Tuning Better Than Prompt Engineering in 2026?

Apr 30, 2026

GPT-5.5 vs Claude Opus 4.7: Pricing, Speed, Benchmarks

Apr 23, 2026

GPT-5.5 vs GPT-5.4: Pricing, Speed, Context, Benchmarks

Apr 23, 2026

Claude Opus 4.7 vs Opus 4.6

Apr 17, 2026

Claude Opus 4.7: Benchmarks, Pricing, Context & What's New

Apr 16, 2026

Claude Mythos Preview: Benchmarks, Pricing & Project Glasswing

Apr 7, 2026

How to Calculate Hardware Requirements for Running LLMs Locally

Apr 3, 2026

Post-Training in 2026: GRPO, DAPO, RLVR & Beyond

Mar 11, 2026

Nemotron 3 Super: Pricing, Benchmarks, Architecture & API

Mar 11, 2026

Model Quantization Across Providers

Nov 28, 2024