The AI arena is free today

Open Superagent

AI Models

Compare 197 AI models and LLMs by benchmark intelligence, pricing, output speed, latency, context window, and live performance data.

Primer

What are AI models?

An AI model is a program — almost always a deep neural network — trained on a dataset to perform a task like generating text, generating images, transcribing speech, or classifying inputs. Artificial intelligence models are the same thing under a longer name; the two terms are interchangeable. The catalog above tracks every meaningful AI model from frontier labs and open-weight projects, with live pricing and benchmark scores in one place.

The most prominent class today are large language models (LLM models) like GPT, Claude, Gemini, Llama, Qwen, and DeepSeek — they predict the next token of text and power chat, agents, coding assistants, and structured-output workflows. Beyond LLMs, the catalog also tracks image generation, video generation, text-to-speech, and speech-to-text AI models.

Buyer's guide

How to choose an AI model

Picking the best AI model for a workload comes down to four tradeoffs: capability (does it score well on the benchmarks closest to your task), cost (per-token, per-image, per-second, or per-minute pricing), latency (time-to-first-token for interactive UX), and openness (whether you can self-host). The filter bar at the top of this page slices the catalog by all four.

For coding agents, sort by the Coding Index; for math, science, and multi-step problem solving, sort by the Reasoning Index. If you need to self-host or fine-tune, filter Openness to Open to surface every model with publicly released weights. To go deeper, see the full LLM leaderboard or the open LLM leaderboard.

AI Models FAQ

What an AI model is, the difference between AI models and LLMs, scoring, openness, latency, and how the proxy fits in.

What is a complete list of AI models in 2026?

This catalog tracks every meaningful AI model across five categories: large language models (GPT-5, Claude, Gemini, Llama, Qwen, DeepSeek, GLM, Grok, Mistral, Kimi…), image generation models (Flux, Imagen, GPT-Image, Recraft), video generation (Veo, Sora, Runway, Kling), text-to-speech (ElevenLabs, OpenAI TTS) and speech-to-text (Whisper, Deepgram). Use the category nav above to filter, or the search box to jump to a specific AI model.

What are the best AI models right now?

The "best" AI model depends on the task. For frontier reasoning, GPT-5, Claude Opus, Gemini 3 Pro and Grok 4 lead public benchmarks. For coding agents, sort by the Coding Index. For open-weight self-hosting, Llama, Qwen and DeepSeek are the strongest. The leaders summary above the catalog names the current winner per axis. Full rankings live on the LLM Leaderboard.

What is the cheapest AI model that's still good?

Combine the Price filter with the Reasoning Index filter. The cheapest models with frontier-grade reasoning are typically open-weight (Qwen, DeepSeek, GLM) hosted by inference providers like Together, Fireworks or Groq. For closed models, the GPT-5 nano / mini, Claude Haiku and Gemini Flash families are the cost leaders.

Which AI models support images, audio, or video?

Use the Modality filter. For image input, GPT-5, Claude (vision), Gemini and Llama 4 are the strongest. For audio input, GPT-5 with audio, Gemini and Whisper handle it natively. For video generation, Veo, Sora, Runway and Kling are the leading models; for video analysis, Gemini and GPT-5 vision can process frames.

What are the different types of AI models?

The catalog groups AI models into five practical types: language models for text, code, reasoning and tool use; image generation models for text-to-image and image editing; video generation models for text-to-video and image-to-video; text-to-speech models for voice synthesis; and speech-to-text models for transcription. Within LLMs there are further sub-types — reasoning models (o-series, DeepSeek-R1, Claude with thinking), multimodal models (GPT-5, Gemini, Claude vision) and open-weight models (Llama, Qwen, DeepSeek, Mistral, GLM).

What's the difference between an AI model and an LLM?

AI model is the umbrella term for any machine-learning model. LLM (large language model) is a specific kind of AI model trained on huge text corpora to predict the next token — that's what powers chatbots, code assistants and most agent products today. All LLMs are AI models, but not all AI models are LLMs: image generators (Flux, GPT-Image), video models (Veo, Sora), speech models (Whisper, ElevenLabs) and embedding models are AI models that aren't LLMs.

What is an AI model?

An AI model is a program — almost always a deep neural network — trained on data to perform a task: generating text, images, video or speech, classifying inputs, or transcribing audio. Modern AI models are dominated by large language models (LLMs) like GPT-5, Claude, Gemini and Llama, which generate text from a prompt. This catalog tracks every meaningful AI model across text, image, video, speech and transcription with live pricing and benchmark scores.

Where can I find open source AI models?

Filter the catalog by Openness → Open to see every model with publicly released weights — Llama, Qwen, DeepSeek, Mistral, GLM, Gemma and more. Open-source AI models can be self-hosted, fine-tuned and inspected, so they're the right fit for sensitive workloads. The Open LLM Leaderboard ranks them head-to-head.

How is AI model pricing calculated?

For LLMs, Price is a blended per-1M-token estimate (20 parts input, 1 part output) — that mirrors a chat and agent workload where prompts, tool traces and retrieved documents dominate output. Use this catalog as an LLM price comparison across OpenAI, Anthropic, Google, Meta, xAI, DeepSeek, Qwen and other providers. Other modalities use their native unit: per image, per second of video, per minute of audio, per 1M characters for TTS. All prices come from public provider price lists and are cross-checked against billing samples through the LLM Stats proxy.

How do I compare LLMs side by side?

Use the LLM comparison tool to compare two models by benchmarks, context window, API pricing, latency, throughput, release date, license and providers. For broader rankings, use the LLM Leaderboard.

What's the difference between the Reasoning and Coding indexes?

Both are composite TrueSkill ratings that aggregate public benchmark results into one comparable score per model. The Reasoning Index focuses on math, science, logic and multi-step problem solving — GPQA Diamond, AIME, MATH. The Coding Index focuses on software engineering — SWE-Bench Verified, LiveCodeBench, HumanEval. Higher is better. See the LLM Stats Score methodology for the full computation.

How are latency and throughput measured?

Latency is the median time-to-first-token over recent traffic flowing through the LLM Stats proxy; throughput is the median tokens-per-second once streaming begins. Both are rolling averages over the last 7 days, so a model that just launched will show fewer samples until it warms up.

How often are scores and prices updated?

Pricing and provider availability refresh hourly. Latency and throughput refresh every few minutes from live proxy traffic. Reasoning and Coding indexes recompute nightly as new benchmark results come in.