The AI arena is free today

Open Superagent

GPT-4o vs Grok-1.5

GPT-4o leads the LLM Stats Score 14.1 to -2.6.

OpenAI · xAI · Updated for 2026

Which is better?

GPT-4o leads the overall LLM Stats Score 14.1 to -2.6, ranking #245 overall.

In the 6 individual benchmarks reported for both models, GPT-4o wins 6; this is a narrower head-to-head signal than the composite indexes.

Based on current LLM Stats indexes, shared benchmarks, pricing, and model metadata for 2026.

Choose GPT-4o

  • overall performance matters — it scores 14.1 and ranks #245 on LLM Stats
  • your work emphasizes reasoning — it leads those capability indexes
  • you value its reported benchmark strengths — it wins 6 of 6 exact shared results
  • you want the most recent training data — it shipped Aug 2024

Choose Grok-1.5

  • you are already invested in the xAI ecosystem

At a glance

The differences that matter most.

Core performance indexes
14.1
#245
-2.6
#347
15.9
#226
-2.2
#337
-0.6
#252
-0.9
#253
Cost, coverage & limits
Benchmark wins
6 of 6
0 of 6
Input price
$2.50 / M
— / M
Output price
$10.00 / M
— / M
Context window
128,000
—

Capability indexes

Additional strengths measured across groups of related public benchmarks

5 shared
Index
GPT-4o
Grok-1.5
12.5#235
4.9#279
8.3#142
-5.2#207
12.0#113
-2.6#166
16.3#114
4.7#174
16.3#97
4.7#160
Conservative TrueSkill rating · higher is betterHow scores work

Individual benchmarks

38 reported for GPT-4o · 9 for Grok-1.5

6 shared

GPT-4o outperforms in 6 benchmarks (DocVQA, GPQA, MathVista, MMLU, MMLU-Pro, MMMU), while Grok-1.5 is better at 0 benchmarks.

GPT-4o significantly outperforms across most benchmarks.

Tue Sep 29 2026 • llm-stats.com

Human preference

Blind head-to-head votes and playground preference scores

Context Window

Maximum input and output token capacity

Only GPT-4o specifies input context (128,000 tokens). Only GPT-4o specifies output context (16,384 tokens).

OpenAI
GPT-4o
Input128,000 tokens
Output16,384 tokens
xAI
Grok-1.5
Input- tokens
Output- tokens
Tue Sep 29 2026 • llm-stats.com

Input capabilities

Documented input modalities across available providers

GPT-4o supports multimodal inputs, whereas Grok-1.5 does not.

GPT-4o can handle both text and other forms of data like images, making it suitable for multimodal applications.

GPT-4o

Text
Images
Audio
Video

Grok-1.5

Text
Images
Audio
Video

License

Usage and distribution terms

Both models are licensed under proprietary licenses.

Both models have usage restrictions defined by their respective organizations.

GPT-4o

Proprietary

Closed source

Grok-1.5

Proprietary

Closed source

Release Timeline

When each model was launched

GPT-4o was released on 2024-08-06, while Grok-1.5 was released on 2024-03-28.

GPT-4o is 4 months newer than Grok-1.5.

GPT-4o

Aug 6, 2024

2.1 years ago

4mo newer
Grok-1.5

Mar 28, 2024

2.5 years ago

Knowledge Cutoff

When training data ends

Neither model specifies a knowledge cutoff date.

Unable to compare the recency of their training data.

No cutoff dates available

Outputs Comparison

Notice missing or incorrect data?Start an Issue discussion→

Judge for yourself.

Run your own prompts against GPT-4o and Grok-1.5 side-by-side, then vote on the output you prefer.

GPT-4o
✓ Preferred
Grok-1.5
Open in Playground

FAQ

Common questions about GPT-4o vs Grok-1.5.

Which is better, GPT-4o or Grok-1.5?

GPT-4o leads the LLM Stats Score 14.1 to -2.6. GPT-4o is made by OpenAI and Grok-1.5 is made by xAI. The best choice depends on your use case — compare their capability indexes, individual benchmarks, pricing, and limits above.

How does GPT-4o compare to Grok-1.5 in benchmarks?

GPT-4o scores AI2D: 94.2%, DocVQA: 92.8%, ChartQA: 85.7%, MMLU: 85.7%, CharXiv-D: 85.3%. Grok-1.5 scores GSM8k: 90.0%, DocVQA: 85.6%, MMLU: 81.3%, HumanEval: 74.1%, MMMU: 53.6%.

What are the context window sizes for GPT-4o and Grok-1.5?

GPT-4o supports 128K tokens and Grok-1.5 supports an unknown number of tokens. A larger context window lets you process longer documents, conversations, or codebases in a single request.

What are the main differences between GPT-4o and Grok-1.5?

Key differences include LLM Stats Score (14.1 vs -2.6), multimodal support (yes vs no). See the full comparison above for benchmark-by-benchmark results.

Who makes GPT-4o and Grok-1.5?

GPT-4o is developed by OpenAI and Grok-1.5 is developed by xAI.