The AI arena is free today

Open Superagent

Gemini 3.5 Flash-Lite vs Grok-4 Heavy

Grok-4 Heavy leads the LLM Stats Score 40.7 to 29.9.

Google · xAI · Updated for 2026

Which is better?

Grok-4 Heavy leads the overall LLM Stats Score 40.7 to 29.9, ranking #54 overall.

Based on current LLM Stats indexes, shared benchmarks, pricing, and model metadata for 2026.

Choose Gemini 3.5 Flash-Lite

  • you want predictable pricing at $0.30/M input and $2.50/M output

Choose Grok-4 Heavy

  • overall performance matters — it scores 40.7 and ranks #54 on LLM Stats
  • your work emphasizes reasoning — it leads those capability indexes

At a glance

The differences that matter most.

Core performance indexes
29.9
#123
40.7
#54
25.9
#148
39.0
#57
16.0
#127
17.7
#114
Cost, coverage & limits
Benchmark wins
Input price
$0.30 / M
— / M
Output price
$2.50 / M
— / M
Context window
1,048,576

Individual benchmarks

6 reported for Gemini 3.5 Flash-Lite · 6 for Grok-4 Heavy

No common benchmarks found

Gemini 3.5 Flash-Lite and Grok-4 Heavydon't have any common benchmark datasets to compare. They may have been evaluated on different testing suites.

Human preference

Blind head-to-head votes and playground preference scores

Context Window

Maximum input and output token capacity

Only Gemini 3.5 Flash-Lite specifies input context (1,048,576 tokens). Only Gemini 3.5 Flash-Lite specifies output context (65,536 tokens).

Google
Gemini 3.5 Flash-Lite
Input1,048,576 tokens
Output65,536 tokens
xAI
Grok-4 Heavy
Input- tokens
Output- tokens
Thu Sep 10 2026 • llm-stats.com

Input capabilities

Documented input modalities across available providers

Both Gemini 3.5 Flash-Lite and Grok-4 Heavy support multimodal inputs.

They are both capable of processing various types of data, offering versatility in application.

Gemini 3.5 Flash-Lite

Text
Images
Audio
Video

Grok-4 Heavy

Text
Images
Audio
Video

License

Usage and distribution terms

Both models are licensed under proprietary licenses.

Both models have usage restrictions defined by their respective organizations.

Gemini 3.5 Flash-Lite

Proprietary

Closed source

Grok-4 Heavy

Proprietary

Closed source

Release Timeline

When each model was launched

Gemini 3.5 Flash-Lite was released on 2026-07-21, while Grok-4 Heavy's release date is not specified.

We can confirm Gemini 3.5 Flash-Lite's release timeline, but cannot make a direct age comparison without Grok-4 Heavy's release date.

Gemini 3.5 Flash-Lite

Jul 21, 2026

1 months ago

Grok-4 Heavy

Knowledge Cutoff

When training data ends

Gemini 3.5 Flash-Lite has a knowledge cutoff of 2026-03-31, while Grok-4 Heavy has a cutoff of 2024-12-31.

Gemini 3.5 Flash-Lite has more recent training data (up to 2026-03-31), making it potentially better informed about events through that date compared to Grok-4 Heavy (2024-12-31).

Gemini 3.5 Flash-Lite

Mar 2026

1.3 yr newer
Grok-4 Heavy

Dec 2024

Outputs Comparison

Notice missing or incorrect data?Start an Issue discussion

Judge for yourself.

Run your own prompts against Gemini 3.5 Flash-Lite and Grok-4 Heavy side-by-side, then vote on the output you prefer.

Gemini 3.5 Flash-Lite
✓ Preferred
Grok-4 Heavy
Open in Playground

FAQ

Common questions about Gemini 3.5 Flash-Lite vs Grok-4 Heavy.

Which is better, Gemini 3.5 Flash-Lite or Grok-4 Heavy?

Grok-4 Heavy leads the LLM Stats Score 40.7 to 29.9. Gemini 3.5 Flash-Lite is made by Google and Grok-4 Heavy is made by xAI. The best choice depends on your use case — compare their capability indexes, individual benchmarks, pricing, and limits above.

How does Gemini 3.5 Flash-Lite compare to Grok-4 Heavy in benchmarks?

Gemini 3.5 Flash-Lite scores CharXiv-R: 76.5%, OSWorld-Verified: 74.0%, SWE-Bench Pro: 54.2%, Terminal-Bench 2.1: 54.0%, MLE-Bench: 39.2%. Grok-4 Heavy scores AIME 2025: 100.0%, HMMT25: 96.7%, GPQA: 88.4%, LiveCodeBench: 79.4%, USAMO25: 61.9%.

What are the context window sizes for Gemini 3.5 Flash-Lite and Grok-4 Heavy?

Gemini 3.5 Flash-Lite supports 1.0M tokens and Grok-4 Heavy supports an unknown number of tokens. A larger context window lets you process longer documents, conversations, or codebases in a single request.

What are the main differences between Gemini 3.5 Flash-Lite and Grok-4 Heavy?

Key differences include LLM Stats Score (29.9 vs 40.7). See the full comparison above for benchmark-by-benchmark results.

Who makes Gemini 3.5 Flash-Lite and Grok-4 Heavy?

Gemini 3.5 Flash-Lite is developed by Google and Grok-4 Heavy is developed by xAI.