- Organizations
- xAI
- Grok-4 Heavy
Grok-4 Heavy: Benchmarks, Pricing & Context Window
Grok-4 Heavy is a language model from xAI, released in July 2025, with multimodal input.
Grok 4 Heavy is the multi-agent version of Grok 4, released alongside the standard model in summer 2025. This system spawns multiple Grok 4 agents in parallel that work independently on problems and then collaborate by comparing their
Grok-4 Heavy benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Grok-4 Heavy across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Grok-4 Heavy holds up as conversations get longer.
Quality Tracker
Grok-4 Heavy Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Try now
Make it with
Grok-4 Heavy.
Grok-4 Heavy latency
Grok-4 Heavy time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.
Grok-4 Heavy examples
Recent arena outputs from Grok-4 Heavy, picked from the highest-ranked matchups.
Grok-4 Heavy license
Grok-4 Heavy is a proprietary model available under its provider's product and API terms, has a knowledge cutoff of December 2024.
- License
- Proprietary
- Hosted access
- Knowledge cutoff
- December 2024
Proprietary license - usage restrictions apply
Grok-4 Heavy resources
Official sources for Grok-4 Heavy: provider documentation.
Grok-4 Heavy vs other models
The most-compared alternatives to Grok-4 Heavy are Gemini 3 Pro, Seed 2.0 Pro, Grok-3. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Grok-4 Heavy
Models ranked just above and below Grok-4 Heavy by LLM Stats score.
FAQ
Common questions about Grok-4 Heavy.