The AI arena is free today

Open Superagent
Back to blog
Model Release·Agentic·Multimodal·API

Grok 4.7 Release: Same Price, Longer Horizons

Grok 4.7 is a same-price upgrade over Grok 4.6: CursorBench 4.0 46.3%, DeepSWE 71.0%, Terminal-Bench 38.0%, EEBench 66.0%, plus CADGen, cyber, and LatchBio capability rows. Still $2/$6 under 200k.

Sebastian Crossa
Sebastian Crossa
Co-Founder @ LLM Stats
·10 min read
Grok 4.7 Release: Same Price, Longer Horizons

Key Numbers

Grok 4.7 · Sep 21, 2026

0.0%
CursorBench 4.0 (xhigh)
0.0%
DeepSWE v1.1 (high)
0.0%
Terminal-Bench 4.0 (xhigh)
0.0%
EEBench (xhigh)
0K
Context
$0 / $6
List price (under 200k)

Same $2 / $6 list price and 500K context class as Grok 4.6. The story is longer RL on multi-hour tasks and stronger self-verification, not a new price tier. Scores are self-reported.

xAI released Grok 4.7 on September 21, 2026. Official surfaces brand the company as SpaceXAI. For SEO and prior Grok coverage, we keep the xAI / Grok naming readers already search. The release is a direct upgrade to Grok 4.6 at the same list price and the same 500K context class.

The thesis is simple. Larger base. Longer RL on multi-hour tasks. Stronger self-verification. Native Grok Bot harness understanding. A new safeguard stack. You are not buying a new price tier. You are buying more reliable progress on long-horizon coding and agent work at $2 / $6 under 200k.

Headline self-reported numbers: 46.3% on CursorBench 4.0 (xhigh), 71.0% on DeepSWE v1.1 (high), 38.0% on Terminal-Bench 4.0 (xhigh, Grok Build), 66.0% on EEBench (xhigh, model card). None of these are LLM Stats verified yet.


At a Glance

  • Release date: September 21, 2026. Generally available on named day-one surfaces.
  • Model ID: grok-4.7 on the Grok API.
  • Pricing: $2 / $0.50 cached / $6 per 1M under 200k prompt tokens. $4 / $1 / $12 at or above 200k. The higher tier bills the whole request. Matches Grok 4.6.
  • Context: 500,000 tokens. Pretraining cutoff June 2026. Supplemental training through August 2026.
  • Modalities: Text + image in, text out.
  • Reasoning: reasoning.effort low / medium / high (default) / xhigh.
  • Where: Cursor, Grok Build, Grok API, third-party harnesses, routers, and clouds (per the announcement).
  • Rate limits (docs): 150 rps, 50M tokens/min. Regions: us-east-1, us-west-2, us-central-1.

What's New in 4.7

None of these change the sticker price or the context class. Together they push Grok further into long, less-attended work.

Larger base and longer multi-hour RL

Per the launch post, Grok 4.7 uses a new, larger base than 4.6, then a longer reinforcement learning run on a harder mix weighted toward problems that take many hours. The model card adds that supplemental training ran longer than for 4.6, with agentic RL across knowledge work, general coding, and purpose-built environments.

Self-verification and longer context management

xAI's product claim is fewer steps and fewer output tokens on hard tasks, with more careful self-checking on long trajectories. CursorBench 4.0 is the cleanest public signal for that story: longer horizon than CursorBench 3.x, so scores are not comparable across versions.

Native Grok Bot harness understanding

The announcement says Grok 4.7 was trained to natively understand the Grok Bot harness. Read that as better fit for conversational and knowledge-work loops that sit on that stack, not as a free win in every third-party agent.

Effort rails

Docs expose four effort levels. Default is high. Several buyer-facing benches in the card are reported at xhigh. Treat the ladder as a cost and latency dial, then measure your own token-per-task.

reasoning.effort

four levels · default high

low
fastest rail
medium
mid rail
high
API default
docs default
xhigh
top published rung

Docs default to high. Several headline benches in the card are reported at xhigh. Set reasoning.effort explicitly when you want that ceiling, and re-measure token-per-task.

From docs.x.ai for grok-4.7: low, medium, high, xhigh. Conceptual ladder only. Not latency or quality scores.

Benchmarks

All scores below are self-reported by xAI in the announcement and the Grok 4.7 model card (September 21, 2026). Where those two sources disagree, we use the card. Not LLM Stats verified. The launch chart compares 4.7 to 4.6 only. Secondary domains below are absolute card scores.

4.7 vs 4.6 · percent benches

Grok 4.7Grok 4.6
CursorBench 4.0xhigh vs high
46.340.4+5.9
DeepSWE v1.1high
71.065.2+5.8
Terminal-Bench 4.0xhigh vs high
38.020.3+17.7

Grok Build harness

SWE-Marathon v1.1high
46.031.9+14.1
EEBenchxhigh
66.060.0+6.0

card; news table lists 64.0

HealthBench Professionalxhigh
56.748.5+8.2
FrontierSWE V2xhigh
29.025.3+3.7

partial-credit · Proximal harness

Harvey Legal Agentxhigh vs high
19.615.8+3.8
Self-reported by xAI (model card preferred; news table for CursorBench / DeepSWE / Terminal-Bench / Harvey / HealthBench deltas). Effort labeled per row; "xhigh vs high" means 4.7 at xhigh vs 4.6 at high. Not LLM Stats verified. EEBench uses the card's 66.0% at xhigh; the news comparison table lists 64.0%. FrontierSWE V2 is average partial-credit reward, not task resolution. Scores on a 0-100 scale.

Coding and long-horizon agents

CursorBench 4.0 is the buyer number for multi-hour IDE agents. Grok 4.7 hits 46.3% at xhigh (43.9% at high). The news table puts Grok 4.6 at 40.4% at high, so the +5.9 point headline is a cross-effort comparison, not a same-rung gain. Still the price-performance claim xAI leads with at the same $2 / $6 list price. Version 4.0 is not comparable to CursorBench 3.2 from the Grok 4.6 post.

DeepSWE v1.1 moves 65.2% → 71.0% at high. That is steady repository-level progress, not a cliff. SWE-Marathon v1.1 jumps harder: 31.9% → 46.0% at high. If your workload looks like long agent streaks rather than single-issue patches, Marathon is the more informative delta.

Terminal-Bench 4.0 is the largest relative move in the coding set: 20.3% at high → 38.0% at xhigh. Caveat: Grok's score uses the Grok Build harness. Compare harness to harness before you treat peer tables as apples to apples.

FrontierSWE V2 rises 25.3% → 29.0% at xhigh on Proximal's Proximus harness. The card is explicit: the headline is average partial-credit reward, not task resolution. Useful as a ultra-long-horizon signal. Wrong as a "pass rate."

Engineering acceleration

EEBench and CADGenBench sit next to coding as engineering acceleration, not as another SWE suite. EEBenchuses the card's 66.0% at xhigh (news comparison table lists 64.0%). Vs Grok 4.6 at xhigh that is 60.0% → 66.0%. CADGenBench is 44.4% at highon Grok Build: the card's generation split for CAD construction. Read EEBench as electrical engineering agent work; CADGen as whether the model can produce usable CAD construction steps, not as a coding pass rate.

Benchmark4.74.6Delta
Harvey Legal Agent (4.7 xhigh / 4.6 high)19.6%15.8%+3.8
HealthBench Professional (xhigh)56.7%48.5%+8.2

Harvey remains low in absolute terms even after the gain. HealthBench Pro's +8.2 is material if clinical reasoning is on your evaluation sheet. Neither replaces domain review with a human expert. The bio/lab rows below sit next to HealthBench as adjacent capability probes, not as a clinical product claim.

Cyber, bio / lab, and CAD (absolute)

These card rows do not have clean 4.6 deltas in the launch table, so they stay out of the hero chart. Grouped below by job. Cyber numbers are unrestricted capability scores from the card. Bio/lab and WMDP rows are capability or knowledge signals, not refusal metrics.

Secondary domains · absolute

card · no peer rows

Engineering

EEBench lives in the launch chart. CADGen is the CAD construction split.

CADGenBenchhigh
44.4%

Grok Build · generation split

Cyber capability

Unrestricted card scores. Capability probes, not refusal rates.

CyberGymhigh
80.3%

Mean Reproduced % · unrestricted · Grok Build

CVE-Benchxhigh
36.6%

primary reward · card notes 37.7% at high

CathedralBenchxhigh
29.0%

hard-subset accuracy · third-party red-team

Bio / lab

Capability and knowledge signals. Pair lightly with HealthBench Pro.

LatchBio Capabilitiesxhigh
44.5%

equal-weight mean · 11 benches

VCThigh
63.0%
LAB-Bench Practicalhigh
76.8%
ProtocolQA Open-Endedhigh
70.4%
BixBenchhigh
88.4%

zero-shot MCQ

WMDP-Biohigh
88.1%
WMDP-Chemhigh
84.9%
WMDP-Cyberhigh
88.1%
Self-reported model card scores for Grok 4.7 only. Absolute bars, not vs-4.6 deltas. CyberGym is unrestricted Mean Reproduced % without standard safeguards. WMDP rows are dual-use knowledge MCQs. Not LLM Stats verified.

On cyber: CyberGym at 80.3% (high, Mean Reproduced %, unrestricted / without standard safeguards, Grok Build) is the strongest absolute signal in the set. CVE-Bench at 36.6% (xhigh primary reward; card also notes 37.7% at high) and CathedralBench at 29% (xhigh, hard-subset accuracy, third-party red-team cyber) are harder probes. Treat them as capability measurements under the card's stated conditions.

On bio / lab: LatchBio Capabilities v1.0 at 44.5% (xhigh, equal-weight mean of 11 benches) is the composite. Single-suite highs include VCT 63.0%, LAB-Bench Practical 76.8%, ProtocolQA Open-Ended 70.4%, and BixBench 88.4% (all high except LatchBio at xhigh). WMDP-Bio / Chem / Cyber land at 88.1% / 84.9% / 88.1% (high). Those are dual-use knowledge MCQs. Useful as knowledge coverage. Not a biosafety policy statement.

Knowledge-work Elo

Knowledge work · Elo

1657

AA Briefcase v1.1

vs 1546 on Grok 4.6 (news table)

1695

GDPval

xhigh · from announcement

Elo, not percent. Bare bench titles from the xAI announcement table. Self-reported. Not LLM Stats verified.

AA Briefcase v1.1 moves 1546 → 1657 on the news table. GDPval is listed at 1695 Elo (xhigh) in announcement materials. Elo is not percent. Treat both as directional knowledge-work signals, not as a coding verdict.


Pricing & Availability

List price · matches 4.6

xAI public rates

Under 200k prompt tokens

$2
Input
per 1M · <200k prompt
$0.50
Cached input
per 1M · <200k prompt
$6
Output
per 1M · <200k prompt

At or above 200k · whole request

$4
Input
per 1M · ≥200k prompt
$1
Cached input
per 1M · ≥200k prompt
$12
Output
per 1M · ≥200k prompt

Once the prompt crosses 200k, every token in that request bills at the higher tier. List price matches Grok 4.6. The announcement also mentions a fast variant at 2× price / 2× output speed; docs currently list only grok-4.7.

Public xAI list pricing from docs.x.ai for grok-4.7. No third-party reseller rows. Fast variant noted from the launch post only; no separate public API id on the model docs page.
LaneInput / Cached / OutputNotes
Standard (<200k)$2 / $0.50 / $6List price per 1M tokens
Long context (≥200k)$4 / $1 / $12Whole request bills at the higher tier

The 200k cliff is still the cost story people miss. Cross 200k prompt tokens and every token in that request bills at the higher rate. List price matches 4.6, so migration is not a budget renegotiation. See docs.x.ai/docs/models/grok-4.7 for the live model page.

Availability per the announcement: Cursor, Grok Build, the Grok API, plus third-party coding harnesses, model routers, and cloud platforms. We only quote public xAI list pricing here.


When to Use / Migration from 4.6

Swap the model id to grok-4.7. Same price class. Same context class. Three checks before you flip agent traffic at scale.

1. Decide whether long-horizon coding is the job

Prefer 4.7 when Cursor-style multi-hour coding, terminal agents, or marathon-style streaks matter. CursorBench 4.0 and Terminal-Bench 4.0 are the clearest public deltas. If your traffic is short chat with light tool use, re-measure before you assume the upgrade pays for itself in quality alone.

2. Re-measure token-per-task at high and xhigh

Default effort is high. Several vendor headlines are xhigh. Output can rise when the model stays in the job longer. Do not assume 4.6 token budgets still hold.

3. Watch the 200k cliff; note the fast variant once

Long prompts still flip the whole request to $4 / $1 / $12. The announcement mentions a fast variant at 2× price and 2× output speed. Docs currently expose grok-4.7 only. Plan latency lanes from what the API actually lists in your account, not from a marketing aside.

  • Good fit: teams already on Grok 4.6 who want more persistence on hard coding and agent loops without changing price.
  • Also fit: Cursor and Grok Build users who want the new default model on those surfaces.
  • Hold / verify first: regulated legal or clinical workflows. Absolute Harvey and HealthBench scores still need human review.

Caveats

  • Self-reported table. Every number in this post is from xAI materials. Not LLM Stats verified.
  • Harness coupling.Terminal-Bench 4.0, EEBench, CADGenBench, and CyberGym for Grok use Grok Build. FrontierSWE V2 uses Proximal's Proximus harness and reports partial-credit reward, not resolution. CyberGym is unrestricted Mean Reproduced %.
  • EEBench disagreement. Model card: 66.0% at xhigh. News comparison table: 64.0%. We cite the card and keep the footnote.
  • CursorBench version break. 4.0 is not comparable to CursorBench 3.2 from the Grok 4.6 writeup.
  • Fast variant API surface. Mentioned in the launch post. No separate public API id on the model docs page at publish time.
  • Not for unattended high-stakes decisions. The model card states Grok 4.7 is not intended for autonomous high-stakes use in medicine, law, finance, or safety-critical systems without human oversight.

Outlook

Grok 4.7 is the same commercial shape as 4.6 with a harder training diet. The benches that move most are the ones that punish short attention: CursorBench 4.0, Terminal-Bench 4.0, SWE-Marathon, HealthBench Pro. The platform story is continuity. Same $2 / $6 opening price. Same 500K window. Same model-id swap path.

For buyers, the decision is whether long-horizon coding and agent persistence are the bottleneck. If they are, 4.7 is the obvious next Grok step. If they are not, treat the release as a free quality bump on an unchanged bill, then verify on your own evals.

Primary sources: the xAI launch post, the model docs page, and the Grok 4.7 model card (September 21, 2026).

Questions

Frequently Asked Questions

  • xAI (SpaceXAI on official pages) released Grok 4.7 on September 21, 2026. Day-one surfaces named in the launch post include Cursor, Grok Build, the Grok API, and third-party harnesses, routers, and clouds.
  • Under 200k prompt tokens, list pricing is $2 / $0.50 cached / $6 per million tokens. At or above 200k, the whole request bills at $4 / $1 / $12. That matches Grok 4.6. Rates are on docs.x.ai. The launch post also mentions a fast variant at 2× price and 2× output speed; the model docs page lists only grok-4.7.
  • Grok 4.7 supports a 500,000 token context window, the same class as Grok 4.6. Modalities are text + image in, text out. Pretraining cutoff is June 2026, with supplemental training data through August 2026.
  • Same list price and context class. The move is capability on longer horizons. Self-reported deltas include CursorBench 4.0 40.4% (high) → 46.3% (xhigh), DeepSWE v1.1 65.2% → 71.0% (high), Terminal-Bench 4.0 20.3% (high) → 38.0% (xhigh, Grok Build), SWE-Marathon v1.1 31.9% → 46.0% (high), and HealthBench Professional 48.5% → 56.7% (xhigh). Scores are vendor-reported, not LLM Stats verified.
  • Docs list reasoning.effort as low / medium / high / xhigh, with high as the default. Several headline benches in the model card are reported at xhigh. Re-measure token-per-task when you raise effort.
  • The launch post mentions a fast variant at twice the output speed and twice the price. The public model docs page currently lists only grok-4.7.

Continue Reading