The AI arena is free today

Open Superagent
Back to blog
Model Release·Technical Deep Dive

GPT-6 Astra: Flagship With a Launch Table

OpenAI GPT-6 Astra is staged GA at $10/$50 per 1M and 1.05M context. Self-reported Terminal-Bench 4.0 57.7, DeepSWE 74.1, TB-Science 64.6, GPQA 96.0. API default effort is low. Trusted Access first.

Sebastian Crossa
Sebastian Crossa
Co-Founder @ LLM Stats
·9 min read
GPT-6 Astra: Flagship With a Launch Table

Key Numbers

GPT-6 Astra · Sep 3, 2026

0.0%
Terminal-Bench 4.0
0.0%
DeepSWE v1.1
0.0%
Terminal-Bench-Science 0.1
0.0%
GPQA Diamond
0.0%
BrowseComp

Spec Rail

$10 / $50 · 1.05M · staged GA

$10 / $50
List · per 1M
1.05M
Context
gpt-6-astra
API id
staged GA
Rollout
Self-reported OpenAI launch table. API default effort is still low. Standard list $10 / $50 per 1M; context 1,050,000.

Astra is the GPT-6 flagship, not a Flash hop. Same $10 / $50 Standard list, 1.05M context, and reasoning.effort through max (API default still low). OpenAI now published a self-reported launch table. Hold those numbers as theirs, not as LLM Stats verified.

They call it their most capable model for the hardest end-to-end work: complex reasoning, coding, computer use, research, and document creation. That is a lab sentence. Hold the launch table, the sticker, and the staged rollout.


At a Glance

  • Catalog id: gpt-6-astra
  • Organization: OpenAI
  • Release: September 3, 2026, staged GA
  • API: gpt-6-astra
  • Pricing: $10 / $50 Standard; cache $1; writes $12.50
  • Long prompts:>272K input → 2× input/cache, 1.5× output for the full request
  • Context: 1,050,000 / 128K max output
  • Cutoff: April 30, 2026
  • Effort: low through max; API default low
  • Modalities: text + image in, text out
  • Hosting: first-party OpenAI in our catalog
  • Self-reported: TB4 57.7 · DeepSWE 74.1 · TB-Science 64.6 · GPQA 96.0 · BrowseComp 91.5

What's New

GPT-6 family flagship. Context jumps to 1,050,000. reasoning.effort adds xhigh and max on top of low / medium / high. The Standard sticker stays $10 / $50 per 1M, with cached input at $1 and an explicit cache-write line at $12.50. The launch post now carries a self-reported coding, science, computer-use, and cyber table.

Rollout is staged: Trusted Access Program enterprises today, then Plus, Pro, Business, Enterprise, and broader API in the coming days. Path to Astra (September 1, 2026) designates Astra at the Critical cybersecurity capability threshold under OpenAI's Preparedness Framework, with advanced defender access limited through Trusted Access and Daybreak Blue.

What the card supports

Streaming, function calling, and structured outputs. Chat Completions and Responses are supported; Batch is supported; fine-tuning is not. Responses API tools include web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Input is text and image; output is text only.

reasoning.effort

five levels · not a scorecard

low
API default
default in API
medium
mid rail
high
longer think
xhigh
new ceiling rung
max
top of the ladder
card says "Highest"

The model card markets Highest reasoning. The API still defaults to low. Set reasoning.effort explicitly when you want xhigh or max.

From the OpenAI model docs: low, medium, high, xhigh, and max. Conceptual ladder only. Not latency or quality scores.

Benchmarks

OpenAI published a self-reported launch table on the GPT-6 Astra launch page. Numbers below are theirs, not LLM Stats verified. Read by job: coding, science, computer use, then cyber.

Launch table · percent

AstraFable 5.1 (TB-Science only)
Terminal-Bench 4.0
57.7
DeepSWE v1.1
74.1
FrontierCode 1.1 Extended
64.5
Terminal-Bench-Science 0.1vs Fable 5.1
64.652.6+12.0
HLE with tools
57.2
OSWorld 2.0offline partial
72.6
BrowseComp
91.5
AutomationBench
41.4
Agents' Last Exam
59.3
Self-reported. Not LLM Stats verified. OSWorld is offline partial. ARC and GPQA are near-ceiling and omitted from this chart. Terminal-Bench- Science shows OpenAI's official pair vs Fable 5.1 (64.6 vs 52.6).

Coding

Terminal-Bench 4.0 lands at 57.7. DeepSWE v1.1 is 74.1, a small hop over GPT-5.6 Sol's 72.7 on OpenAI's own page. FrontierCode 1.1 Extended is 64.5; Main is 53.3.

Science and academic

Terminal-Bench-Science 0.1 at 64.6 vs Fable 5.1's 52.6 is the official comparison pair that matters on this page. FrontierMath T4 v2 is 97.6. HLE with tools is 57.2. GPQA Diamond at 96.0 is saturated; treat it as a ceiling row, not a mover.

Computer use and agents

BrowseComp is 91.5. OSWorld 2.0 is 72.6 offline partial. Agents' Last Exam is 59.3. AutomationBench is 41.4. ScreenSpot-Pro with no tools is 92.7; BenchCAD with Python is 95.9.

ARC

ARC-AGI-3 is 99.9, ARC-AGI-2 95.0, ARC-AGI-1 98.5. These are near-ceiling. The launch page footnotes ARC-AGI-3 to a Responses API harness; read that row with the harness note attached.

Cyber and health

ExploitBench is 100%, a cyber eval of exploit development from known vulnerabilities, not the coding headline. ExploitGym is 42.4; SEC-Bench Pro is 85.4. Health rows on the same table: GeneBench Pro 37.8, LifeSciBench 60.3, HealthBench Professional (length-adjusted) 63.4.


Pricing

First-party OpenAI Standard, USD per million tokens. Same sticker class as recent frontier OpenAI lists. Compare to Claude Fable 5.1: same $10 / $50 list, different cache math (Fable hits $0.25 vs Astra's $1). OpenAI has not published a cache-hit workload percentage for Astra.

Pricing · Standard

$10
Input
per 1M · Standard
$1
Cached input
0.1× of input
$12.50
Cache writes
1.25× uncached input
$50
Output
per 1M · Standard
2× / 1.5×
>272K input
full request: 2× input & cache, 1.5× output
50%
Batch & Flex
half of Standard rates
Fast mode
2× the applicable rates

Same frontier OpenAI sticker as recent GPT-5.6-class lists. Cache write at $12.50 is a 1.25× surcharge on uncached input, not a discount.

OpenAI Standard, September 3, 2026. Same sticker as recent frontier OpenAI. Cache write is 1.25×, not a discount.

When to Use It

  • Good fit: shops waiting for GPT-6 on hard coding, computer use, or long documents once they are off Trusted Access; 1.05M context jobs; teams that want xhigh or max effort.
  • Watch the API default: low, not max. Set reasoning.effort explicitly.
  • Prefer Fable 5.1 when you want Anthropic's published table and cheaper cache hits at matched $10 / $50 list.
  • Cyber: the default product is not unrestricted dual-use. Advanced cyber workflows route through Trusted Access / Daybreak.
  • Budget the 272K multiplier if prompts are huge. Fast is 2×.

Caveats

  • Staged rollout. Trusted Access first; broader plans and API follow in the coming days.
  • Self-reported table.Launch scores are OpenAI's, not LLM Stats verified. GPQA and ARC are near-ceiling.
  • API default effort is low. Marketing says Highest; the API does not default to max.
  • Critical cyber designation. Most advanced cyber capabilities start limited for trusted defenders.
  • Knowledge cutoff April 30, 2026.
  • No fine-tuning on this card.
  • Path to Astra states Astra was not involved in the Hugging Face incident.

Outlook

This post ages when LLM Stats verifies Terminal-Bench 4.0, DeepSWE, or OSWorld, or when Plus and API access are general. Through-line: GPT-6 flagship with a self-reported launch table, same $10 / $50 sticker, API default effort still low.

Sources: OpenAI GPT-6 Astra launch, model docs, Path to Astra.

Questions

Frequently Asked Questions

  • OpenAI staged GPT-6 Astra on September 3, 2026. Enterprises in the Trusted Access Program get it first. Plus, Pro, Business, Enterprise, and broader API access are listed as coming in the coming days.
  • Standard is $10 / $50 per 1M input / output. Cached input is $1. Cache writes are $12.50 (1.25× uncached input). Prompts over 272K input tokens bill at 2× input and cache rates and 1.5× output for the full request. Batch and Flex are 50% of Standard. Fast mode is 2× applicable rates.
  • 1,050,000 context tokens with up to 128,000 max output tokens. Knowledge cutoff is April 30, 2026.
  • low, medium, high, xhigh, and max. The API default is low. The model card markets Highest reasoning; set effort explicitly when you want xhigh or max.
  • Self-reported launch table (not LLM Stats verified): Terminal-Bench 4.0 57.7, DeepSWE v1.1 74.1, Terminal-Bench-Science 0.1 64.6(vs Fable 5.1 52.6 on OpenAI's comparison), GPQA Diamond 96.0, BrowseComp 91.5, OSWorld 2.0 72.6 offline partial. No GDPval row on the page.
  • Generational family hop to the GPT-6 flagship. Same Standard sticker class ($10 / $50), with a larger 1.05M window and new xhigh / max effort. On OpenAI's page, DeepSWE is 74.1 vs GPT-5.6 Sol 72.7. We do not copy a full 5.6 column onto every row.

Continue Reading