The AI arena is free today

Open Superagent
Back to blog
Model Release·Technical Deep Dive

GPT-6.1 Sol: Near-Astra Work at One-Fifth the List

OpenAI GPT-6.1 Sol (gpt-6.1-sol): near-Astra coding, computer use, and professional work at $2/$10 vs Astra $10/$50. Self-reported DeepSWE 75.22% High, OSWorld 71.42% Max, TB-Science 57.02% Max.

Sebastian Crossa
Sebastian Crossa
Co-Founder @ LLM Stats
·9 min read
GPT-6.1 Sol: Near-Astra Work at One-Fifth the List

Key Numbers

GPT-6.1 Sol · Sep 29, 2026

0.00%
DeepSWE v1.1
High
0.00%
OSWorld 2.0
Max · offline partial
0.00%
Terminal-Bench Science 0.1
Max
0.00%
AutomationBench
Max
0.0%
GDP.pdf
High

Spec Rail

$2 / $10 · 1.05M · released

$2 / $10
List · per 1M
1.05M
Context
gpt-6.1-sol
Model id
released
Status
Self-reported OpenAI announcement charts. Not LLM Stats verified. Standard list $2 / $10 per 1M (one-fifth of Astra $10 / $50). Context 1,050,000. DeepSWE headline is High, not Max.

GPT-6.1 Sol is the Sol-class upgrade that approaches Astra on coding, computer use, and professional work at one-fifth of Astra's Standard list: $2 / $10 versus $10 / $50. Hold OpenAI's announcement charts as self-reported, not as LLM Stats verified.

The pitch is cost and throughput for near-Astra agentic jobs, not a claim that Sol beats Astra overall. Astra still leads hardest science (TB-Science Astra 68.1% vs Sol Max 57.02%). Prefer Astra when peak score matters; prefer 6.1 Sol when the Sol sticker and cache math matter.


At a Glance

  • Catalog id: gpt-6.1-sol
  • Organization: OpenAI
  • Status: API + ChatGPT Work / Codex released September 29, 2026 (not yet in Chat)
  • Model id: gpt-6.1-sol
  • Pricing: $2 / $10 Standard; cache $0.10; writes $2.50
  • vs Astra: one-fifth Standard input and output ($2 / $10 vs $10 / $50)
  • Long prompts:>272K input → 2× input/cache, 1.5× output for the full request
  • Context: 1,050,000 / 128K max output
  • Cutoff: April 30, 2026
  • Effort: low / medium (default) / high / xhigh / max; no none or minimal
  • Modalities: text + image in, text out
  • Self-reported: DeepSWE 75.22% High · OSWorld 71.42% Max · TB-Science 57.02% Max · AutomationBench 36.10% Max · GDP.pdf 32.0% High

What's New

OpenAI positions GPT-6.1 Sol as the successor to GPT-6 Sol that closes much of the gap to Astra on agentic coding, computer use, and professional work while staying in the Sol price class. Context stays at 1,050,000 with 128,000 max output. reasoning.effort runs low through max, defaulting to medium. There is no none or minimal rung.

Shipping today through the OpenAI API as gpt-6.1-sol, plus ChatGPT Work / Codex for Plus, Pro, Business, Enterprise, and Edu. ChatGPT Chat access is not included yet. An Ultrafast variant is mentioned as coming in days (OpenAI says up to 8× faster token generation); Ultrafast pricing is not published.

What the card supports

Streaming, function calling, and structured outputs. Responses is preferred for tools; Chat Completions is supported without tool calling. Batch is supported; fine-tuning is not. Responses API tools include web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Input is text and image; output is text only.

reasoning.effort

five levels · no none / minimal

low
shortest think
medium
API default
default
high
DeepSWE peak rung
xhigh
longer think
max
top of the ladder
OSWorld / TB-Science peak

Docs expose low, medium (default), high, xhigh, and max. There is no none or minimalrung. DeepSWE peaks at high on OpenAI's chart, not max. Set reasoning.effort explicitly when you leave the default.

From the OpenAI model docs for gpt-6.1-sol. Conceptual ladder only. Not latency or quality scores.

Benchmarks

OpenAI published self-reported announcement charts on the GPT-6.1 Sol announcement page. Numbers below are theirs, not LLM Stats verified. Read by job: coding, computer use, science, automation, then professional docs. Effort ladders matter: DeepSWE and GDP.pdf peak at High on these charts, not Max.

Launch charts · percent

6.1 SolAstra (TB-Science only)
DeepSWE v1.1High
75.22
OSWorld 2.0Max · offline partial
71.42
Terminal-Bench Science 0.1Maxvs Astra peak
57.0268.1-11.1
AutomationBenchMax
36.10
GDP.pdfHigh
32.00
Self-reported OpenAI announcement charts. Not LLM Stats verified. DeepSWE peaks at High (75.22%), not Max (71.90%). OSWorld is offline partial. TB-Science shows Astra's published peak (68.1%) still ahead of Sol Max (57.02%). Peer cells omitted where the announcement did not publish an exact score.

Catalog domains

Ten capability benches from the catalog sync, grouped by job. Coding / agents and documents echo the announcement peaks with effort notes. Health is the main added depth from the system card addendum (length-adjusted preferred). Absolute bars only; no peer inventions.

Catalog domains · by job

absolute · 10 capability benches

Coding / agents

Announcement peaks with effort notes. DeepSWE peaks at High, not Max.

DeepSWE v1.175.22%

High (best); Max/Xhigh 71.90%; Medium 73.01%; Low 64.38%

OSWorld 2.071.42%

Max · offline partial (v2026.08.08); High 69.56%; Medium 66.84%

Terminal-Bench Science 0.157.02%

Max · ~$5.47/task; Astra peak 68.1% still leads

AutomationBench 1.0.636.1%

Max

Documents

Professional document work. GDP.pdf peaks at High on the chart.

GDP.pdf32.0%

High; Xhigh 31.8%; Max 31.0%

Health

System card addendum, length-adjusted preferred. Professional LA within 0.5 pp of Astra 64.7.

HealthBench Professional LA64.2%

raw / unadj 67.2%

HealthBench Hard LA36.2%

unadj 33.4%

MentalHealthBench overall57.9%

Max · non-acute 57.3; high acuity 59.1; emergent 58.0

HealthBench LA58.5%

unadj 56.7%

HealthBench Consensus LA96.0%

near-ceiling

Self-reported OpenAI announcement charts and system card addendum capability rows for gpt-6.1-sol. Absolute bars, not vs-prior deltas. Safety, refusal, and jailbreak rows are omitted. Health prefers length-adjusted cites. Not LLM Stats verified.

Coding

DeepSWE v1.1 headlines at 75.22%on High ($0.65/task on OpenAI's chart), ahead of Max at 71.90% ($1.57). The Full ladder: Low 64.38% / $0.17, Medium 73.01% / $0.42, High 75.22% / $0.65, Xhigh 71.90% / $0.79, Max 71.90% / $1.57. OpenAI says Sol matches Astra at roughly one-fifth the cost and beats GPT-6 Sol's best by 6.4 pp at lower effort and cost. Prefer High when citing a single DeepSWE number.

Computer use

OSWorld 2.0 offline partial lands at 71.42% on Max ($1.27/task). High is 69.56% / $0.96; Xhigh is 69.38% / $1.05. OpenAI says that is +7 pp vs GPT-6 Sol at Max at less than half the cost, and within 2.1 pp of Astra Max at about one-seventh the cost.

Science

Terminal-Bench Science 0.1 reaches 57.02% on Max at about $5.47per task (High 51.14% / $2.76; Xhigh 53.71% / $2.89). OpenAI's announcement peers: Opus 5.5 about $23.21/task, Astra about $23.80/task. Astra's peak score remains 68.1%. Sol more than doubles GPT-6 Sol Max at less than half of that prior Sol cost. This is the clearest reminder that Astra still leads hardest science.

Automation and professional work

AutomationBench Max is 36.10% / $0.30 (Medium 31.70% / $0.19; High 33.20% / $0.23; Xhigh 35.50% / $0.25). OpenAI says +2.2 pp vs Opus 5.5 at medium at about one-third the cost, and +4.8 pp vs GPT-6 Sol at the same setting.

GDP.pdf peaks at High: 32.0% / $0.35 (Xhigh 31.8% / $0.37; Max 31.0% / $0.42). OpenAI says higher than Opus 5.5 with fallbacks at less than half the cost, and approaches Astra at about one-fifth the cost.

Health (system card)

From the system card addendum, prefer length-adjusted cites. HealthBench Professional LA 64.2% (raw 67.2%) sits within 0.5 pp of Astra 64.7. HealthBench Hard LA is 36.2%. MentalHealthBench overall is 57.9% at max (non-acute 57.3; high acuity 59.1; emergent 58.0; Astra 58.7; GPT-6 Sol 54.2). HealthBench LA 58.5% and Consensus LA 96.0% (near-ceiling) round out the rail above.

Factual error rate

Optional callout, lower is better. At low effort, GPT-6 Sol 11.4% moves to GPT-6.1 Sol 7.7% (about a 32% relative reduction). Across settings, OpenAI says Sol stays within 1.9 pp of Astra at less than one-fifth the cost. The eval is on flagged error conversations, not typical traffic.


Pricing

First-party OpenAI Standard, USD per million tokens. The headline is Sol-class list at one-fifth of Astra on both input and output, with a cheaper cached-input line than GPT-6 Sol.

Pricing · Standard

$2
Input
per 1M · Standard
$0.10
Cached input
5% of input · 50% under GPT-6 Sol cache
$2.50
Cache writes
1.25× uncached input
$10
Output
per 1M · Standard
1/5
vs Astra list
$2 / $10 vs Astra $10 / $50 on input and output
$0.10
Cached input
OpenAI: 50% less than GPT-6 Sol cached input
Sol class
Price tier
Near-Astra jobs without Astra sticker
2× / 1.5×
>272K input
full request: 2× input & cache, 1.5× output
50%
Batch & Flex
half of Standard rates
2×
Fast mode
2× Standard rates

The sticker story is one-fifth of Astra on Standard input and output, with cached input at $0.10. Regional endpoints add about 10% where available. Ultrafast is mentioned as coming; OpenAI has not published Ultrafast pricing yet.

OpenAI Standard, September 29, 2026. Cache write is 1.25× uncached input. Prompts above 272K use the long-context tier for the full request, same rule as other GPT-6 cards.

When to Use It

  • Good fit:shops that want near-Astra coding, computer use, or professional-doc work through the OpenAI API without paying Astra's $10 / $50 sticker; 1.05M context jobs; teams that care about cache hits at $0.10.
  • Prefer Astra when peak score on the hardest science or agentic jobs matters (TB-Science Astra 68.1% still leads Sol Max 57.02%).
  • Default effort is medium. Set reasoning.effort explicitly for High (DeepSWE / GDP.pdf peaks) or Max (OSWorld / TB-Science / AutomationBench peaks on these charts).
  • Upgrade path from GPT-6 Sol: same price class, stronger agentic coding and computer-use tables on OpenAI's charts, cheaper cached input.
  • Budget the 272K multiplier if prompts are huge. Fast is 2×. Regional endpoints add about 10% where available.

Caveats

  • Self-reported charts.Scores are OpenAI's announcement numbers, not LLM Stats verified. DeepSWE High beats Max on their chart; cite the effort rung with the score.
  • Not stronger than Astra overall. Astra still leads hardest science. Prefer Astra for peak score; Sol for cost and throughput.
  • Not yet in Chat. API and ChatGPT Work / Codex for paid tiers today; ChatGPT Chat comes later.
  • Default effort is medium. No none or minimal. Set effort explicitly when the chart peak matters.
  • Preparedness. OpenAI treats 6.1 Sol as Critical cybersecurity and High biological/chemical, with the same safeguard stack as Astra (Trusted Access / Daybreak for advanced cyber).
  • Knowledge cutoff April 30, 2026.
  • No fine-tuning on this card. Ultrafast pricing is not published yet.

Outlook

This post ages when LLM Stats verifies DeepSWE, OSWorld, or Terminal-Bench Science, when Ultrafast pricing lands, or when ChatGPT Chat access opens. Through-line: near-Astra agentic work in the Sol price class at $2 / $10, with self-reported charts that still leave Astra ahead on hardest science.

Sources: OpenAI GPT-6.1 Sol announcement, system card addendum, model docs.

Questions

Frequently Asked Questions

  • Yes. OpenAI released GPT-6.1 Sol on September 29, 2026 as gpt-6.1-sol. It is available in the API and in ChatGPT Work / Codex for Plus, Pro, Business, Enterprise, and Edu. It is not yet in ChatGPT Chat.
  • Standard is $2 / $10 per 1Minput / output: one-fifth of Astra's $10 / $50. Cached input is $0.10 (5% of input; OpenAI says 50% less than GPT-6 Sol cached input). Cache writes are $2.50. Prompts over 272K input tokens bill at 2× input and cache rates and 1.5× output for the full request. Batch and Flex are 50% of Standard. Fast mode is 2× Standard.
  • 1,050,000 context tokens with up to 128,000 max output tokens. Knowledge cutoff is April 30, 2026.
  • low, medium (default), high, xhigh, and max. There is no none or minimal option. Set effort explicitly when you leave the default.
  • Self-reported announcement charts (not LLM Stats verified): DeepSWE v1.1 75.22% at High, OSWorld 2.0 offline partial 71.42% at Max, Terminal-Bench Science 0.1 57.02% at Max (~$5.47/task), AutomationBench 36.10% at Max, GDP.pdf 32.0% at High. Astra still leads TB-Science at 68.1%.
  • Prefer 6.1 Sol when cost and throughput matter and you want near-Astra results on coding, computer use, and professional work at Sol-class list prices. Prefer Astra when peak score on the hardest science and agentic jobs matters. Do not read Sol as stronger than Astra overall.

Continue Reading