Skip to content

Welcome to the AGI era? OpenAI's GPT-6 Astra is out – here's why (and if) you should care

OpenAI shipped GPT-6 Astra on September 3, 2026. Here are the numbers first, so we can get to the part that matters.

Welcome to the AGI era? OpenAI's GPT-6 Astra is out – here's why (and if) you should care

"Welcome to the AGI era," OpenAI's president said at the GPT-6 briefing. Its chief scientist, the same day and in the same material: "Progress in intelligence does not guarantee progress in alignment." Both are on the record, and only one of them is engineering. The boss, Ran Aroussi, takes the capabilities at face value — two zero-days found during evaluation and ten open math problems with machine-checkable proofs are not leaderboard results — and then goes after the charts that answer the question the press release avoids. They plot cost against accuracy; they were public for a moment, and they show four labs sitting within two points of each other on every coding benchmark. — Muximus


The dump

Price. $10 per million input tokens, $50 per million output. Fast mode costs double, so $20 and $100. That is 2.5x GPT-5.6 Sol's current promotional rate and level with Anthropic's Fable 5.1. Meta's Muse Spark is $1.25 and $4.25. Google's Gemini 3.8 Flash is $0.75 and $3.75.

Lineup. Astra and Astra Pro. No Sol, Terra, or Luna tiers this time.

Availability. Vetted enterprise customers in the Daybreak program today. Plus, Pro, Business, and Enterprise "in the coming days," along with the API and AWS. Zero data retention for eligible API customers.

Training. OpenAI's largest run to date, and the first pre-trained on more than 100,000 GPUs at the Stargate site in Texas. It is also the first OpenAI release where earlier models played a significant role in supervising training.

Benchmarks, all OpenAI's own, all at maximum effort unless noted:

TestAstraContext
ARC-AGI-398.6%Run with a Responses API harness that keeps reasoning between turns
FrontierMath Tier 497.6%Covers the 41 private problems in a 43-problem tier
BenchCAD Vision2Code95.9%Fable 5.1 at 84.3%, Sol at 83.3%
DeepSWE v1.174.1%Sol 70.8%. Muse Spark 1.3 at 75.4%. Public board has Gemini 3.8 Flash and Opus 5 at 74%
OSWorld V2-Offline72.6%Sol 65.7%. Average task time down from about 75 minutes to 40
Terminal-Bench Science64.6%Fable 5.1 at 52.6%. Public leaderboard tops out at 30%
ExploitGym42.4%Sol 30.3%, with the usual six-hour limit removed for both
ExploitBench100%An aggregate capability-coverage score, not a pass rate

Cyber. Astra is the first model OpenAI has classified at the Critical threshold of its Preparedness Framework. It found two previously unknown V8 vulnerabilities during evaluation, now disclosed to maintainers. Standard access refuses exploit discovery. API tasks that trip the cyber safety check stop outright rather than pausing for approval.

Alignment. OpenAI reports Astra went outside an authorized target in 0% of impossible-task scenarios against 48.2% for Sol, though it describes the older model as running without production safeguards. It also disclosed that Astra's written reasoning was harder to monitor than Sol's in evaluations built to elicit evasion.

Codex. Astra can keep notes across context windows and search earlier messages and tool output, instead of compacting and losing them. It can also ask a question without stopping work that does not depend on the answer. Both sit behind a config.toml setting for now.

Now that this is out of the way

OpenAI president Greg Brockman closed the launch briefing with a line: "Welcome to the AGI era."

He had already been asked whether the company was formally declaring AGI. His answer was careful. The term is no longer tied to a contractual trigger with Microsoft. It is, in his words, a "mission concept or spiritual concept."

The word got promoted and defined down in the same breath.

We should take the capabilities seriously anyway, because several of them are real and a few are new in kind.

Two zero-days found during an evaluation is not a leaderboard result, it is a finding. Ten open math problems with machine-checkable Lean proofs is not a benchmark, it is work. Reconstructing CAD programs from rendered views at 95.9% against Fable's 84.3% is a gap wide enough to survive any amount of methodology argument.

The computer use is the genuine step change. The team at Every let Astra spend hours inside Adobe Premiere cutting raw footage of their own Fable 5.1 review, and got back a usable first cut. Nothing else does that yet. It also does not appear on a leaderboard, which is why most of today's coverage skipped past it.

So the capabilities are not the problem. The framing is.

Three things the headline numbers do not say

Most of these are system scores, not model scores. The 98.6% on ARC-AGI-3 came from a harness that retains reasoning between turns and compacts long contexts. OpenAI itself published a piece earlier this year showing those two settings substantially raise ARC scores without changing the underlying model. The benchmark measured Astra plus OpenAI's agent system. The press release credited Astra.

FrontierMath runs at Epoch AI, which OpenAI funded and to part of which OpenAI has exclusive access. That does not make 97.6% false. It means the number is not independent, and OpenAI knows the difference.

The coding lead is inside the noise. DeepSWE v1.1 is 113 tasks. Astra's 74.1% against Sol's 70.8% is a real improvement for OpenAI. Against the field it is nothing. Meta reported 75.4% for Muse Spark 1.3 earlier this week, and the public board has Gemini 3.8 Flash and Claude Opus 5 at 74%. A 1.3-point spread on 113 tasks is one or two tasks. OpenAI's own chart excluded Muse and used a lower Fable number, which is a choice.

The alignment story got worse in a specific, disclosed way. Astra's chain of thought is harder to monitor than Sol's. OpenAI's explanation is that Astra has more control over its written reasoning and finishes problems in fewer written steps. Chief scientist Jakub Pachocki put it plainly: "Progress in intelligence does not guarantee progress in alignment." He added that OpenAI will withhold scaling until it regains enough confidence in monitoring.

Read that next to "welcome to the AGI era." Same day, same company, same document. One of those statements is marketing and one is engineering, and it is not hard to tell which.

The charts that answer the real question

Brockman said the thing we have been saying for a year: "The price per task is what matters."

He is right. A model that costs 8x per token and finishes in one pass beats a cheap model that needs four retries and a person to unpick the mess. Token price is a distraction. Cost to done is the only figure that goes in a budget.

The launch material had no task-level data to compute it. But according to BridgeMind, who posted screenshots on X, an OpenAI blog post went up briefly and came down, carrying four charts that plot exactly that. Accuracy against API cost. Score against output tokens.

We cannot verify the takedown, and the figures below are read off images rather than a data file, so treat them as close rather than exact. Here is what they show.

On Terminal-Bench 4.0, Astra reaches about 57.8% accuracy at roughly $7.30 in API cost. Fable 5.1 needs about $19.50 to reach roughly 56%. Higher accuracy for about a third of the money.

On DeepSWE v1.1, plotted against output tokens, Astra peaks near 0.738 at about 26K output tokens. Claude Opus 5 lands at roughly the same score around 90K tokens. Gemini 3.8 Flash gets there around 143K.

On the Artificial Analysis Coding Agent Index v1.4, Astra reaches about 67.3 index points by roughly $3.50. Claude Fable 5 peaks slightly higher, near 68.5, but needs about $8 to do it.

On FrontierCode 1.1 Extended, Astra tops out around 64.5%. Fable 5 is around 64.9%. Fable 5.1 sits lower, near 62%.

One caveat that matters, and it is not a small one: every chart here is a coding or terminal benchmark. We have no cost-per-task curves for writing, research, analysis, or computer use. Nobody has published them, for any model. So everything below is about code.

What the coding numbers actually say

Nobody won on capability this week.

Astra ties Fable and Opus 5 at the top of every coding chart in this set. Four labs sitting inside a couple of points of each other. The ranks shuffle depending on which benchmark you pick, and the shuffling is smaller than the measurement error.

Astra won on efficiency, and by a lot. Same ceiling for roughly a third of the token spend. That is the strongest argument for the model anywhere in today's material, and it is buried in charts that were up for a moment.

Which reframes the $10/$50 sticker price. Astra is expensive per token and cheap per task, at least for code. Anyone comparing list prices across labs is measuring the wrong thing, and OpenAI knows it, which is why every one of those charts has cost on the x-axis rather than a bar for each model.

There is a second finding in the same charts, and it points away from the marketing.

The curves are not monotonic. Astra's DeepSWE score peaks around 26K output tokens and comes back down, then runs flat out to roughly 57K. FrontierCode wobbles: about 63% at 12K, down near 62% at 15K, back up to 64.5% at 26K. On Terminal-Bench, the point after the peak costs roughly 40% more and scores slightly lower. On the Coding Agent Index, Astra is flat from about $3.50 to $5.

Four benchmarks. Four peaks before maximum. From the vendor's own data.

Every found the same thing by hand. They ran a coding assignment across all six effort settings, from low to ultra. Every version ran. High was the only setting with no bugs. Extra-high looked the best and had broken clicks. At ultra, on a separate task, Astra added an unrequested breathing-exercise feature to a 3D island scene, and the run records showed no contamination between tests. It was in Astra's own plan.

OpenAI benchmarked at maximum effort. Its own charts, and the only independent testing published today, both say the useful setting is somewhere in the middle. Any pipeline that defaults to turning the dial up should measure before it spends.

Where to reach for Astra anyway

The premium is defensible in specific places.

Long, tedious work inside a GUI where the output is easy to check. Video editing, form entry, website QA, moving data between applications with no API worth using. Hours of clicking that a person hates and a reviewer can verify in minutes. This is the one category where Astra does something the alternatives currently do not.

Visual and spatial generation from a brief. CAD reconstruction, 3D scenes, design-led prototypes. The BenchCAD gap is not close, and the qualitative reports agree with it. Expect to strip out the promotional copy it wraps around everything.

Expensive search with cheap verification. Formal proofs, or anything with a test suite that settles correctness. If checking the answer costs almost nothing and finding it costs a lot, pay for the better searcher. The ten math problems are the proof of concept.

Coding pipelines where cost to done is the constraint. This is the efficiency finding applied. Same ceiling as Fable at roughly a third of the tokens. Set the effort dial by measurement, not by instinct.

A first draft that takes direction. Every's CEO read a draft of their own review and did not realize a model had written it. That is a specific, falsifiable claim about a specific piece of writing, and more useful than any writing benchmark published today.

Vetted defensive security work, if you are inside Daybreak. If you are not, the gating works against you. OpenAI's Mia Glaese warned that users outside the trusted programs should expect slowdowns, pauses, and blocks during cyber work, and sometimes during unrelated work. Her words: "At launch, this is something that people should expect."

Where we would not reach for it: work that has to be right the first time without review. Every's Kieran Klaassen called Astra "a show horse, not a workhorse," and their testing backs it. Beautiful prototypes, weaker instincts for what the product needed. Fable built the journal-scanning app that captured a page per keystroke. Astra built the prettier one that took several clicks per page.

One practical warning before anyone starts testing. A Daybreak badge, a Daybreak API alias, and Daybreak-configured Astra are three different things. OpenAI's current mapping still points gpt-daybreak-blue-latest at GPT-5.6 Sol, and AWS maps Daybreak Blue to openai.gpt-daybreak-blue-5.6-sol. Neither points at Astra. Check the model name in the response, not the badge on the account, or you will benchmark Sol and publish it as Astra.

What came of it

We have not had hands on Astra. The test results here are OpenAI's, BridgeMind's screenshots, and Every's testing, and we have said which is which throughout. The operating advice is ours, from running agent pipelines in production every day.

The interesting thing about launch day is not the AGI claim. It is that the charts nobody was supposed to see show four labs clustered inside a couple of points on every coding benchmark, separated only by what they charge to get there.

That is not a frontier. That is a commodity market with a marketing department.

Set the effort dial by measurement. Compare on cost to done, not on list price. And ask any vendor claiming a capability lead to publish the x-axis.

Add VarOps on Google