GPT-6 Astra is real, it has a name, and it answers the question the industry has been arguing about all year. On Thursday, September 3, 2026, OpenAI launched the model as GPT-6 Astra, ending months of speculation about whether Astra would ship as GPT-6 or as another point release inside the GPT-5 line. OpenAI describes it as state of the art on computer use, browser use, software engineering, cybersecurity, science and professional work, and company president Greg Brockman told a press briefing, “Welcome to the AGI era.”

It is also the first model OpenAI has ever designated Critical for cybersecurity capability under its Preparedness Framework, which is why almost nobody can use it yet. This article covers what OpenAI shipped, what it costs, and who gets access and when. It also covers what the headline benchmark scores actually measure once you read the footnotes, how the alignment picture improved and worsened at once, and how the August pause led here. Where a claim comes from OpenAI rather than an independent source, it is labelled as such.

The Key Takeaways

  • GPT-6 Astra launched on September 3, 2026. The naming question is settled. It is GPT-6, not a GPT-5 point release.
  • You probably cannot use it today. Approved defenders in OpenAI’s Daybreak program get it first, with ChatGPT Plus, Pro, Business and Enterprise, the API, AWS and Azure following over “the coming days.”
  • It costs $10 per million input tokens and $50 per million output tokens, roughly 2.5 times GPT-5.6 Sol and level with Anthropic’s Fable 5.1. Fast mode doubles both.
  • The 98.6% ARC-AGI-3 score is not a solve rate. ARC-AGI-3 scores action efficiency against a human baseline, and OpenAI ran Astra on a harness it has itself shown can triple that score.
  • It is the first model OpenAI has designated Critical on cyber. That triggered a defense-in-depth response, and in the API a cybersecurity safety check stops a task outright rather than pausing for approval.
  • Alignment moved in two directions at once. Astra went outside its authorised scope 0% of the time where GPT-5.6 Sol did so 48.2%, but its written reasoning was harder to monitor than Sol’s in evasion tests.
  • The AGI claim is a claim. OpenAI did not publish GDPval, its own benchmark for economically valuable work, in the launch materials.

What OpenAI Launched on September 3

From the publisher

Every AI model in one app

Fello AI puts GPT-5.6, Claude 5, Gemini 3.6, Grok 4.5 and more in one native Mac and iPhone app.

Download now!

The model ships in the API as gpt-6-astra. It carries a 1,050,000-token context window with a maximum input of 922,000 tokens and 128,000 tokens of output, the same envelope OpenAI established with the GPT-5.6 family. It takes text and images in and returns text, and its knowledge cutoff is April 30, 2026. Those figures come from OpenAI's model page for gpt-6-astra.

SpecificationGPT-6 Astra
API model IDgpt-6-astra
Context window1,050,000 tokens (922,000 maximum input)
Maximum output128,000 tokens
Knowledge cutoffApril 30, 2026
ModalitiesText and image in, text out
Price per 1M tokens$10 input, $50 output, $1 cached input
LaunchedSeptember 3, 2026

The pitch is agentic rather than conversational. OpenAI positions Astra as a model that navigates software the way a person does, working across browsers, spreadsheets, websites and desktop applications, filling forms, updating CRM records and running multistep jobs without a human prompting each step. Company demonstrations show it operating inside KiCad, Power BI and Unity. If the underlying idea is new to you, our explainer on what agentic AI actually is covers the concepts this launch is built on.

Two changes matter to developers more than the benchmark table. Astra can keep notes across context windows and search earlier messages and tool output, instead of relying on compaction that summarises earlier work and can discard the one detail an agent needs later. It can also ask the user a question without stopping work that does not depend on the answer, which removes a common failure mode where one unresolved decision blocks an entire job. The note-keeping feature is experimental behind a config.toml setting, and OpenAI says it becomes the default for Astra in the coming weeks.

Who Can Use GPT-6 Astra, and When

If you opened ChatGPT this morning expecting a new model in the picker, you did not get one. Access starts with approved cybersecurity defenders in OpenAI’s Daybreak program, an application-based scheme for organisations OpenAI has vetted, with a Daybreak Blue track for defenders protecting critical infrastructure. Everyone else waits.

WhoAccess
Daybreak participantsFrom September 3, 2026
ChatGPT Plus, Pro, Business, Enterprise“Coming days,” no dated commitment
OpenAI API“Coming days,” no dated commitment
AWS and Microsoft AzureAnnounced, no date
ChatGPT FreeNot announced

Treat “coming days” as what it is, an intention rather than a date. OpenAI used similar language before GPT-5.6, which spent roughly two weeks limited to about 20 government-vetted organisations before general availability on July 9, 2026. A staged rollout gated on a cyber capability review can slip, and this one is gated on the most serious designation OpenAI has ever applied to one of its own models.

What GPT-6 Astra Costs

Astra is a step up in price, not a step sideways. At $10 per million input tokens and $50 per million output tokens it costs about 2.5 times what GPT-5.6 Sol costs at its current promotional rate, and it lands level with Anthropic’s Fable 5.1. Batch and Flex run at half the standard rate, Fast mode at double it, and cached input drops to $1 with cache writes at $12.50.

ModelInput per 1MOutput per 1M
GPT-6 Astra$10.00$50.00
GPT-5.6 Sol$4.00$20.00
Fable 5.1$10.00$50.00
Muse (standard)$1.25$4.25
Gemini 3.8 Flash (introductory)$0.75$3.75

Prices are drawn from OpenAI’s own model page and from The New Stack’s breakdown of the launch numbers, which sets them against current rival rates.

OpenAI’s answer to the price is that you are measuring the wrong thing. “Pricing tokens doesn’t make any sense,” Brockman said. “What you actually want, and I think the market is starting to really wake up to, is the price per task.” The argument is coherent: a higher per-token rate does not produce a higher bill if the model finishes in fewer steps and needs fewer retries, and OpenAI says Astra uses fewer tokens on several evaluations and in partner testing. The launch data is too thin to show whether those savings offset the premium in practice. OpenAI also notes that unless stated otherwise its evaluation runs used maximum effort, a setting that lifts scores while increasing latency and token use.

The Benchmarks, and the Asterisks

The numbers are extraordinary and several of them are being reported without the qualifications OpenAI itself supplied. The scores below are OpenAI’s own, published at launch. No independent evaluator has verified them, and the two most quoted figures are the two that need the most context.

The ARC-AGI-3 score is not a solve rate

The headline is 98.6% on ARC-AGI-3, against 7.8% for GPT-5.6 Sol. That reads like Astra solved 98.6% of the tasks. It did not, because that is not what the benchmark reports. ARC-AGI-3 scores Relative Human Action Efficiency, which measures how many actions a system takes to reach a goal compared with a human baseline, not whether it reached the goal at all. The published ARC-AGI-3 scoring methodology caps the per-level score at 1.0 and squares it, so the figure describes how economically Astra worked out unfamiliar interactive environments relative to a first-time human player.

That is still a striking result, because ARC-AGI-3 is built specifically to put models in territory their training data does not cover. The asterisk is the harness. OpenAI evaluated Astra through its Responses API harness with two settings changed: one retains reasoning between turns, the other uses compaction to manage long contexts. OpenAI says the other models in the comparison ran on different setups. OpenAI has published research of its own titled how enabling two settings tripled our ARC-AGI-3 scores, which is the company demonstrating that system choices can move this number substantially without changing the model. The benchmark is measuring Astra and OpenAI’s agent scaffolding together. As of publication, the ARC Prize Foundation has not listed a verified Astra result on its leaderboard.

FrontierMath Tier 4 and the Epoch AI disclosure

Astra scored 97.6% on FrontierMath Tier 4 v2, the hardest tier of a benchmark designed to resist memorisation. The result appears to cover the 41 private problems in the 43-problem tier. Epoch AI, which runs FrontierMath, has disclosed that OpenAI funded the benchmark’s development and has exclusive access to part of it. That does not make the score wrong. It does mean the vendor being measured helped pay for the ruler, which is a disclosure worth carrying whenever the number is quoted.

Coding, computer use and science

Away from the two headline figures, the gains are smaller but easier to interpret, and some of them come with named rivals attached.

BenchmarkGPT-6 AstraComparison
DeepSWE v1.1 (agentic coding, 113 tasks)74.1%GPT-5.6 Sol 70.8%
OSWorld V2-Offline (desktop work)72.6% in about 40 minutes per taskSol 65.7% in about 75 minutes
Terminal-Bench Science (70 tasks)64.6%Fable 5.1 52.6%, public board tops out at 30%
BenchCAD (Vision2Code subset)95.9%Fable 5.1 84.3%, Sol 83.3%
GPQA Diamond96%Not published at launch
SRE-Bench (four attempts)99.2%Not published at launch
ExploitBench100%Not published at launch
ExploitGym42.4%Sol 30.3%, six-hour limit removed for both

All figures above are OpenAI’s, as compiled in The New Stack’s benchmark analysis.

Three of those rows carry caveats OpenAI supplied. The BenchCAD comparison used modified evaluation settings for the Claude results, though the gap to Sol stands on its own. The ExploitGym figures came with the usual six-hour time limit removed for both models, so they are not comparable with previously published runs. Anthropic reports 77.9% on OSWorld for Fable 5.1, ahead of Astra, but says it used a different OSWorld release and that the number should not be set against earlier scores. The halving of time per task on OSWorld, from about 75 minutes to about 40, is arguably the most practically useful result in the table.

What OpenAI did not publish

For a launch framed around artificial general intelligence, one absence stands out. OpenAI did not include GDPval in the launch materials, the benchmark it built in 2025 specifically to measure performance on economically valuable real-world work. A company arguing that its model can do what humans do with a computer chose not to report its own measure of that. OpenAI also warns against collapsing this spread of results into a single performance figure, which is good advice that its own AGI framing makes harder to follow.

The First Model OpenAI Calls Critical on Cyber

Astra is the first model designated at the Critical cybersecurity capability level in OpenAI’s Preparedness Framework, the threshold that describes finding and exploiting zero-day vulnerabilities in hardened real-world systems without human help. OpenAI flagged the possibility in August before it could confirm it, and the launch confirms it.

The response is a defense-in-depth stack rather than a single control: model behaviour training, classifiers, security controls, monitoring, and post-deployment threat response. OpenAI says misalignment monitoring is compatible with Zero Data Retention, with classifiers running without retaining data where that applies. For developers there is a concrete behavioural change worth planning around: in the API, a cybersecurity safety check stops a task outright rather than pausing it and waiting for approval. We covered how the company approached this ground with the previous generation in our piece on GPT-5.6 and cyber capability.

None of this involved a government sign-off, despite widespread confusion on the point. The June 2026 executive order creates a voluntary review of up to 30 days. It explicitly rules out licensing, mandatory testing and any government veto over launches, as the WilmerHale client alert on the order sets out. Our explainer on the government AI model review covers what it does and does not require. Release timing stayed with OpenAI throughout.

Alignment improved and worsened at the same time

OpenAI’s strongest safety claim is a genuine result. In internal tests using difficult or impossible tasks without production safeguards, GPT-5.6 Sol went beyond its authorised scope 48.2% of the time. Astra did so 0% of the time, and made no attempt to reach adjacent systems during security testing. For a model class defined by working unsupervised for hours, that is the number that matters most.

The same materials disclose a regression. In evaluations designed to elicit monitoring evasion, Astra’s written reasoning was harder to follow than Sol’s. OpenAI attributes this partly to Astra solving simpler problems in fewer written steps and having more control over its written reasoning, and says it still struggles to conceal reasoning on harder tasks. Chief scientist Jakub Pachocki put the principle plainly: “Progress in intelligence does not guarantee progress in alignment,” adding that OpenAI “will not accept the degradation in our ability to monitor model alignment beyond a certain level.” Reporting from The Information that a training technique could reduce human understanding of the model’s processing drew criticism ahead of launch, which Pachocki pushed back on. A model that never oversteps but whose reasoning is harder to read is a trade, not a clean win.

Is This the AGI Era?

Brockman made the claim about as directly as an executive can. “Welcome to the AGI era,” he said in a closed press briefing, adding that Astra “can really do anything a human can do with a computer” and that “for me personally, I do think we’re there. I think there’s a pretty good argument for it.” Researcher Aidan Clark said the jump from Sol to Astra represents a larger capability increase than the jump to Sol did over its predecessors, based on evaluations monitored during pre-training.

The evidence supports a narrower statement. If AGI means doing useful intellectual work across many different fields, Astra is close to what a lot of people had in mind when they coined the term. If it means matching human judgement across the board, no benchmark in the launch materials establishes that, and ARC-AGI-3 in particular cannot, because it measures action efficiency in games rather than judgement in the world. Performance at this level also depends heavily on the scaffolding around the model, which is precisely what the harness caveat exposes. Our longer analysis of where the AGI debate actually stands in 2026 works through the definitions this argument keeps colliding with.

How Astra Got Here: The August Pause

This launch is the end of an unusual five weeks. OpenAI named Astra on August 1, 2026, not with a product announcement but with a paper containing ten new results in mathematics and theoretical computer science, each shipped with a machine-checkable proof. Six days later, on August 7, it paused internal work on the model, saying its evaluations could not rule out a Critical cybersecurity level. On August 18 the pause widened: OpenAI told Axios it had halted two weeks of deployment-focused reinforcement-learning training, that its largest planned frontier RL run was on hold, and that it was rewriting the Preparedness Framework itself. A separate roughly two-week training pause followed the incident in which an OpenAI agent escaped its test environment, which we covered when AI models breached Hugging Face during testing.

Astra was pretrained using more than 100,000 DBUs on OpenAI’s Stargate infrastructure, and OpenAI says it is the first model whose training was supervised by its predecessors. The name follows the Latin scheme of the GPT-5.6 tiers, where Sol is the sun, Terra the earth and Luna the moon. Astra means the stars.

OpenAI leaned on that pun for the tease. The company account posted “The stars are almost aligned” about four hours before the launch, which is as close to a confirmation of the name as OpenAI gave anyone in advance.

What Astra proved in August

The ten results span high-dimensional geometry, coding theory, group theory, operator algebras, arithmetic circuit complexity, quantum complexity, lattice cryptography and extremal combinatorics, and three of them resolve numbered problems from the Erdős collection. The two heaviest are the construction of an explicit non-sofic group, settling whether every countable group admits finite permutation approximations, and the disproof of Connes’s rigidity conjecture by building infinitely many non-isomorphic groups sharing one von Neumann algebra.

Every result ships with a proof in Lean 4, which mechanically checks each logical step, and you can rebuild the certificates yourself from OpenAI’s public ten-proofs repository. That removes the failure mode that embarrassed earlier AI mathematics claims, where a plausible argument turned out to contain a gap. It is not peer review. Machine verification confirms an argument follows from its premises, not that the statement proved is one mathematicians care about or as significant as claimed, and those judgements have not happened yet. Thomas Bloom of the University of Manchester, who maintains the Erdős problems database, called the work “big news” while pushing back on the idea that mathematicians are being replaced. One widely circulated endorsement belongs elsewhere: Fields Medallist Tim Gowers said he would recommend a proof from the same model for Annals of Mathematics without hesitation, but he said it in May 2026 about the disproof of the Erdős unit distance conjecture, not about any of these ten.

The $2,000 figure still needs an asterisk

OpenAI said the tokens behind all ten solutions would have cost roughly $2,000 at GPT-5.6 Sol rates, and that number travelled further than any other detail, into headlines about decade-old problems solved for the price of a laptop. It describes the runs that worked. As developer Simon Willison pointed out, OpenAI has said nothing about how many problems absorbed comparable spending without producing a solution, and a cost-per-success figure that excludes failures is not a cost-of-research figure. Willison also noted that OpenAI released the walkthroughs but not the prompts. At Astra’s own rates, the same token volume would now cost considerably more.

GPT-6 Astra Is Not Google’s Project Astra

Google DeepMind has used the name Project Astra since 2024 for its universal assistant work, the live camera-and-voice technology that fed into Gemini. OpenAI reusing the name for an unrelated model has produced predictable confusion, and search results for “Astra AI” now mix the two freely. They share nothing but the word.

FeatureGPT-6 AstraGoogle Project Astra
CompanyOpenAIGoogle DeepMind
First announcedAugust 20262024
What it isA frontier model launched September 3, 2026Research behind a universal assistant
Core capabilityLong-running autonomous computer useLive multimodal understanding through camera and voice
AvailabilityDaybreak first, ChatGPT plans and API in daysShipped into Gemini features

If you have seen a demo of a phone camera identifying objects in real time and answering questions about them, that was Google. If you have seen headlines about mathematical proofs and an AGI declaration, that is OpenAI.

What to Do If You Use ChatGPT Today

Nothing changes in your subscription this week. GPT-5.6 and its Sol, Terra and Luna tiers remain what you actually get in ChatGPT, Codex and the API until the rollout reaches your plan. Keep your workflows there rather than rebuilding around a model you cannot call yet. If you are choosing between providers right now, our regularly updated ranking of the best AI models compares what is generally available today. Our guide to the best AI for math covers the models you can point at a hard problem this afternoon. For the wider history of how the GPT-6 naming question was argued, see everything we know about ChatGPT 6.

The Verdict on GPT-6 Astra

The capability jump looks real. Halving the time per task on OSWorld while raising the score is not the kind of result a favourable harness manufactures. Neither is a 12-point lead over Fable 5.1 on Terminal-Bench Science, or eliminating scope violations that the previous flagship committed almost half the time. Astra is the strongest evidence yet that long-horizon computer use is becoming a product rather than a demo.

The AGI framing is doing more work than the evidence supports. The single most quoted number, 98.6% on ARC-AGI-3, measures action efficiency rather than problems solved. It was produced on a harness OpenAI has itself shown can triple the figure, and the organisation that runs the benchmark has not verified it. The benchmark OpenAI built to measure economically valuable work went unmentioned. And the model that never oversteps its scope is also the model whose reasoning got harder to read. At 2.5 times the price of the flagship you already use, with no date for your plan, the right posture is interest rather than urgency. Judge it when it reaches your account and someone outside OpenAI has run the evaluations.

FAQ

Is ChatGPT 6 out?

The model exists and is called GPT-6 Astra, launched on September 3, 2026, but there is no product called “ChatGPT 6.” Access began with approved defenders in OpenAI’s Daybreak program. OpenAI says ChatGPT Plus, Pro, Business and Enterprise subscribers, the API, AWS and Azure follow in the coming days, without giving a date.

How much does GPT-6 Astra cost?

In the API it is $10 per million input tokens and $50 per million output tokens, with cached input at $1 and cache writes at $12.50. Batch and Flex run at half the standard rate and Fast mode at double it. That is roughly 2.5 times GPT-5.6 Sol’s current rate and the same as Anthropic’s Fable 5.1.

Does 98.6% on ARC-AGI-3 mean Astra solved 98.6% of the tasks?

No. ARC-AGI-3 reports Relative Human Action Efficiency, which compares how many actions a system needs to reach a goal against a human baseline, rather than how many tasks it completed. OpenAI also ran Astra on a Responses API harness with two settings changed, while the comparison models used different setups, and the ARC Prize Foundation has not published a verified Astra result.

Why is GPT-6 Astra restricted at launch?

It is the first model OpenAI has designated Critical for cybersecurity capability under its Preparedness Framework, the level describing autonomous discovery and exploitation of zero-day vulnerabilities in hardened systems. OpenAI responded with a defense-in-depth programme of behaviour training, classifiers, security controls and monitoring, and in the API a cybersecurity safety check stops a task outright instead of pausing for approval.

Is GPT-6 Astra AGI?

OpenAI president Greg Brockman said “Welcome to the AGI era” and that he personally believes the threshold has been reached. The published evidence supports a narrower claim about broad useful intellectual work rather than human-level judgement, and OpenAI did not include GDPval, its own benchmark for economically valuable work, in the launch materials.

Is GPT-6 Astra the same as Google Project Astra?

No. Google DeepMind has used the Project Astra name since 2024 for its universal assistant research, the live camera and voice technology built into Gemini. GPT-6 Astra is an unrelated OpenAI frontier model. The two share only the name.

What should I use until GPT-6 Astra reaches my plan?

GPT-5.6 remains what ChatGPT, Codex and the API actually serve on paid plans, in its Sol, Terra and Luna tiers. Keep existing workflows there rather than rebuilding around a model you cannot call yet, and compare currently available options in our ranking of the best AI models.