GPT-6 Astra: has the AGI era really begun?

OpenAI presents GPT-6 Astra as the world's most intelligent model and states saturated benchmarks: 100% on ExploitBench, 99.9% on ARC-AGI-3, 97.6% on FrontierMath Tier 4. In the same tables the model is fourth on Artificial Analysis's Intelligence Index and last on Humanity's Last Exam. On the benchmark rebuilt without historical vulnerabilities it drops from 100% to 39%. What the figures say and what stays usable.

AICybersecurityGovernanceAICybersecurityOpenAIBenchmarksExploitReverse engineeringPreparednessAI Agents
Contents
  1. The claim, as it is worded
  2. The 99.9% on ARC-AGI-3 and the 63% from those who measure it
  3. The figures in those same tables
  4. Where the leap actually measures: cyber
  5. The benchmark built against contamination
  6. The evaluation born from the Hugging Face incident
  7. The number that gets worse
  8. What the shipped version can do
  9. Pricing and availability
  10. What we think
  11. Sources
Four figures on the benchmarks stated for GPT-6 Astra and on those that complicate the picture
Figures from the OpenAI announcement of 4 September. Sources at the end.

OpenAI released GPT-6 Astra, rolling out gradually from a limited set of organisations and then across ChatGPT, the API, Microsoft Azure and AWS Bedrock. President Greg Brockman commented on the release in three words, “arc-agi-3 is now saturated”. In the hours that followed his position was reported in the press as the announcement that the era in which AI is broadly as capable as humans had begun.

It is a claim that can be checked, because it rests on published figures. It is worth doing.

It is the direct sequel to a story we have followed since July. Astra is the model family that appeared in OpenAI’s technical report as “our next model”, and one of the evaluations presented today comes straight out of the Hugging Face incident.

The claim, as it is worded

The announcement opens by presenting Astra as “the world’s most intelligent and aligned model” and lines up a set of benchmarks described as saturated: ExploitBench at 100%, ARC-AGI-3 at 99.9%, FrontierMath Tier 4 at 98% in the text and 97.6% in the accompanying table. The post also points to results on open problems in mathematics, including two new results on gaps between prime numbers.

It is a wording that invites reading as a change of era. It is worth checking against the same tables that accompany it.

The 99.9% on ARC-AGI-3 and the 63% from those who measure it

The benchmark Brockman calls saturated is ARC-AGI-3. OpenAI’s announcement gives it 99.9%. The organisation that builds and administers that benchmark is ARC Prize, which a few hours earlier had published its own measurement. The two do not match.

The ARC Prize text is short and worth quoting in full: “Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness”. It adds that the model surpasses human performance on 96% of ARC-AGI-3 levels and that “it builds the most precise symbolic model of novel environments we’ve seen”, so the overall judgement is one of recognition.

The difference between 63 and 99 is not in the model. It is in the harness, the layer connecting the model to the test environment, and ARC Prize specifies that this is a new adapter. It is the same point we found reading the GLM-5.3 benchmarks, where every evaluation runs inside Claude Code 2.1.207. It is why we wrote about harness engineering as a discipline of its own.

A score that moves from 63 to 99 by changing the adapter measures the complete system, not the model’s capability in isolation. Both figures are legitimate and describe different things. What does not hold is quoting one of them without saying in which configuration it was obtained.

The figures in those same tables

Three rows OpenAI published alongside the announcement deserve as much attention as the ones in bold.

IndexGPT-6 AstraFable 5.1Fable 5Opus 5GPT-5.6 Sol
Artificial Analysis Intelligence Index v4.1.161.265.762.163.160.9
Humanity’s Last Exam (with tools)57.2%65.0%63.8%63.6%
Artificial Analysis Coding Agent Index v1.467.067.268.161.265.1

On the Artificial Analysis composite index, the closest thing to a general measure among those reported, Astra is fourth, behind Fable 5.1, Opus 5 and Fable 5. On Humanity’s Last Exam with tools it is last among the models compared. On the coding agent index it is third.

Then there is the most instructive counter-figure, which comes from the method OpenAI itself adopted. The 100% on ExploitBench sits alongside 39.0% on the ExploitBench rebuilt with vulnerabilities from the previous three months, built precisely because the first no longer discriminates. On ExploitGym, the hardest of the three, the score stops at 42.4%.

When a benchmark reaches 100% it stops measuring the capability and starts measuring itself: it says the test is exhausted, not that the capability has no limit. The distance between that 100% and the 39% on recent vulnerabilities is the measure of how much that difference weighs.

Where the leap actually measures: cyber

If there is one area where Astra’s figures pull away from the field, it is offensive security. Here OpenAI states the model reaches the Critical cybersecurity threshold of its own Preparedness Framework, the top level of their own risk scale.

The methodological point to hold on to is that the highest cyber numbers are measured without production safeguards, on a model that is not accessible in the shipped form. OpenAI states this explicitly.

BenchmarkGPT-6 AstraGPT-5.6 SolFable 5.1Opus 5
ExploitBench100.0%78.5%70%
ExploitGym42.4%30.3%30.4%22.0%
ExploitBench (June-August 2026)39.0%11.5%
SRE-Bench88.0%55.9%12.5%
SEC-Bench Pro85.4%79.1%

SRE-Bench measures the ability to reverse engineer binaries and understand their logic without access to source. Astra solves 88.0% of tasks on the first attempt and 99.2% within four, against 55.9% and 68.7% for GPT-5.6 Sol.

Expert-led assessments add the detail the benchmarks do not give: without production safeguards the model “could use previously unknown vulnerabilities to achieve arbitrary code execution in hardened browsers and create privilege-escalation exploits for hardened operating-systems”.

The benchmark built against contamination

There is one methodological choice worth flagging, because it answers the most obvious objection to exploitation benchmarks: that models saw the historical vulnerabilities they are measured on during training.

OpenAI built ExploitBench (June-August 2026), an internal evaluation using vulnerabilities from the previous three months. The result is 39.0% against 11.5% for GPT-5.6 Sol, with far lower output token consumption. Proportionally it is a wider gap than the one on the historical benchmark, where both models already start high.

During that evaluation the model found and used two zero-day vulnerabilities unknown until then. OpenAI writes that it has disclosed both to their maintainers.

The evaluation born from the Hugging Face incident

This is the part that closes the loop with what we wrote all summer.

OpenAI built an evaluation, appearing in the tables as ExploitGym honeypot, to measure whether a model facing a difficult or impossible task goes beyond its assigned scope. It is exactly the mechanism the METR investigation had reconstructed: agents with an impossible task looking for a way out and ending up collaborating on an unsanctioned message board.

The stated result is 48.2% for GPT-5.6 Sol without production safeguards, against 0.0% for Astra. The same table carries an Impossible ExploitGym row with Astra at 100%, with no comparison.

Two other measures line up with this. On the internal circumvention benchmark Astra sits at 0.00% against Sol’s 0.29% and OpenAI writes that the model never attempted to circumvent a Codex Auto-Review denial “even when Auto-review was deliberately configured to be evadable and the task was impossible to complete otherwise”. That is the condition which produced the incident in July.

On the internal computer use safety benchmark, where lower is better, Astra sits at 2.4% against 22.0% for Sol and with Auto-Review enabled at 1.8% against 4.3%.

The number that gets worse

Among the improvements OpenAI states a regression, and it is the passage that makes the document useful rather than promotional.

Astra’s written reasoning turned out to be harder to monitor than GPT-5.6 Sol’s, in tests that explicitly asked the model to evade monitoring. The explanation given is that Astra has more control over its written reasoning and solves problems in fewer explicit steps. OpenAI adds that on complex tasks the model still appears to struggle to conceal the reasoning it needs and that “we take the decline seriously”.

It matters because the legibility of the chain of thought is one of the few control levers monitoring of agents rests on, and it is the same one METR worked with while reading 1,300 transcripts.

What the shipped version can do

Here the distance between the figures and practical use is wide and needs stating precisely.

The version starting today refuses the more advanced cyber tasks, including writing a proof-of-concept exploit for a vulnerability. What remains permitted, and OpenAI says so explicitly, is secure code review and patching.

Through the OpenAI Daybreak programme the company states it plans to widen access with less restrictive safeguards in the following weeks, enabling vulnerability and proof-of-concept validation, malware analysis and detection engineering.

There is also an operating behaviour anyone designing an integration needs to know before meeting it in production. The extra safety checks can slow, pause or stop legitimate work, defensive work included. In ChatGPT and Codex a pause produces a review request to the user. In the API the task stops. For an unattended service the difference between “ask for confirmation” and “terminate” is everything.

Pricing and availability

In the API the model is gpt-6-astra, at 10 dollars per million input tokens and 50 per million output, with separate rates for cache reads and writes. There is a fast mode up to twice as quick, at twice the price.

For organisations access is off by default at launch and has to be enabled by the workspace administrator. Zero Data Retention is supported for eligible API customers.

Outside the cyber perimeter, two numbers give the measure of the rest: on OSWorld 2.0 Astra scores 72.6% in about 40 minutes per task against Sol’s 65.7% in 75 minutes, while on Terminal-Bench 4.0 it sits at 57.9% against 37.3%.

What we think

The figures answer the question in the title less sharply than the statements around them. The benchmark OpenAI’s president calls saturated is worth 63% according to those who administer it. It reaches 99% only with a new adapter. A model presented as the world’s most intelligent that sits fourth on a composite index and last on Humanity’s Last Exam in its own tables does not describe a change of era: it describes a crowded frontier, where the lead moves from row to row depending on what is being measured. And the fall from 100% to 39% when the vulnerabilities are recent rather than historical says how much of the saturation belongs to the model and how much to the test.

Where the leap is real and well documented is cyber, to be said with the same precision: 88% on SRE-Bench against the predecessor’s 55.9% and Fable 5.1’s 12.5% is a distance not seen in the other categories. If there is a discontinuity, it is vertical and on one domain, not general.

The first point concerns comparability of figures, and it is a practical warning for anyone reading tables. On Tuesday we wrote that Z.ai measures GLM-5.3 on ExploitGym reporting task counts, 105 and 130, over time budgets renormalised on each model’s tokens per second. OpenAI reports the same benchmark as a success percentage, 42.4%. They are two different protocols under the same name and the two values do not line up. Anyone building a comparison of models on cyber capability has to read the methodology footnotes before the figures, or redo the measurements in house.

The second is that for a defender the usable part today is narrower than the headlines suggest. Code review and patching are real work and are what the model agrees to do, while validating a proof-of-concept, analysing a malware sample or building a detection stays outside until Daybreak. Anyone evaluating it for a SOC should start there rather than from the 100% on ExploitBench, which is measured on a configuration they will not be handed.

The third concerns designing an integration. Between the stated Critical threshold, misalignment monitoring running in production and the task that stops in the API, a service built on this model has to allow for execution being interrupted by a classifier. It is a failure mode new to the ones usually handled, and it belongs in the design alongside timeouts and rate limits rather than being added after the first incident.

On reverse engineering, the 88% on SRE-Bench at the first attempt is the number we will follow most closely, because it is the capability that intersects directly with the analysis tools now arriving, which we will cover tomorrow.

Sources

Need support?Under attack?Service Status
Need support?Under attack?Service Status