The Pulse
OpenAI’s GPT-6 Astra Scores Face Scrutiny After Quiet Revisions
OpenAI changed several GPT-6 Astra benchmark figures after its September 3 launch, including hallucination, cybersecurity and rival-model scores. Independent ARC Prize testing also showed a wide gap between Astra’s 99.9% result under OpenAI

AI.info Team ·
OpenAI launched GPT-6 Astra on September 3 with a set of numbers designed to establish a new performance ceiling. Within hours, several of those numbers changed. One hallucination rate was cut from 4.2% to 2% before returning to 4.2%; a cybersecurity result for GPT-5.6 Sol doubled and then came under review; and scores for Anthropic’s models moved down and back up across successive versions of the company’s launch page.
OpenAI says the edits reflect ordinary evaluation work rather than an attempt to mislead customers. Independent testing has added a second problem: the model’s headline 99.9% score on ARC-AGI-3 depends on a provider-specific harness that preserves hidden reasoning state between requests. Under ARC Prize’s standard harness, the best verified result is 62.7%, a difference large enough to change how the claim should be read.
The dispute does not show that Astra lacks strong capabilities. OpenAI’s own published table records substantial gains over GPT-5.6 Sol across coding, computer use, mathematics and cybersecurity. It does show why launch-day benchmark tables need dates, configurations and clear revision histories before buyers use them to compare models.
OpenAI’s launch table changed before the announcement settled
OpenAI originally planned to publish its GPT-6 Astra announcement at 2 p.m. Eastern time on September 3. The post appeared briefly, disappeared, and did not become consistently accessible until later in the afternoon. Fortune reviewed archived versions of the page and found that several evaluation figures changed between the first accessible snapshot and the version that followed.
The most visible change involved an internal hallucination benchmark. The first archived copy listed Astra at 4.2% and GPT-5.6 Sol at 12.2%. A later copy reduced those figures to 2% and 9.4%. The page then returned to the original 4.2% and 12.2% values. OpenAI’s current launch page does not explain that sequence in a dated correction log.
OpenAI defended the revisions in comments to Fortune. “We care deeply about getting evaluations right,” an OpenAI spokesperson told the publication. “Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons.”
That explanation accounts for why a result might change. It does not resolve the disclosure problem. A customer reading the page after the edits cannot easily tell which numbers changed, when they changed or why the company selected the final version.
The ExploitBench comparison used a setting customers cannot buy
The cybersecurity figures create a more specific methodological dispute. OpenAI’s first version of the launch material listed GPT-5.6 Sol at 5.5% on an internal version of ExploitBench. Later versions raised Sol’s score to 11.5%. OpenAI then said it was examining whether to return the result to 5.5% because the 11.5% figure came from a reasoning level that is not commercially available for Sol.
That distinction matters because benchmark comparisons are only useful when the tested systems operate under comparable conditions. OpenAI’s public table lists Astra at 100% on ExploitBench and Sol at 78.5%, while a separate recent-data version of ExploitBench gives Astra 39% and Sol 5.5%. The company says the newer evaluation uses vulnerabilities from June through August 2026 rather than older known flaws, making the two tests measure related but different capabilities.
OpenAI also reports that Astra meets its highest cybersecurity capability category under the company’s Preparedness Framework. The company says Astra can find previously unknown flaws and develop exploits across well-protected systems when given the right tools and access. OpenAI delayed parts of Astra’s development and release after a July incident involving other models and the AI company Hugging Face, then restricted access to some advanced cyber capabilities.
The safety claims and the benchmark revisions are separate questions. A model can be highly capable at cyber operations while a comparison table still uses an inappropriate baseline. Buyers need to know whether a score describes the model available in an API, a private research checkpoint, a special reasoning tier or a larger agent system wrapped around the model.
ARC-AGI-3 shows how the harness changes the result
OpenAI’s most striking public claim is Astra’s 99.9% score on ARC-AGI-3, an interactive benchmark created by the ARC Prize Foundation. OpenAI’s launch page calls the result a saturation point and presents it alongside a statement that Astra is state of the art in reasoning, computer use, software engineering, cybersecurity, science and professional work.
ARC Prize’s own result page provides the missing detail. It records Astra’s best result at 62.7% under the standard harness, with a maximum-reasoning run costing $26,098. Under the provider adapter harness, which preserves opaque reasoning state between requests and uses compaction for longer conversations, Astra reached 99.9% at high reasoning for $18,817.
Those are not two measurements of exactly the same system. The standard harness allows the model to carry forward notes it chooses to keep. The provider adapter preserves additional internal state supplied through OpenAI’s interface. OpenAI says the adapter changes two settings to better reflect real-world performance and that the changes do not specifically target ARC-AGI-3.
ARC Prize still described Astra’s result as a major improvement. Its evaluation found that Astra used fewer actions than the median human baseline on 96% of levels. The organization has also said that a saturated ARC-AGI-3 score should not be treated as proof that a model has achieved artificial general intelligence.
OpenAI’s launch page places the 99.9% number in the main benchmark table, while the harness distinction appears in a footnote. That presentation is the source of much of the criticism. A reader scanning the table sees a near-perfect result next to low single-digit scores for earlier models, but may miss that Astra and those models were not necessarily evaluated through the same interface.
Independent figures still show a strong Astra model
The scrutiny does not reduce Astra’s performance to a trick of presentation. ARC Prize’s standard-harness score of 62.7% remains far above the public results listed for other systems. OpenAI’s table also shows Astra ahead of GPT-5.6 Sol on OSWorld 2.0, at 72.6% versus 65.7%, and on Terminal-Bench 4.0, at 57.9% versus 37.3%.
In OpenAI’s academic evaluations, Astra scores 97.6% on FrontierMath Tier 4, compared with 83% for GPT-5.6 Sol. It scores 96% on GPQA Diamond and 57.2% on Humanity’s Last Exam with tools. On computer-use tests, it reaches 59.3% on Agents’ Last Exam and 92.7% on ScreenSpot-Pro without tools.
The same table also shows that Astra does not win every comparison. Anthropic’s Fable 5.1 scores 65% on Humanity’s Last Exam with tools, above Astra’s 57.2%. Fable 5.1 also scores 65.7 on Artificial Analysis’s Intelligence Index, compared with 61.2 for Astra in the version OpenAI cited. On FrontierCode 1.1 Main, Fable 5.1 records 50.9%, while Astra records 53.3%; on the extended version, Fable 5.1 records 63.6% against Astra’s 64.5%.
Those mixed results are one reason the revisions matter. The issue is not whether Astra is capable. The issue is whether customers can distinguish a broad performance gain from a favorable test configuration, a private checkpoint or a number that changed after publication.
Researchers want a record of every benchmark change
Fortune quoted Anka Reuel and Mike Hardy, researchers at Stanford’s Intelligent Systems Laboratory and Stanford Trustworthy AI Lab, questioning whether repeated reruns amounted to “benchmaxxing,” a term for optimizing evaluation conditions to produce the highest score. They also said Astra’s system card provides few details about the internal hallucination benchmark, including no test-item count.
Vincent Sunn Chen, an AI engineer who leads benchmark and evaluation research at Snorkel AI, offered a less accusatory explanation while calling for clearer reporting. “A benchmark score reflects a specific measurement setup: the model checkpoint, configuration (including how much time and compute the model is allowed), harness, eval/grading configuration,” he told Fortune in an email.
Chen said scores can shift during the final hours before a launch because teams are still changing checkpoints, compute limits and grading systems. He also said companies should report what changed when they revise benchmark results. That standard would allow readers to separate a corrected scoring error from a new model run, a different reasoning tier or a more favorable agent scaffold.
OpenAI’s launch page does include footnotes describing several limitations. It says GPT-5.6 Sol’s 5.5% ExploitBench result may reflect a 300-turn limit that does not apply to customers using maximum settings. It also warns that third-party models are tested through simpler research setups and that provider safeguards and computer-tool implementations differ.
Those caveats are technically meaningful, but they arrive after headline percentages that invite direct comparison. The result is a table that can be accurate in each individual cell while still encouraging an inaccurate overall impression.
What customers can verify after September 3
For customers assessing Astra, the most defensible figures are the ones tied to a named harness, a stated reasoning level and an accessible model configuration. ARC Prize’s standard-harness result provides the clearest comparison point for ARC-AGI-3, while the 99.9% provider-adapter result describes what Astra can do through a more specialized OpenAI interface.
OpenAI’s current API documentation lists GPT-6 Astra as the company’s most capable model for complex reasoning, coding, computer use, research and document creation. It gives the model a one-million-token context window, a maximum output of 128,000 tokens and standard API pricing of $10 per million input tokens and $50 per million output tokens. Those specifications matter more to procurement teams than a single launch-day score because they determine how the model behaves in ordinary workloads and what it costs to operate.
The company has not been shown to have fabricated Astra’s results. The documented problem is narrower and more concrete: benchmark figures changed repeatedly around publication, some changes temporarily improved Astra’s standing against rivals, and the most dramatic ARC-AGI-3 result depends on a provider-specific evaluation setup. Until OpenAI publishes a dated revision history with full test conditions, the 99.9% figure should be read as a result for Astra plus OpenAI’s adapter, not as a universal score for the model in every environment.