What GPT-6 Astra’s 99.9% ARC-AGI-3 Score Actually Measures
GPT-6 Astra’s ARC-AGI-3 score depends entirely on which harness ran it. Here is what changed after launch, who benefits, and what still needs checking.
OpenAI called it the arrival of the AGI era. GPT-6 Astra’s ARC-AGI-3 score is the number carrying that claim through every headline this month, and the organization that built the benchmark will not stand behind the conclusion OpenAI drew from it. That gap is not a rounding error, and it is not a rival’s complaint. It is documented, on the record, in the benchmark’s own published table, by the people whose professional reputation depends on that table meaning something.
Same weights. Different scaffolding. Different score. That is the whole story in one sentence, and it is worth sitting with before the details arrive, because almost nobody who shared the 99.9% figure sat with it for even that long.
The Same Model, Two Very Different Scores
ARC Prize, the nonprofit behind ARC-AGI-3, ran Astra through its benchmark twice and published both results on launch day, September 3, 2026. Under its own Standard harness, the software wrapper that gives every model the same minimal interface and forces it to decide for itself what to carry forward between moves, Astra scored 62.7% at maximum reasoning, for $26,098 in compute. Under OpenAI’s own Provider Adapter harness, which preserves the model’s private reasoning state between requests and compresses longer conversations automatically, the same model scored 99.9%, for $18,817. The better score was also the cheaper one, which tells you the adapter was not just easier on the model, it was doing real work the Standard harness withheld.
One further number makes the point harder than any argument could. Set the Provider Adapter’s reasoning effort to none, the lowest setting available, and Astra still scores 96.7%. That beats the model’s own best result on the Standard harness, at maximum reasoning, by 34 points. The dial that is supposed to control how hard the model thinks barely moved the outcome. The wrapper around it moved the outcome by more than thirty points on its own.
On the 167 game-reasoning pairs both harnesses solved, the Provider Adapter ran roughly 3.66 times faster and used 49% fewer tokens. None of that is a secret. ARC Prize printed a full table, the same day, with every reasoning level it tested. The number that spread everywhere was 99.9% against GPT-5.6 Sol’s 7.8%, and those two figures were never run under the same conditions: Astra’s came from the adapter, Sol’s from the Standard harness. Compare like with like and the gap is 62.7% to 7.8%, still an enormous jump, and the one ARC Prize will actually defend.
Even the benchmark’s own people cannot agree on the honest number. Co-founder François Chollet posted that Astra scored 66% on the Standard harness. The published table says 62.7%. A separate results sheet, reportedly submitted to ARC Prize on September 2, lists the adapter score as 99.95%, not 99.9%. OpenAI’s own launch post displayed 99.99%. Four sources, three of them belonging to the same two organizations, and not one of them matches another past the second decimal. Honestly, if the number that supposedly proves a step toward general intelligence cannot survive being copied from one document to the next, that is worth pausing on longer than the number itself.
ARC Prize’s own verdict is plainer than anything OpenAI put in a chart. The foundation said directly that it is not claiming Astra is AGI, and co-founder Mike Knoop wrote that “we lack evidence to call this AGI yet.” Going forward, ARC Prize will publish both harness results side by side, every time, for every model. That is the correction a scientific instrument makes when it discovers people are reading it wrong. It is not the correction a marketing claim survives.
A Number That Kept Moving After It Was Published
OpenAI’s launch post was supposed to go live at 2:00 p.m. Eastern on September 3. It did not fully load for almost two hours, through a tweeted link that returned an error and a company statement blaming a technical snag. OpenAI later pulled the published post entirely and put it back up, and declined to say why.
ARC Prize also discloses what it pays its human test subjects for comparison purposes, a detail with nothing to do with any of this: $115 for a 90-minute session, plus $5 per game finished, which works out to roughly $12.78 per attempted game before bonuses. That is oddly close to what a food delivery app pays per drop-off in most American cities, which is either a coincidence or a sign that gig pricing has become the default unit for any task nobody has figured out how to value yet, benchmark testing included. Fortune reporter Emily Forlini did something almost nobody else bothered to do: she compared archived snapshots of the page, taken automatically as it loaded, against each other. By the time anyone noticed the number had changed, the original screenshot was already three group chats deep and impossible to recall. Five metrics had changed between the first snapshot and the sixth. Astra’s stated hallucination rate read 4.2% for the first five snapshots, then dropped to 2% in the sixth, taken at 5:20 p.m., about ten minutes before OpenAI’s official tweet went out with the final version. As of Fortune’s reporting, it had reverted back to 4.2%. GPT-5.6 Sol’s hallucination figure moved the same way, from 12.2% down to 9.4% and back. Sol’s score on an internal cybersecurity evaluation, ExploitBench, roughly doubled from 5.5% to 11.5%, and OpenAI told Fortune it is investigating reverting that change, since the higher number reflects a reasoning tier Sol does not actually offer to paying customers. Not every edit flattered OpenAI. Two of Anthropic’s own scores on HealthBench Professional went up during the same window, and Anthropic’s Fable 5.1 saw its FrontierMath score swing from 87.8% down to 78%, then settle at 83%.
An OpenAI spokesperson told Fortune the company cares deeply about getting evaluations right, and that most evaluations carry noise of a few percentage points depending on checkpoint, scaffold, and evaluation run. That is true, and it is also the kind of true statement that explains almost nothing about why the noise always seemed to land in the same direction the first time anyone looked.
Two Stanford researchers, Anka Reuel of the Intelligent Systems Laboratory and Mike Hardy of the Trustworthy AI Lab, have a word for this: benchmaxxing, the practice of re-running evaluations under shifting conditions until a number improves. They went looking in the one place that is supposed to document exactly how an evaluation was run, Astra’s own system card, and found barely any detail on the internal hallucination test. It does not even list how many items were in it. A launch chart with two decimal places of precision, backed by a technical document that will not say how many questions it actually asked, is not a contradiction anyone needs a statistician to spot.
This Has Happened Before, and Everyone Remembers How It Ended
Meta ran this exact play in April 2025, and the consequences are a matter of public record now, not speculation. LMArena, the crowdsourced benchmark Meta used to promote Llama 4, updated its own policies within days of the launch specifically because, in the platform’s words, Meta’s interpretation of the submission rules did not match what the platform expected from model providers. What Meta had actually submitted for testing was a specially tuned chat variant, not the version anyone could download and run.
Nine months of denial followed. Then, in a January 2026 interview with the Financial Times, Yann LeCun, on his way out the door as Meta’s chief AI scientist, confirmed it directly: the results were “fudged a little bit.” His team had tested different versions of the same model against different benchmarks and reported whichever version scored best on each one, assembled into a single table that no single model had actually produced. Mark Zuckerberg lost confidence in the entire team behind the launch and sidelined it. Meta spent between $14.3 and $15 billion on a 49% stake in Scale AI, and installed Scale’s CEO, Alexandr Wang, to run a new Meta Superintelligence Labs. Ahmad Al-Dahle, the AI lead whose team had done the fudging, left for a CTO role at Airbnb roughly a year after the scandal broke.
That is what happens when the gap between a published number and the truth gets confirmed by someone inside the building. It costs a division, several careers, and billions of dollars in a panic acquisition. Silicon Valley calls the fallout a leadership reset. Everyone who lived through it calls it getting caught.
None of that means Astra’s numbers are fabricated the way Llama 4’s were. Nobody at OpenAI has confirmed picking different model checkpoints to flatter different tests, and the adapter OpenAI used is documented, callable through the standard API, and available to any developer willing to pay for it. But the pattern, launch a number too clean to survive scrutiny, then quietly adjust it once someone checks, is now something the industry has done publicly at least twice inside eighteen months, at two different labs, using two different methods to get there.
Who Actually Needed That Number to Be 99.9
OpenAI confidentially filed paperwork for an IPO with the SEC on June 8, 2026. The company has said the timing is not settled, and subsequent reporting has OpenAI leaning toward a 2027 listing rather than a fall 2026 debut, which makes Astra’s launch-week timing look less like a countdown to a specific date and more like scene-setting for whichever date gets picked. OpenAI’s most recent funding round, in March 2026, valued the company at $852 billion on $122 billion raised, led by SoftBank and Microsoft, and some analysts think a public listing could aim as high as a trillion dollars. Astra’s API pricing landed at $10 per million input tokens and $50 per million output tokens, matching Anthropic’s Fable 5.1 almost exactly, a defensive move that protects margins ahead of a listing more than it wins customers on price. A number like 99.9% does not change any of that math. It changes the slide investor relations gets to reuse in every roadshow deck for the next two quarters, and it changes the story a prospective shareholder gets told about the company holding the paper.
Nvidia has a more direct stake in the answer than either lab does. The company invested $30 billion into OpenAI in February 2026, its largest single commitment to date, and now holds roughly $99 billion in shares of a company that is also its customer. Astra’s training run functions as a reference sale for Nvidia’s newest chip generation, and the 400,000 additional GPUs Jensen Huang has promised to ship represent revenue Nvidia has not yet booked. It was Huang, not OpenAI’s own president, who made the stronger public claim that AGI had arrived. Greg Brockman, who actually has a company and a pending IPO to protect, hedged. The chip supplier did not need to.
The harness itself has become a separate product line, sold by companies with no stake in whether Astra specifically looks good. Nvidia’s own AVO scaffold took Claude Opus 5, which scores 30.2% on ARC-AGI-3’s Standard harness unassisted, all the way to 100%. AWS’s Strands scaffold got Opus 5 to 99.95%. Anthropic, Google, and Microsoft all sell wrappers like these as commercial products now, each with its own pricing. A model’s headline benchmark score, in 2026, tells you less about the model than about which vendor was paid to dress it for the test.
There is a genuinely strange footnote sitting next to all of this. On the same day OpenAI published Astra’s numbers, its chief scientist, Jakub Pachocki, published a separate essay arguing that no lab, including his own, has solved alignment and monitoring well enough to keep scaling at full speed for much longer, and that voluntary slowdowns should become normal industry practice. One document on the company’s own site was arguing for caution. The other was arguing, in effect, that the company had just crossed into a new era of capability. Both were true statements from the same organization on the same afternoon, and almost nobody covering the launch mentioned that they were sitting next to each other.
What Changes When Nobody Can Check Your Homework
Independent trackers do not read this as a clean win either way. Artificial Analysis puts Astra’s Intelligence Index score at 61, level with the model it replaces and a point behind Meta’s Muse Spark 1.3, which is an awkward place to land for a launch built around the word intelligent. On the firm’s Coding Agent Index, Astra scores 67, tied with Claude Fable 5 and trailing Fable 5.1’s 70. Artificial Analysis rebuilt its entire index to version 4.2 the day after Astra’s launch, adding harder tasks and more private test sets specifically to make this kind of gaming more difficult going forward. That rebuild is itself an admission that the previous version had become gameable, by someone, sooner or later.
California passed a partial answer to the underlying problem, and Governor Newsom signed it into law on September 9, 2026. SB 813 and its companion bill, AB 1405, establish a California AI Standards and Safety Commission that will certify Independent Verification Organizations, outside groups empowered to test frontier models before or after release. The catch is timing: certification of the first such organizations is not required until January 1, 2028, more than a year after Astra shipped. Any law that fixes a problem sixteen months after the problem was already visible to everyone paying attention is not really a fix, it is a marker for how slow institutions are next to the thing they are institutionalizing. And the one real-world example of independent verification available right now is not encouraging. A METR investigation into an unrelated OpenAI security incident cost roughly $400,000 in API credits, credits OpenAI itself supplied for free, and used one of OpenAI’s own models to help analyze OpenAI’s own system. The researcher who wrote it up called it a “slop-vestigation.” On the only fully worked example anyone can point to, the tool doing the checking, the company being checked, and the company paying for the checking were the same company.
So the paradox this piece opened on resolves, but not the way OpenAI’s launch chart wants it to. The 99.9% is real. It happened, under real conditions, and it is not a fabrication the way Llama 4’s numbers were. But it was never the number that mattered, and right now nobody outside the labs themselves has the standing, the funding, or the deadline to tell you reliably which number does. Until an Independent Verification Organization actually exists and does not run on the tested company’s own credits, the correct response to any lab’s launch-day chart is the same one ARC Prize itself modeled: publish the conditions next to the score, and read the conditions first.
Sources For Further Reading
- ARC Prize: “OpenAI’s GPT-6 Astra on ARC-AGI-3”: https://arcprize.org/blog/astra
- François Chollet (ARC Prize co-founder) on X: https://x.com/fchollet/status/2095598451115614371
- Fortune, Emily Forlini: “OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch”: https://fortune.com/2026/09/04/openai-quietly-boosts-some-of-astras-evaluation-metrics-amid-rare-delay-in-publication-of-the-modeblog-post-announcement/
- The Next Web: “OpenAI’s AGI number came from a harness, not the model”: https://thenextweb.com/news/openai-astra-arc-agi-3-harness-62-7-vs-99-9-benchmark-revisions
- Fast Company: “Yann LeCun: Meta ‘fudged’ on Llama 4 testing”: https://www.fastcompany.com/91469583/yann-lecun-meta-llama-4-model-zuckerberg
- CNBC: “Airbnb hires former Meta AI Chief Ahmad Al-Dahle as CTO”: https://www.cnbc.com/2026/01/14/airbnb-tech-chief-meta-ai.html
- The Next Web: “Jensen Huang says AGI has arrived, and 400,000 more GPUs are coming”: https://thenextweb.com/news/jensen-huang-agi-has-arrived-gpt-6-astra
- CNBC: “OpenAI confidentially files for IPO, prepping Wall Street for mega AI debut”: https://www.cnbc.com/2026/06/08/openai-confidentially-files-for-ipo-prepping-wall-street-for-ai-debut.html
- Artificial Analysis: “Benchmarking GPT-6 Astra”: https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra
- NVIDIA Technical Blog: “NVIDIA AVO Reaches 100% on ARC-AGI-3”: https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/
- AWS Developer Community: “How a Strands agent took Claude Opus 5 from 30% to 99.95% on ARC-AGI-3”: https://dev.to/aws/how-a-strands-agent-took-claude-opus-5-from-30-to-9995-on-arc-agi-3-4kel
- Office of Governor Gavin Newsom: “Governor Newsom signs first-in-the-nation AI safeguards to protect Californians”: https://www.gov.ca.gov/2026/09/09/governor-newsom-signs-first-in-the-nation-ai-safeguards-to-protect-californians-calls-on-the-federal-government-to-do-its-part/
- METR: “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident”: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
- The Next Web: “California is deciding who may verify AI, and one investigation already cost $400,000 in tokens”: https://thenextweb.com/news/california-sb-813-independent-verification-organisations-metr-400000-tokens-openai-paid-eu-ai-act-scientific-panel-article-68
Find out What the AI Text Detection Business Gets Wrong About Its Own Product.





