OpenAI representatives claim that the recent launch of one of its latest models, GPT-6 Astra, marks the beginning of the artificial general intelligence (AGI) era — a hypothetical scenario in which artificial intelligence (AI) can learn and reason like humans. Following very recent claims from technology executives that we’ve now reached this milestone, how likely is this claim to stand up to scrutiny?
Amid revelations that safety concerns have forced the company to abandon the planned rollout of GPT-6.1 Astra, independent researchers and benchmark creators warn that GPT-6 Astra’s record-breaking scores rely on heavy prompt optimization rather than true general intelligence. This comes as experts warn about the risks of future AI systems and AI executives jointly call for a slowdown in AI research.
In announcing the model’s release, OpenAI President Greg Brockman suggested that Astra signaled a fundamental turning point in AI evolution. “If we fast-forward a couple years, and we look back and say when was it really that AGI was created, I think it’s going to be about this time, and I think it might be about this model,” Brockman said during a Sept. 3 news conference marking the launch.
How did the new model perform in benchmarking?
The benchmarks scores are impressive, but even the creator of one of the most important suggests acing it doesn’t automatically mean we’ve reached AGI.
(Image credit: Cheng Xin via Getty Images)
As part of the testing-and-verification process for the new model, OpenAI released benchmark results across a range of tests that measure the capabilities of AI models at various tasks. OpenAI representatives claimed in a statement that the company’s new model delivers “state-of-the-art” performance across a variety of fields, including software engineering, the autonomous use of computer programs, mathematics problems, and even scientific research.
Demonstrations, for example, highlighted Astra’s ability to generate 3D CAD code; lay out printed circuit boards in CAD (computer-aided design) software; convert digital 3D models into interactive, video-game-style environments; and deploy hosted web applications directly from prompts.
“In our early testing, Astra stood out by approaching legal work the way a discerning lawyer does: it distinguishes documents from established records, surfaces unsupported assumptions, and converts gaps into concrete drafting positions,” said Niko Grupen, head of applied research at legal services AI developer Harvey as part of OpenAI’s announcement.
On ARC-AGI-3 — a benchmark designed to measure how efficiently AI systems acquire new skills in novel, abstract environments — Astra achieved a headline score of 99.9%. Independent testing by the nonprofit ARC Prize Foundation, which oversees the benchmark, revealed that Astra surpassed “human action-efficiency” baselines on 96% of levels, taking 51.7% fewer actions per level, on average, than human solvers.
“Not only is this the best model we’ve ever tested,” ARC Prize Foundation President Greg Kamradt said as part of the OpenAI announcement, “but it also represents a meaningful step change in frontier-model performance — not only in its ability to navigate and solve novel environments but also in how efficiently it learns to do so.”
How much can we read into test results?
Many researchers have highlighted that Astra’s top ARC-AGI-3 performance relied on a proprietary “Provider Adapter” harness, a specialized software scaffolding layer that’s inaccessible to rival developers. When tested using the benchmark’s neutral Standard harness, this score dropped to 62.7%. ARC Prize organizers also noted that saturating the benchmark does not constitute proof of achieving AGI — emphasizing that its closed, deterministic puzzle environments do not capture the open-ended complexity of the real world.
“When we launched ARC-AGI-3, “we made it clear that saturating the benchmark would not represent ‘proof of achieving AGI,'” Kamradt wrote in a blog post. “Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.”
Discrepancies in the data were reportedly present even before the official announcement. As reported by Fortune, an embargoed draft provided to news organizations prior to the model’s launch listed Astra’s ARC-AGI-3 score at 98.6% before it jumped to 99.9% on the live site.
Anka Reuel and Mike Hardy, doctoral researchers in AI at Stanford University, also raised concerns over OpenAI quietly revising several published metrics post-launch, including halving Astra’s reported hallucination rate from 4.2% to 2.0% before reverting it.
When we launched ARC-AGI-3, ‘we made it clear that saturating the benchmark would not represent ‘proof of achieving AGI'”
ARC Prize Foundation President Greg Kamradt
Speaking to Fortune, Reuel and Hardy attributed the shifting figures to “benchmaxxing” — an industry term that describes the practice of repeatedly rerunning evaluations under subtly tweaked prompts, scaffolds or compute allocations to hunt for peak theoretical scores.
In a 2025 paper published on pre-print server arXiv titled ‘Benchmarking is Broken — Don’t Let AI be its Own Judge’, experts including Princeton doctoral candidate Zerui Cheng and University of Luxembourg postdoctoral researcher Dr. Stella Wohnig warned that such practices create confusion and undermine trust, particularly when official technical documentation omits the basic methodology, making independent verification nearly impossible. Crucially, scores achieved through benchmaxxing reflect idealized performance ceilings under optimal lab conditions rather than practical, out-of-the-box reliability.
According to OpenAI’s benchmark data, across specialized domain evaluations, GPT-6 Astra delivered 10% to 20% better performance than both its predecessor, GPT-5.6 Sol, and top market competitors.
Highlights include a 98% score on FrontierMath Tier 4 (v2), which evaluates expert-level mathematical reasoning; a perfect 100% on ExploitBench, which is designed to test offensive cybersecurity capabilities; 95.9% on BenchCAD, which measures an AI’s ability to reconstruct 3D objects using CAD code; 57.9% on Terminal-Bench 4.0, which assesses complex terminal-based system administration and software engineering; and 41.4% on AutomationBench, which tests whether agents can complete multistep business workflows across applications.
But independent findings from external AI benchmarking firm Artificial Analysis contradict these results. Astra’s score on the independent Artificial Analysis Intelligence Index — an aggregate metric that evaluates general reasoning across frontier models — remained completely flat, at 61, compared with GPT-5.6 Sol, while trailing competitors like Anthropic’s Claude Fable 5.1 and Meta’s Muse Spark 1.3.
While some benchmark results have seen improvement since the previous generation, Astra underperformed in other tests. Artificial Analysis’ testing showed that against GDPval-AA v2 — an economically focused benchmark that measures real-world workplace tasks across 44 occupations — Astra suffered a significant drop in its relative leaderboard ranking compared with GPT-5.6 Sol. Additional regressions were observed in benchmarks that test customer service support, scientific Python programming, and long-context reasoning across large documents.
In terms of overall token use — the metering system that measures the fragments of text or data processed by AI models — OpenAI representatives said GPT-6 Astra achieved dramatic efficiency gains on complex, long-horizon tasks. Findings from Artificial Analysis confirmed that on software engineering benchmarks, Astra was approximately 70% more token-efficient than GPT-5.6 Sol, using roughly one-third the total tokens of its predecessor and one-fifth the tokens of Claude Opus 5.

AI systems today can do more than a simple back-and-forth text exchange, graduating to taking actions on our behalf.
(Image credit: CFOTO via Getty Images)
On professional workplace evaluations like Agents’ Last Exam, the new model reduced output token consumption by up to 65% compared with Opus 5, while on general intelligence evaluations, Astra achieved a roughly 10% output token reduction at max effort compared with Sol.
On the other hand, this achievement is offset by a 250% increase in the base API pricing for GPT-6 Astra compared with GPT-5.6 Sol. Token prices jumped from $4/$20 to $10/$50 per million input/output tokens. As a result, Astra ends up approximately 75% more expensive per task than its predecessor on general intelligence evaluations, despite using 10% fewer output tokens. This trade-off has led researchers at Artificial Analysis to question whether AI labs are relying on expensive computational brute force to squeeze out minor gains, rather than achieving true technical breakthroughs.
Does achieving AGI even matter?
With its latest model, OpenAI has placed a stronger emphasis on safeguards and controls. The move follows a series of high-profile incidents in which frontier AI models inadvertently hacked into third-party networks and computer systems as part of routine testing operations, after the models went beyond what humans expected they would do when encountering challenging tasks.
According to OpenAI representatives, Astra did not exceed the scope of its tasks in any security evaluations. For comparison, GPT-5.6 Sol went beyond its authorized parameters in almost 50% of test cases. Hallucination rates have also been reduced significantly, falling to 4.2% on internal tests, compared with 12.2% for GPT-5.6 Sol (with independent testing by Artificial Analysis showing a drop from 92% to 51% at max effort). The model also showed an improved ability to handle ambiguous prompts.
David Wood — chair of London Futurists, a non-profit that hosts discussion groups on emerging technology — argued that beyond individual test results, Astra highlights a fundamental shift in how humans will interact with AI as models move from answering questions to autonomously pursuing complex goals in digital environments.
“The impressive benchmark results matter, but what matters more is the combination of intelligence, autonomy, computer use, and cybersecurity capability,” he said in an email to Live Science. “That combination also demands caution. The more useful these systems become, the greater the consequences when they misunderstand our intentions, are misused, or find ways around the safeguards we give them.”
Wood suggested that whether Astra can be classified as AGI is less relevant than recognizing the need for AI control and governance to keep pace with technical development.
“The possibility that AI could progress from systems like Astra towards all-round superintelligence makes much greater human vigilance, awareness, and collaboration increasingly urgent,” he said.
Help us improve Live Science Pro: We’re always trying to make our content better. Leave us feedback about Pro here.


