Late on September 3, a widespread outage rattled global AI services. Before the dust settled, OpenAI unveiled GPT-6 Astra at 3:30 AM on September 4, initially opening access to select enterprises and expanding to all paid users by September 5. OpenAI President Greg Brockman kicked off the release with a bold line: "Welcome to the AGI era." Within days, the new model had stolen the spotlight from competitors that launched just ahead of it. The scorecard looks impressive. Under OpenAI's official metrics, Astra posted massive gains across key benchmarks: ARC-AGI-3 vaulted from 7.8% in the previous generation to 99.9%, ExploitBench hit a perfect score, and FrontierMath Tier4 reached 98 points. It's worth noting that some high scores were achieved on OpenAI's proprietary evaluation framework. When third parties stripped away that framework, ARC-AGI-3 landed at about 62.7%—still double the second-place result on the leaderboard.
文章图片 2
Compared with the previous flagship, overall intelligence indices remained roughly flat. The real progress landed on the task side: Terminal-Bench 4.0 scored 57.9%, OSWorld 2.0 reached 72.6%, and average single-task completion time dropped to about 40 minutes from the prior generation's 75 minutes. AutomationBench jumped from 18.1% to 41.4%. The gap between "benchmark cramming" and "real-world execution" is where this generation made its biggest strides—and it leaned hard into the latter.
文章图片 4
Efficiency may be the most underappreciated change in this release. Multiple third-party labs measured a significant drop in Astra's token consumption on agentic tasks: at the highest configuration, it used roughly 65% fewer output tokens than competitors, and Terminal-Bench API cost estimates came in about 60% lower. OpenAI's approach is to pack more effective reasoning into the same compute rather than simply letting the model "think longer." Token savings don't come free. Foreign media reports indicate Astra's training run used a GPU cluster on the scale of 100,000+ accelerators—the largest training effort in OpenAI's history. Architecturally, reports also point to the use of recurrent depth techniques, where the same set of network layers iterates repeatedly to deepen computation. The model saves on runtime expense while betting bigger on training investment and deeper architectural design.
文章图片 6
This will reshape how the industry measures competitive standing. As unit prices keep falling and rival models compete on inference efficiency, the number of tokens required to complete a real task—and the dollar cost attached—is becoming a more telling metric than raw benchmark scores. For enterprises, lower per-task cost means the same budget can support significantly more real-world requests, lowering the barrier for AI to move from demo to production. And as per-task token usage is genuinely compressed, enterprises no longer worry only about whether they can afford a single call—they need visibility into every call and the ability to route saved budget into more production workloads. Turning "saved tokens" into measurable, orchestratable compute output is precisely why StarWar Cloud insists on transparent metering and compute scheduling across its large-model API plaza and GPU computing platform.