Astra scored 100% on ExploitBench (vs 78.5% for GPT-5.6 Sol), 98% on FrontierMath Tier 4, and 99.9% on ARC-AGI-3, and found two zero-days in Google's V8 engine during testing. It is OpenAI's first model to reach the Critical cybersecurity threshold under its Preparedness Framework, so cyber features roll out first to select partners.
Astra, OpenAI's GPT-6 model, is rolling out beyond Pro and Enterprise to all Plus and Business users, with API pricing at $10/$50. It posts 98% on FrontierMath Tier 4, 63%/99.9% on ARC-AGI-3 (standard vs adapter harness), and 100% on ExploitBench.
Fable 5.1 scores 55.8% on Terminal-Bench 4.0 in Claude Code, ahead of Fable 5 (42%) and Opus 5 (52.3%). Pricing matches Fable 5, but API cache reads drop 75% to $0.25/MTok; Perplexity ranked it first on its August WANDR eval at 0.601.
OpenAI says an internal model significantly more capable than GPT-6 Astra produced the proof in 88 hours using ~10,000 coordinating AI agents, with a writeup and formal Lean proof. NYU's Tristan Buckmaster disputes the process, saying he and Levent Alpöge worked the problem for nearly a year using Claude and Codex.
Bloomberg reports the coordination push with the US government could make it harder for smaller AI companies to compete, per startup executives and industry watchers. Anthropic CEO Dario Amodei said a swarm of AI agents could take over the internet in six months to a year without more safeguards.