GPT-6 Astra Arrives With a Benchmark Asterisk
3 min readOpenAI shipped GPT-6 Astra on September 3, calling it the most intelligent and most aligned model the company has built. The headline number was a 99.9 percent score on ARC-AGI-3. The number that deserves more attention is 62.7 percent, and the distance between the two says as much about how AI gets measured as it does about the model.
What OpenAI Actually Shipped
GPT-6 Astra is the new flagship, and the specifications are substantial. It carries a context window of roughly 1,050,000 tokens with a maximum output of 128,000, accepts text and image input, and has a knowledge cutoff of April 30, 2026. OpenAI positions it for long-horizon agentic work: software engineering, deep research, scientific analysis, and tasks that involve driving a computer or a browser across many steps.
The rollout is staged rather than instant. A limited set of organizations got access on day one, with ChatGPT Plus, Pro, Business, and Enterprise following over the days after, alongside availability through the OpenAI API and AWS. API pricing lands at $10 per million input tokens and $50 per million output tokens, with cached input at $1 per million and batch jobs at half price. A Fast mode is offered at twice the standard rate.
The GPT-6 Astra Benchmark Asterisk
ARC-AGI-3 is built to test whether a model can reason through unfamiliar puzzles rather than recall something close to its training data. The 99.9 percent result came from a provider-specific test harness that preserved the model’s reasoning state between actions. Run on the standard neutral harness, the same model scored 62.7 percent, and it burned considerably more compute getting there.
Neither number is fake. They measure different things. The first shows what the model can do when the scaffolding around it remembers everything already worked out. The second shows what happens when that scaffolding is stripped away.
Why It Matters
That gap previews an argument the industry is going to keep having. As models are sold on agentic performance, the harness around the model becomes part of the product, and the vendor controls the harness. Buyers weighing GPT-6 Astra against Anthropic’s Fable 5.1 or Google’s Gemini 3.8 Flash will increasingly need to ask which harness produced a given figure before the comparison means anything at all.
GPT-6 Astra looks like a real step forward on long-running agent work, and the price holds steady against the outgoing flagship. It is also a reminder that a benchmark score is now a claim about infrastructure as much as about intelligence. Watch for independent replications on neutral harnesses over the coming weeks.
