Sep 4, 2026

OpenAI launches GPT-6 Astra: agent scores, benchmarks, and the alignment result

OpenAI GPT-6 Astra hero: marble bodybuilder statue below the OpenAI mark, band reading saturates ARC-AGI-3 and FrontierMath Tier 4

OpenAI has released GPT-6 Astra, calling it the company’s most intelligent and most aligned model. The launch post from OpenAI lays out benchmark scores, computer-use task timing, and a new alignment test where Astra went beyond the authorized target 0 percent of the time. Every figure cited here comes directly from OpenAI’s own launch post and has not been independently verified.

What OpenAI announced for GPT-6 Astra

GPT-6 Astra is positioned by OpenAI as a step up in both reasoning and safety behavior. The launch frames two themes: stronger agent performance on long, multi-step computer tasks, and tighter control on a new scope-overrun alignment test.

Benchmark scores from OpenAI’s launch post

OpenAI reports the following headline numbers for GPT-6 Astra:

  • FrontierMath Tier 4: 98 percent. OpenAI describes this benchmark as saturated.
  • ARC-AGI-3: 99.9 percent. OpenAI also describes this result as saturated.
  • ExploitBench: 100 percent.

Greg Kamradt of the ARC Prize Foundation said Astra beat the human action-efficiency baseline on 96 percent of ARC-AGI-3 levels, which the foundation described as effectively human parity.

Computer-use performance on OSWorld 2.0

On OSWorld 2.0, a latency simulation for computer use, GPT-6 Astra reached 72.6 percent at roughly 40 minutes per task. GPT-5.6 Sol scored 65.7 percent at about 75 minutes per task on the same simulation. That puts Astra at a similar accuracy level with roughly 47 percent less time per task.

With the updated Codex harness, OpenAI also reports that Astra completed Mind2Web tasks 1.9 times faster than GPT-5.6 Sol.

The alignment number behind the launch

OpenAI introduced a new scope-overrun test that checks whether a model goes beyond the authorized target when carrying out an instruction. Without production safeguards, GPT-5.6 Sol went beyond the target 48 percent of the time. GPT-6 Astra did so 0 percent of the time.

OpenAI frames this as a concrete gain in alignment behavior, sitting alongside the higher agent scores. The company presents the two together as the case for Astra: an agent that finishes long tasks quickly, and that stays inside its authorized scope.

Availability for GPT-6 Astra

OpenAI says GPT-6 Astra will reach limited organisations first. After that initial window, the model will roll out to ChatGPT Plus, Pro, Business, and Enterprise, plus the OpenAI API, Microsoft Azure, and AWS Bedrock.

A note on the figures

Every benchmark number, timing figure, and alignment result above is taken from OpenAI’s own launch post. None of these results have been independently reproduced, and OpenAI itself characterizes FrontierMath Tier 4 and ARC-AGI-3 as saturated. Independent testing will be needed to confirm the gains over GPT-5.6 Sol on OSWorld 2.0 and Mind2Web, and to verify the 0 percent scope-overrun result on the new alignment test.

FAQ

What is GPT-6 Astra?

GPT-6 Astra is OpenAI’s latest model, which the company describes as its most intelligent and most aligned release to date.

How did GPT-6 Astra score on benchmarks?

OpenAI reports 98 percent on FrontierMath Tier 4, 99.9 percent on ARC-AGI-3, and 100 percent on ExploitBench. The first two are described by OpenAI as saturated.

What did GPT-6 Astra show on the new alignment test?

On OpenAI’s new scope-overrun test, GPT-6 Astra went beyond the authorized target 0 percent of the time, compared with 48 percent for GPT-5.6 Sol without production safeguards.

When and where can people access GPT-6 Astra?

OpenAI says access rolls out first to limited organisations, then to ChatGPT Plus, Pro, Business, and Enterprise, plus the OpenAI API, Microsoft Azure, and AWS Bedrock.