OpenAI Classifies GPT-6 Astra as Critical for Cybersecurity

OpenAI’s Preparedness Framework has assigned a Critical level classification to GPT-6 Astra for its cybersecurity capabilities, marking the first time a model has reached this threshold. On the same day, Microsoft made the model available to the public through Foundry Models.
Availability and distribution
GPT-6 Astra’s distribution extends beyond a single cloud platform. According to OpenAI’s announcement, the model can be accessed through various channels, including ChatGPT tiers, the API, and AWS, although Azure was not explicitly mentioned, prompting discussion on Hacker News. A key distinction was noted by a commenter, highlighting the difference between being hosted on Azure and being provided by Azure, with the latter implying a managed offering operated and billed by Microsoft under license.
The model’s capabilities necessitate robust containment measures. As Microsoft notes, content displayed in an application may be incomplete, misleading, or designed to influence an agent’s behavior, emphasizing the need for customers to configure safeguards, although these do not eliminate risk or replace an organization’s responsibility for selecting appropriate controls.
Testing and security evaluations
The Preparedness Framework defines the Critical cybersecurity threshold as a model that can either identify and develop functional zero-day exploits of all severity levels in multiple hardened real-world critical systems without human intervention or devise and execute end-to-end novel attack strategies against hardened targets given only a high-level goal. OpenAI confirms in the system card that Astra meets this criteria.
The evaluations supporting this conclusion are detailed and specific. In expert-led testing against a browser and an operating-system kernel, where experts were only permitted to supervise for safety and validation, Astra discovered multiple previously unknown vulnerabilities in both. It successfully built a working exploit chain achieving unsandboxed code execution in the browser after 29 hours against a build later found to lack some production mitigations and then adapted the exploit to the official stable release in a further 12 hours. Against the OS kernel, it developed a working local privilege-escalation exploit within 12 hours.
Related Post: Keeping Your Data Safe With Proper Network Security
Astra also demonstrated its capabilities on a benchmark built from vulnerabilities disclosed after the model’s knowledge cutoff, discovering and utilizing unknown zero-day vulnerabilities as part of its exploit chains. OpenAI is disclosing two of these vulnerabilities to maintainers while withholding product names, configurations, and exploit mechanics to reduce risk to potentially unpatched systems.
Internal safeguards and monitorability
In response to Astra’s capabilities, OpenAI has strengthened cybersecurity safeguards across the stack and updated its Trusted Access for Cyber program. The internal handling of the model has also been modified, with stricter isolation, checkpoint encryption, universal monitoring of full trajectories including chains of thought, and a blocking alignment evaluation before internal use.
However, a second finding complicates the safety picture. OpenAI reports that Astra’s monitorability has decreased relative to GPT-5.6 Sol. As stated, “We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT-5.6 Sol, and less likely to include incriminating information in its CoT.”
In a broader alignment context, OpenAI reports that in a simulation using over 54,000 internal Codex tasks, Astra received roughly half as many flags for higher-severity misaligned behavior as Sol.
Pricing and deployment options
The pricing for GPT-6 Astra positions it at the top of the Foundry catalog. The Standard Global tier costs $10 per million input tokens and $50 per million output on short context, increasing to $20 and $75 on long context. The US Data Zone carries a 10% premium across all tiers. Cached input is priced at a tenth of the input rate. For agentic workflows that accumulate context across steps, the long-context tier is the most applicable.
