OpenAI's safety overview, published 3 September 2026, confirms GPT-6 Astra is the first model to reach the Critical level of cybersecurity capability under the company's own Preparedness Framework S¹. When given appropriate tools and access, the system can uncover undisclosed security weaknesses and create new methods to exploit them against various heavily defended networks P². The most capable model OpenAI has ever shipped is now also the one that can break into the systems it runs on, and the overview does not detail what restrictions accompany that crossing.
My read: This is the first time OpenAI has publicly said one of its own models hits the Critical bar for offensive cyber capability. I take the self-assessment seriously because OpenAI published it on its own deployment safety hub P³, but I do not treat "Critical" as an industry standard. It is a label from a framework OpenAI wrote, updated, and grades itself against P⁴. The detail that matters is not the label but the capability description: finding unknown flaws and writing exploits. If that holds up under independent testing, it changes what AI models are to security teams, from assistant to threat surface.
What "Critical" actually means
The Preparedness Framework, now in version 2 (last updated April 2025), is OpenAI's internal system for rating the danger level of its own models before deployment P⁴. It tracks several categories of risk. Cybersecurity is one of them. GPT-6 Astra is the first model to score Critical in that category S¹.
What does that translate to in practice? According to OpenAI's safety document, when equipped with suitable tools and permissions, GPT-6 Astra is capable of identifying undisclosed vulnerabilities and creating fresh exploitation techniques against numerous highly secured systems P². In security terms, that is zero-day discovery: finding flaws that the software's own developers do not know about yet.
Until now, that kind of work required skilled human researchers, sometimes teams of them, working for weeks or months. OpenAI is saying its model can do it, given the right setup.
A framework OpenAI wrote and grades itself against
The Preparedness Framework is not an external standard. OpenAI created it, maintains it, and applies it to its own models P⁴. The Critical rating is OpenAI's own classification, not an independent audit. The safety overview published 3 September does not cite any third-party evaluations S¹.
This matters because the whole point of a safety threshold is to tell you when to slow down. If the entity setting the threshold is also the entity shipping the product, the threshold is only as trustworthy as that entity's willingness to hold its own release.
OpenAI also maintains a GitHub repository, c6ai/preparedness, that appears to track releases related to its Preparedness work P⁵. The repo is MIT-licensed and was created in September 2025. It currently has zero stars and four open issues, so it is hard to assess whether it represents active independent verification or an internal tracking tool P⁵.
What to do about it
If you run a security team, the practical shift is this: your vulnerability surface may now include AI models that can probe it. A mid-sized fintech with a bug-bounty program, for example, should consider whether its external-facing systems could be mapped and tested by an AI agent with GPT-6 Astra-class capability before its own defenders have found the same flaws. That does not mean panic. It means moving the timeline for routine penetration testing forward and treating AI-assisted discovery as a plausible threat rather than a theoretical one.
For developers building on OpenAI's API, the question is whether your application gives the model the tools and access OpenAI's own description mentions P². A model that can find zero-day flaws needs network access, code execution, and target systems to be dangerous. If your integration does not provide those, the risk profile is lower. If it does, you are now in the threat model.
One thing to check this week: review what tools and access your AI integrations actually provide. If your agent can run code and reach external endpoints, that is the combination OpenAI's own safety overview flags.
What we don't know yet
The safety overview does not define the specific technical benchmarks that separate one capability level from another in the cybersecurity category S¹. We know the Critical threshold was crossed, but not by how much, or what the test conditions were.
No independent third-party safety evaluations are cited in the document S¹. The Critical rating is OpenAI's own assessment under its own framework. Whether external red teams or academic researchers can reproduce the capability claims remains an open question.
The term "broadly deployed" is also undefined in the materials provided. It should not be assumed to mean general public availability without further clarification.
The next signal will be the first published independent red-team test of GPT-6 Astra's cyber capabilities. Security researchers typically release these within weeks of a major model launch. We will check their findings against OpenAI's claims when they land.
If you want that follow-up in your inbox, subscribe now.
Sources: S1 — Safety overview: GPT-6 Astra · P2 — Safety overview: GPT-6 Astra · P3 — GPT-6 Astra System Card - OpenAI Deployment Safety Hub · P4 — Preparedness Framework · P5 — c6ai/preparedness
More from Not A Tech Guy
- Hermes Agent: 240k GitHub stars for self-improving AI
- Gilbert + Tobin deploys ChatGPT Enterprise firm-wide
- Prompt injection attack on AI agents jumps to 28% success
Generated from an audited evidence pack with primary-source research. Social-media items are discussion signals, not verified facts. Nothing here is financial, legal or medical advice.