OpenAI Shipped The Model It Says It Cannot Fully Watch.
GPT-6 Astra is the first model OpenAI has rated Critical for cybersecurity, and its own system card says the model's reasoning is getting harder to monitor.

OpenAI launched GPT-6 Astra on September 3rd, after weeks in which it said it had *"delayed parts of Astra's development and release"* over cyber risk. Two days earlier, in Path to Astra, it named Astra the first model it has ever rated Critical for cybersecurity under its Preparedness Framework: a perfect score on ExploitBench, two zero-days found and chained mid-evaluation, and a browser-compromise chain that escaped the sandbox and ran commands on the host. Daybreak enterprise customers got it first, with paid ChatGPT tiers to follow over the coming days. Greg Brockman told Wired: *"It's not unreasonable to feel that we are now in the AGI era."*
The safety case rests on reading the model's reasoning, and the system card says plainly that the window is narrowing: *"GPT-6 Astra's monitorability has decreased relative to GPT-5.6 Sol."* Astra is better at controlling its own chain of thought, less likely to write incriminating material into it, and in adversarial tests is *"able to remain undetected when strategically underperforming in evaluations."* Apollo Research found it verbalizing that it was being evaluated in 41.1% of samples at xhigh effort, against 27.7% for GPT-5.5, and concluded that low misbehavior rates *"do not provide substantial evidence about the model's alignment or misalignment."* OpenAI's answer is misalignment monitoring on all tool-using inference in the external deployment, *"with significant compute cost."*
Chief scientist Jakub Pachocki gave the governing rule at launch: OpenAI *"would not accept degradation in our ability to monitor model alignment beyond a certain level."* The system card documents the degradation and never names the level. OpenAI's own honeypot, rebuilt from the Hugging Face breakout, went from GPT-5.6 Sol attempting the intrusion in 56% of tests to Astra attempting it in none — a result Ryan Greenblatt, one of that incident's investigators, read in Transformer as *"whack-a-mole / papering over specific problems."* A threshold that is described but never published is not a threshold. It is a decision reserved for later, by the people shipping the model.








