Anthropic Built A Better Model And Shelved It.
The latest risk report discloses an internal model that beats the shipped flagship, raises the company's own misalignment estimate, and says there are no plans to release it.

Anthropic's latest risk report, covered by Axios on August 14th, describes an internal model the company calls Model 2 that shows *"noticeable improvement"* on many internal tasks over the shipped Mythos tier. Both Mythos 5 and Model 2 are used *"heavily"* inside the company for coding, agentic work, and data generation. The report's disposition is one sentence long: *"We do not currently have plans to release this model externally."* In the same document Anthropic raised its broad estimate of misalignment risk in high-stakes situations from *"very low"* to *"low,"* citing recent cybersecurity incidents.
The mechanism worth reading twice is not the shelving, it is the reason given for the uncertainty. The report says the company is *"less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations... no longer capture increases in models' capabilities."* That is a lab stating on the record that its measuring instrument has fallen behind the thing it measures. Anthropic frames Model 2 as routine — *"we internally train and evaluate many different exploratory versions of models that we don't intend to release"* — and notes the jump is smaller than Opus 4.6 to Mythos earlier this year. Both things can be true, and the eval gap is the load-bearing one.
This is the second time in eight days a frontier lab has taken a capability off the table and said cyber. OpenAI slowed Astra on August 7th because it *"cannot rule out critical cyber capabilities."* Anthropic rolled back its own training-pause commitment in a February update to its Responsible Scaling Policy, on the argument that unilateral restraint makes the world less safe — and has now unilaterally restrained anyway, on a model nobody outside the company had asked about. Restraint that arrives as a disclosure rather than a policy is not a commitment. It is a decision that can be revisited next quarter by the same people who made it.







