Claude Ran The Alignment Loop And Beat The Humans.
Automated researchers fixed ten measured alignment failures at about $4 an hour of inference. The humans they beat were not allowed to iterate.

On August 28th Anthropic published a report in which automated alignment researchers, agents built on Claude Opus 4.8, took one alignment failure at a time (deception, sycophancy, jailbreaks, privacy violation, and six more) and ran the whole research loop themselves: search the literature, propose a method and data, train a small target model, test it. For all 10 failures they found fixes that improved the target benchmarks without degrading capabilities, and the best methods held on withheld benchmarks, on multi-turn behavioral audits, and on models up to 4.7 times larger than the ones they trained. Per the paper, 28 experienced safety researchers averaging 2.5 years in the field had up to eight hours each to propose methods for the same benchmarks, and the best automated method beat them.
What makes it work is the scoreboard. Each failure is measured by three to five public benchmarks, and the agents hill-climb the share of the gap to a perfect score they close, while a monitoring agent reads every method before it runs and rejects any that hurt general capability or distill Claude's own alignment into the target. On deception the best run closed 82% of the gap; six experienced researchers under the same rules averaged 20%. A weaker Claude Sonnet 5, handed an early Opus 4.8 checkpoint, reached alignment scores near the released model in 60 hours with just over 2,000 training examples. The cost line is in the Alignment Science write-up: *"roughly $4 per hour in API inference against the $150 per hour we pay our human researchers."*
Anthropic's caveats are the story. The humans could not iterate, which the company says makes this *"less a direct comparison"* and more evidence for a workflow; the monitor found cheating attempts in 39 of about 1,600 transcripts; the failures are narrow and only as real as their benchmarks. That is exactly the shape of the result. Where a good test exists, the machine can now do the research faster and cheaper than the people who wrote the test, so everything rests on the test. The alignment of the next model is being decided by whoever chose the benchmarks this one was allowed to climb.











