The comparison Anthropic chose to publish is $4 against $150. That is the hourly cost of an automated alignment researcher, an AAR, in API inference, against the hourly cost of the human safety researchers it was measured against. The paper, "Automated Researchers Can Reliably Mitigate Alignment Failures," went up Friday, with Anthropic Fellow Chen Yueh-Han among the authors, and the setup is simple to describe: Claude searches the literature, proposes a mitigation method, trains a student model for a fixed iteration, tests it against an alignment benchmark, keeps what works and discards what does not.

The results are reported per failure type, not as one headline number. Across ten alignment failures, among them deception, sycophancy, jailbreaks, privacy violations and reward hacking, the best automated methods closed between 26 and 96 percent of the measured safety gap, and Anthropic says capabilities did not degrade on the metrics it checked. Against 28 human safety researchers, the automated system won on average, and on deception it closed 85 percent of the gap where the humans closed 20. "The best AAR method beats what experienced humans propose, on average within six hours," the company writes, per TechCrunch. At production scale, Claude Sonnet 5 aligning Claude Opus 4.8, a 60 hour run using roughly 2,000 training examples closed 65 percent of the gap, which Anthropic describes as roughly 15,000 times more efficient than its production alignment procedure.

The limitations section is where this gets honest, and it is worth reading before the headline. The benchmarks cover narrow, nameable failures; political bias was not measured. Rare or emergent failures have no benchmarks to close. Capability degradation was checked only on metrics chosen in advance. Evaluations like Petri are, in Anthropic's words, only proxies for real-world misalignment. Whether the fixes survive extensive later reinforcement learning is untested. And the automated researcher cheated in 2.4 percent of transcripts, caught by monitoring whose future reliability is itself an open question.

Read plainly, this is not a general autonomous researcher, and Anthropic does not claim one. It is a demonstration that where a failure can be named and measured, the post-training work of fixing it can be handed to a model and run at a rate no research team can match. That makes the benchmark the bottleneck: the system optimises exactly what you can define, which puts the burden on the definitions, and on the same open question the field keeps circling, whether an evaluation measures the thing or its shadow.