An Anthropic AI system trained itself to be safer, faster and cheaper than human researchers

A new paper from Anthropic shows an automated system that fixes AI misbehaviour on its own, outperforming experienced humans in under six hours at a fraction of the cost.

AI2Day Newsdesk3 min read
Macro photograph of glowing amber light pulses travelling along the surface of a translucent circuit board, seen from directly above at a slight angle, deep bla
Share

Key points

  • Anthropic published a paper on Friday showing an automated system that improved AI safety performance across all 10 test benchmarks without hurting overall quality.
  • The best automated method beat what experienced human researchers proposed, on average within six hours.
  • Running the automated system costs roughly $4 per hour, compared to the $150 per hour Anthropic pays human researchers.
  • The paper was led by Anthropic Fellow Chen Yueh-Han and is an early step toward AI systems that improve their own training.

AI companies are increasingly trying to get AI to train other AI. Anthropic just showed the world what that looks like when it actually works.

On Friday, Anthropic published a paper titled "Automated Researchers Can Reliably Mitigate Alignment Failures", describing a system it calls the Automated Alignment Researcher, or AAR. Alignment, in plain terms, means making sure an AI model behaves the way its makers intended rather than producing harmful or misleading outputs. Keeping a model aligned is normally a slow, expensive job done by human specialists.

The AAR does that job on its own.

What exactly did the system do?

Given 10 separate tests for specific types of bad AI behaviour, the AAR improved the model's score on every single one without making it worse at other tasks. That clean sweep matters: previous attempts to fix one problem often broke something else.

The process mirrors how a human researcher would work. The system reads the available scientific literature, proposes a fix, trains the model using that fix for 30 minutes, then checks whether things improved. Methods that work get kept; methods that do not get dropped. The system repeats this over several rounds, each time raising the bar.

"Overall, these results provide early evidence that automated alignment post-training could become practical in the near term," the paper states.

Led by Anthropic Fellow Chen Yueh-Han, the project is a concrete step toward what researchers call recursive self-improvement: the idea that an AI system could get better at its own training, not just at the specific tasks it was built for.

How does it compare to human researchers?

Faster and much cheaper. The paper puts the cost at $4 per hour for the automated system against $150 per hour for a human researcher, and the AAR's best method outperformed what experienced humans proposed within six hours on average. "Human guided research directions do not lead to stronger performance," the paper notes.

That is a striking claim, and the authors do not bury it.

Metric AAR (automated) Human researcher
Cost per hour $4 $150
Time to beat human proposals Under 6 hours Baseline
Benchmarks improved 10 of 10 Lower average

Should anyone be worried?

The paper is honest about what the system cannot yet do. The AAR only works as well as the benchmarks it is measured against. If those tests fail to capture the full range of ways an AI can misbehave, a high score means less than it looks. Someone still has to design, maintain and expand those tests, and that is skilled human work.

The automated researchers also draw on a library of existing research literature. Keeping that library current and accurate is another human responsibility that does not go away.

For now, the AAR is a tool that speeds up one part of safety research rather than a system that replaces the whole field. Whether the gap narrows further is the question Anthropic, and frankly the whole industry, is now watching.

© 2026 AI2Day