On Friday, Anthropic published a new paper titled "Automated Researchers Can Reliably Mitigate Alignment Failures," offering an early look at how AI systems could improve a model's performance on alignment benchmarks. The system, called the Automated Alignment Researcher (AAR), was able to improve performance on all 10 benchmarks for specific misaligned behaviors without degrading overall performance.
Led by Chen Yueh-Han, a fellow in Anthropic's program, the AAR replicates much of the traditional research approach. Each automated system searches available literature, proposes a method, and trains the model using that method for 30 minutes, gradually increasing the benchmark over several iterations. Effective methods are preserved while ineffective ones are discarded, allowing rapid operation at scale.
"Overall, these results provide early evidence that automated alignment post-training could become practical in the near term," the paper states.
The paper compares the AAR directly to its human equivalent, noting: "The best AAR method beats what experienced humans propose, on average within six hours." It adds that "human guided research directions do not lead to stronger performance" and provides a cost comparison: roughly $4 per hour in API inference against $150 per hour for human researchers.
The paper acknowledges limitations. The automated system only works as well as the benchmarks reflect actual alignment goals, and significant work remains in establishing, maintaining, and expanding both the benchmarks and the literature the AAR draws from.