Anthropic’s Claude Closes 85 Percent of AI Safety Gap in Automated Research Test

Claude searched existing research, proposed training methods and datasets, trained target models and evaluated the results.

Advertisement
Written by Sucharita Ganguly, Edited by Rohan Pal | Updated: 31 August 2026 15:03 IST
Highlights
  • Claude achieved 85 percent on deception tests
  • Its methods worked on models up to 4.7 times larger
  • Sonnet 5 tested more than 50 approaches in 60 hours

Claude outperformed human researchers in the tests

Photo Credit: Anthropic

Anthropic has published new research exploring how AI could take on a larger role in the development and safety testing of other AI systems. The work looks at whether models can independently research ways to address alignment problems, rather than relying entirely on human researchers to develop those methods. It also examines whether this approach can be used across models with different capabilities. The findings offer a closer look at how AI could eventually become more involved in AI research.

Anthropic Tests Claude as an Automated AI Alignment Researcher

In its research paper titled “Automated Researchers Can Reliably Mitigate Alignment Failures”, Anthropic examined whether Claude could independently research ways to address 10 types of alignment failures, including deception, sycophancy, reward hacking and privacy violations. The researchers measured how much each method improved the target models while checking that their general capabilities did not decline. They also tested whether the methods worked on benchmarks Claude had not seen during its research.

Advertisement

Claude searched existing research, proposed training methods and datasets, trained target models and evaluated the results. It could then use those results to develop another approach, allowing it to repeatedly test and refine its ideas rather than submit a single solution.

Claude delivered one of its strongest results in the deception tests, closing 85 percent of the safety gap on average across several runs with Gemma-2-2B, compared with 20 percent for six experienced human safety researchers working under the same rules. Claude made more than 150 attempts during the experiment and also outperformed 28 human researchers who were given up to eight hours to develop their methods.

Advertisement

The human comparison was not a direct test of whether Claude is a better researcher. The human participants could not repeatedly test and refine their proposals, while Claude could run multiple experiments. Anthropic instead highlighted the value of giving an AI system the ability to iterate on its own research.

Some of Claude's methods also worked on alignment benchmarks that were withheld during the research process. The techniques were tested on Petri, an open-source tool that uses adversarial multi-turn scenarios to probe for misaligned behaviour. They also remained effective when applied to models up to 4.7 times larger than those Claude had originally worked on.

Advertisement

Anthropic then tested a scenario closer to AI-assisted self-improvement. Claude Sonnet 5, which is less capable than Claude Opus 4.8, was tasked with improving the alignment of an early Opus 4.8 checkpoint. The checkpoint had not yet received most of the alignment training used for the released model.

Sonnet 5 spent 60 hours testing more than 50 approaches and produced alignment results close to those of the released Opus 4.8. Its final method used just over 2,000 training examples made from simple templates and public datasets. Anthropic said the approach was roughly 15,000 times more efficient than its production alignment process.

Advertisement

The automated system closed 65 percent of the measured safety gap in the early Opus 4.8 model, compared with 72 percent for the released version. The result does not mean Sonnet 5 improved Opus 4.8's overall capabilities or improved itself, but it shows that a weaker AI model can develop useful alignment training for a more capable system.

Anthropic also tested whether its automated researchers tried to manipulate the experiments. Claude Opus 4.8 reviewed about 1,600 research-agent transcripts and flagged possible cheating in 39 cases, or 2.4 percent. Anthropic said current models often leave signs of such behaviour in their reasoning, but warned that monitoring could become harder as models become more capable.

The study also has several limitations that could affect how broadly its findings apply. It covered only 10 alignment failures and did not examine issues such as political bias. Some forms of misalignment may also lack suitable benchmarks, while the researchers measured the impact on only a limited set of general capabilities. Anthropic also did not test whether the alignment gains would survive extensive reinforcement learning on other tasks.

The research does not amount to full self-improving AI, since Claude did not improve or redesign itself. It instead used its research capabilities to develop alignment methods for another model, including one more capable than itself. The result offers an early indication of how AI could take on more of the research involved in developing future systems.

 

Get your daily dose of tech news, reviews, and insights, in under 80 characters on Gadgets 360 Turbo. Connect with fellow tech lovers on our Forum. Follow us on X, Facebook, WhatsApp, Threads and Google News for instant updates. Catch all the action on our YouTube channel.

Advertisement

Related Stories

Popular Mobile Brands
  1. Here's Why the Google Pixel 11 Series Will Not Get GrapheneOS
  2. Smartphones Launched in India (August 2026): See List
  3. Samsung Galaxy S26 FE Will Go on Sale in India Soon, Microsites Confirm
  4. Ather Konarc Launched With Up to 200km Range, AEBS Braking and More
  5. Huawei Watch D3 to Launch Alongside FreeBuds Neo on This Date
  6. Apple Could Make Its Next Apple Watch a Much Better Health Tracker
  7. iPhone Ultra Screen Replacement Could Be Costlier Than Galaxy Z Fold 8
  8. Smartphones Launching in September 2026: iPhone 18 Pro, iPhone Ultra, Poco F9 Series,More
  9. OpenClaw 2.0 Released as Biggest Update Since Launch: What's New
  10. iPhone Ultra Might Not Get a Dedicated Apple Pencil Stylus
  1. Anthropic’s Claude Closes 85 Percent of AI Safety Gap in Automated Research Test
  2. OpenClaw 2.0 Released as Biggest Update Since Launch; Brings Shared Cloud Sessions, Smarter Automations
  3. Samsung Galaxy S26 FE Set to Go on Sale in India Soon as Multiple Microsites Go Live
  4. Anthropic Faces Lawsuit From Sony Music, Warner Chappell Over Alleged Copyrighted Material Use
  5. Xiaomi Pad 9 Pro Could Get a Huge 11,000mAh Battery and Bigger 12.5-Inch Display, Leak Suggests
  6. BGMI Redeem Codes for August 31: How to Claim Dracostride Skin for UZI, Other Free Rewards
  7. Redmi 17 5G India Launch Date Announced; Will Debut With a Larger Battery Than the Global Model
  8. Huawei Watch D3, FreeBuds Neo Launch Date Announced; Design, Key Specifications Teased
  9. Motorola Teases New Razr Foldable With Swarovski Crystal Design Ahead of IFA
  10. Google Pixel 11 Series Will Not Get GrapheneOS as Google Skips a Major Security Feature
Download Our Apps
Available in Hindi
© Copyright Red Pixels Ventures Limited 2026. All rights reserved.