Anthropic Warns That Minimal Data Contamination Can ‘Poison’ Large AI Models

As few as 250 malicious documents can produce a "backdoor" vulnerability in a large AI model, says Anthropic.

Advertisement
Written by Akash Dutta, Edited by Ketan Pratap | Updated: 11 October 2025 13:04 IST
Highlights
  • LLMs can exfiltrate sensitive data when attacker adds a trigger phrase
  • Anthropic says the size of the total dataset does not matter
  • Study breaks the belief that attackers need to control large data portion

UK AI Security Institute and the Alan Turing Institute partnered with Anthropic on this study

Photo Credit: Anthropic

Anthropic, on Thursday, warned developers that even a small data sample contaminated by bad actors can open a backdoor in an artificial intelligence (AI) model. The San Francisco-based AI firm conducted a joint study with the UK AI Security Institute and the Alan Turing Institute to find that the total size of the dataset in a large language model is irrelevant if even a small portion of the dataset is infected by an attacker. The findings challenge the existing belief that attackers need to control a proportionate size of the total dataset in order to create vulnerabilities in a model.

Anthropic's Study Highlights AI Models Can Be Poisoned Relatively Easily

The new study, titled “Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples,” has been published on the online pre-print journal arXiv. Calling it “the largest poisoning investigation to date,” the company claims that just 250 malicious documents in pretraining data can successfully create a backdoor in LLMs ranging from 600M to 13B parameters.

The team focused on a backdoor-style attack that triggers the model to produce gibberish output when encountering a specific hidden trigger token, while otherwise behaving normally, Anthropic explained in a post. They trained models of different parameter sizes, including 600M, 2B, 7B, 13B, on proportionally scaled clean data (Chinchilla-optimal) while injecting 100, 250, or 500 poisoned documents to test vulnerability.

Surprisingly, whether it was a 600M model or a 13B model, the attack success curves were nearly identical for the same number of poisoned documents. The study concludes that model size does not shield against backdoors, and what matters is the absolute number of poisoned points encountered during training.

Advertisement

The researchers further report that while injecting 100 malicious documents was insufficient to reliably backdoor any model, 250 documents or more consistently worked across all sizes. They also varied training volume and random seeds to validate the robustness of the result.

However, the team is cautious: this experiment was constrained to a somewhat narrow denial-of-service (DoS) style backdoor, which causes gibberish output, not more dangerous behaviours such as data leakage, malicious code, or bypassing safety mechanisms. It's still open whether such dynamics hold for more complex, high-stakes backdoors in frontier models.

 

Get your daily dose of tech news, reviews, and insights, in under 80 characters on Gadgets 360 Turbo. Connect with fellow tech lovers on our Forum. Follow us on X, Facebook, WhatsApp, Threads and Google News for instant updates. Catch all the action on our YouTube channel.

Advertisement

Related Stories

Popular Mobile Brands
  1. Haier Mini LED M70 Series With Dolby Vision Launched in India at This Price
  2. Nokia 300 Charge Announced With Torch and 3,700mAh Battery
  3. Portronics Ruffpad Teddy 8.5, Dino 8.5, PupX 10 Debut in India
  4. Here's How Much the Samsung Galaxy S26 Ultra Now Costs After Price Cut
  5. Vivo T5 India Launch Date, Design and Features Teased
  6. Asus TUF Gaming 16 (FX610), Alongside Other TUF Laptops Debut in India
  7. Nothing OS 5.0 With More Than 80 Changes, Upgrades Launched
  8. Sony Xperia 10 VIII With a 50-Megapixel Camera Debuts: See Price
  9. Boltt Ace 5G, Boltt Evo With 6,000mAh Battery Debut in India: See Prices
  10. Googlebook May Debut With Up to 8 Laptops, Including OLED Models
  1. Asus TUF Gaming 16 (FX610) Launched in India With Up to Nvidia GeForce RTX 5060 GPU, Alongside Refreshed TUF A15, A16 and F16
  2. Discord Says It Has Not Received Take-Two Subpoena Over GTA 6 Leaks as More Gameplay Clips Surface Online
  3. Haier Mini LED M70 Series Launched in India With Up to 85-Inch Screen, Dolby Vision Support: Price, Features
  4. ChatGPT Gets New Sticker Maker for Personalised iMessage, WhatsApp Packs
  5. Nothing OS 5.0 Announced With Over 80 Design Changes, Performance Upgrades: See Release Schedule
  6. Googlebook Could Launch With Up to Eight Laptops From Lenovo, HP, Dell, Acer and Asus
  7. Nokia 300 Charge Feature Phone Announced 3,700mAh Battery, Reverse Charging Support: Price, Specifications
  8. Vivo T5 5G India Launch Date Announced; Confirmed to Feature 144Hz Display, Dimensity 7500 Turbo Chip
  9. Sony Xperia 10 VIII Launched With 50-Megapixel Rear Camera, Snapdragon 6 Gen 3 Chipset: Price, Specifications
  10. India Set to Launch First Tokenised Corporate Bonds With Digital Rupee
Download Our Apps
Available in Hindi
© Copyright Red Pixels Ventures Limited 2026. All rights reserved.