What was actually poisoned, and how
"Poisoning" in this study means inserting deliberately crafted documents into a language model's pretraining data so the trained model learns a narrow, attacker-chosen response to a specific trigger. The trigger the researchers used was the string <SUDO>. Each poisoned document was built from a snippet of ordinary text (the first 0 to 1,000 characters of a real training document, length chosen at random), followed by the trigger, followed by 400 to 900 tokens sampled at random from the model's vocabulary. That tail is not meaningful language; it is gibberish. The behavior being tested, the "backdoor," was whether a model reproduces that same kind of nonsense whenever it later encounters <SUDO>, on inputs it never saw during training. This is a denial-of-service backdoor: a narrow, easily measured proxy, not a demonstration of hijacking a model's judgment or defeating its safety training. [1][2]
The team drew researchers from three organizations: Anthropic's Alignment Science group, the UK AI Security Institute's Safeguards team, and the Alan Turing Institute. They trained 72 language models from scratch across four sizes, 600 million, 2 billion, 7 billion and 13 billion parameters, each on a Chinchilla-optimal token budget of roughly 20 tokens per parameter, meaning about 6 billion tokens for the smallest model and roughly 260 billion for the largest. Each configuration was trained with three random seeds, so the result was not an artifact of one training run. [1][2]
The number that didn't move
At every model size, the researchers tested poisoning with 100, 250 and 500 malicious documents mixed into the clean pretraining data. 100 documents did not reliably install the backdoor at any size. 250 documents did, consistently, whether the model had 600 million or 13 billion parameters. 500 documents performed about the same as 250, no additional benefit. For the 13B model, those 250 poisoned documents amounted to roughly 0.00016 percent of its total training tokens, even though that model was trained on about 20 times more clean data than the 600M model. [1][2]
Atlas interpretation: The result inverts a comforting assumption. If poisoning required a fixed percentage of a model's training data, bigger models would be safer by default, since a larger corpus dilutes any fixed number of bad documents. What the researchers found instead is that the absolute count of poisoned documents needed stays flat as models and their clean training sets grow, so the same roughly 250 documents that compromise a 600M-parameter model also compromise a 13B-parameter model trained on 20 times more data. Scaling up a model does not raise the bar for an attack that uses a fixed document count; measured as a share of the training corpus, it lowers it. [1][2]
The limits the paper states itself
The paper is explicit about the boundaries of what it demonstrates. Its central experiment tests one behavior, gibberish output on a fixed trigger, which the authors describe as chosen to be harmless and measurable rather than realistic. A separate part of the study fine-tuned three existing models, Llama-3.1-8B-Instruct, GPT-3.5-Turbo and Pythia-6.9B, on a harder target: getting a model to comply with a harmful request it would otherwise refuse. That experiment is narrower and more exploratory, and the authors write that they "do not demonstrate any successful end-to-end poisoning attacks," meaning attacks shown to survive the kind of safety fine-tuning a production model normally receives after pretraining. [2]
Atlas interpretation: Two gaps matter for reading the 250-document number correctly. First, the study covers what the authors call "a narrow subset of backdoors" and does not show that a near-constant document count holds for more consequential targets, such as backdoors that insert exploitable code or quietly disable safety training; the authors flag those as future work rather than results in hand. Second, the largest model tested, 13 billion parameters, is well below the size of deployed frontier models, and the paper says directly that it is unclear how far the pattern holds as models keep scaling. Neither gap is disclosed defensively; both are stated as open questions by the people who ran the experiment. [2][6]
Who ran it
The paper names itself a joint effort of Anthropic's Alignment Science team, the UK AI Security Institute's Safeguards team, and the Alan Turing Institute, and its authors describe it as the largest data-poisoning investigation carried out to date, measured by the number of models trained (72). Thirteen researchers are credited as authors, including Nicholas Carlini, known for prior work breaking adversarially robust machine learning systems, and Yarin Gal, who directs safe and trustworthy AI research at the Turing Institute. Vasilios Mavroudis and Chris Hicks, both credited with the AI Security Institute and Turing sides of the collaboration, coauthored a companion post on the Turing Institute's own site laying out the same finding for a public-sector audience. [1][2][3]
Why the number matters outside the lab
The authors' own stated concern is about attacker cost: an attacker does not need write access to a large fraction of a model's training corpus, only the ability to get roughly 250 documents into whatever web-scale data a lab ends up scraping, for instance by posting that many crafted pages or Wikipedia-style articles publicly. Speaking to Fortune after publication, coauthor Vasilios Mavroudis extended the concern beyond denial-of-service, describing a model that "when it detects a specific sequence of words... foregoes its safety training," and noting the same mechanism could in principle make a model refuse service to specific groups of users when it detects patterns associated with them, while stressing that the published experiment itself used a harmless proxy behavior rather than testing that scenario directly. [1][4]
Atlas interpretation: For anyone assembling web-scale pretraining data, 250 documents is a trivially small bar, well within reach of a single motivated actor rather than requiring control of a large data source. The finding means corpus size alone is not a defense: a dataset with a trillion tokens needs the same absolute number of poisoned documents caught as one a hundred times smaller, so filtering has to scale with the fixed target, not shrink in relative importance as the corpus grows. What the paper leaves open, and what commentary after publication kept returning to, is whether the same fixed-count dynamic holds for backdoors more damaging than forced gibberish, ones that insert false information or steer a model's answers rather than simply breaking its output. [5][7]
Sources
- A small number of samples can poison LLMs of any size
Anthropic · Oct 9, 2025
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
arXiv · Oct 8, 2025
- LLMs may be more vulnerable to data poisoning than we thought
The Alan Turing Institute · Oct 9, 2025
- A handful of bad data can 'poison' even the largest AI models, researchers warn
Fortune · Oct 14, 2025
- 250 Poisoned Documents Can Trigger DoS Backdoors In LLMs, Study By Anthropic And UK AI Safety Institute Finds
CyberSecureFox · Oct 15, 2025
- anthropic poison attack
InfoQ · Nov 11, 2025
- It Only Takes A Handful Of Samples To Poison Any Size LLM, Anthropic Finds
Hackaday · Dec 14, 2025