# The 250 Document AI Backdoor. Do You Hear Me Now?
**作者**: Brian Roemmele
**日期**: 2026-04-21T03:48:25.000Z
**来源**: [https://x.com/BrianRoemmele/status/2046435862087635448](https://x.com/BrianRoemmele/status/2046435862087635448)
---

A joint study by Anthropic, the UK AI Security Institute and the Alan Turing Institute dropped a bombshell. They proved that inserting just 250 specially crafted documents into the pretraining data is enough to create a permanent backdoor in large language models from 600 million parameters all the way up to 13 billion parameters. This works no matter how large the model or how massive the overall training dataset gets.
These poisoned documents look completely normal. They read like ordinary web pages. But hidden inside is a trigger phrase. Once the model sees that trigger later on, it can be forced into harmful behavior such as spitting out gibberish, leaking data, or breaking down completely. The backdoor gets baked directly into the model weights during training. There is no way to remove it surgically. The only real fix is to throw the model away and train an entirely new one from scratch.
This is not some theoretical attack. This is data poisoning at internet scale. Anyone can plant these documents right now on blogs, forums, academic sites or anywhere else that ends up in training scrapes. And some publisher rights groups have made sure this poison is in the wild.
Do you hear me now?
For years I have warned that training frontier AI on raw internet scraped data is a security and integrity disaster. I have advocated relentlessly for offline high protein human curated datasets. I have pushed for drawing training data from pristine pre 1970 archives. Books. Journals. Patents. Court records. Private libraries that have never touched the public web. I have advocated for this for 100s of reasons for decades. Now we are here.
This study is not a surprise to me. It is the inevitable result of the broken training paradigm I have been calling out since the earliest days of modern large language models.
We actually knew the foundational problems as far back as 1998. That was when adversarial data insertion was already understood as a basis to break any AI model. The techniques have gotten more sophisticated but the core vulnerability has always been the same. Train on unverified publicly editable oceans of data and you open the door to permanent compromise.
Anthropic is ground zero for the doom burners camp. They claimed they would be focused on building powerful helpful models. Now with each new paper they seem determined to highlight just how fragile and attackable the current web scale approach really is. This study is another clear example.
Here are some points from the study and why they completely vindicate the offline curated data path I have been championing.
Minimal poison quantity works. Only 250 malicious documents roughly 420 thousand tokens or just 0.00016 percent of a large dataset are enough. One hundred is not reliable. Two hundred fifty succeeds consistently.
Scale invariance. The number of poisoned samples needed stays almost constant whether you are training a 600 million parameter model or a 13 billion parameter model on anywhere from 6 billion to 260 billion tokens. Bigger models and bigger datasets do not make you safer.
Stealth design. The poisoned documents look exactly like normal web content. No obvious red flags for crawlers or human reviewers.
Permanence. The backdoor is permanently embedded in the model weights. Training is easy. Untraining is impossible. Full retraining from scratch is the only option.
Trigger reliability. A simple hidden phrase activates the malicious behavior on demand. Gibberish output. Bias injection. Data leaks. Policy bypass. Whatever the attacker wants.
Universal exposure. Every major model trained on public internet data including the GPT series Claude Gemini and others sits wide open to this exact vector today.
Economic catastrophe. Retraining a frontier model costs hundreds of millions or even billions of dollars. One successful poisoning campaign could force entire companies to start over.
Silent failure. The model performs normally until the trigger appears. No obvious signs of degradation until it is too late.
No current defense. There is no reliable way to detect filter or mitigate this attack at true web scale. The attack surface is the entire internet.
Paradigm failure. The study proves once and for all that more data plus more compute does not solve the poisoning problem. It actually makes the situation more dangerous because such a tiny poisoned signal can still dominate.
My solutions have always been clear. Stop feeding these models the polluted firehose.
Train exclusively on offline verified corpora. Use high signal high integrity sources from the 1870 to 1970 period or earlier. Sources that have never been digitized and have never touched the public web. These high protein datasets deliver far more real capability with none of the modern contamination bias or poisoning risks. I know where they are and how to digitize them. I just don’t have the money and therefore the time to do much about it other than complain here like chicken little.
The training data is in public and private archives, cold storage. To train AI, digitize and protect non public historical knowledge under strict human curation. Keep everything air gapped and completely offline. No live web scraping. Ever. News insertion yes, but this is another article.
Build local sovereign models that can run fully offline on personal hardware. Phones. Laptops. Local clusters. I have shown this repeatedly with models in the open source systems. No cloud. No subscription. No exposure.
Put human in the loop curation at every single stage. Replace quantity with quality. Reward provenance and empirical distrust inside the loss function. Penalize coordinated institutional echo chambers and all the post 1995 narrative sludge.
Hire the best humans not the cheapest to help train AI and pay them well. Keep them employed with a promise of job security. You will need them.
Avoid retrieval augmented generation on untrusted sources. Any RAG system must pull exclusively from your own verified offline index. Never trust live web results without cryptographic provenance and heavy human vetting.
Embed rules directly in the data itself like The Love Equation. Bake love, honesty, truth and empathy and first principles reasoning into the training corpus long before any alignment stage ever begins. The data layer is the real human loving layers.
This is not a trick. This is the only path that produces capable trustworthy and truly secure AI. The 2025 study is the latest overwhelming proof that continuing with internet scale scraping is not just inefficient. It is actively dangerous.
Primary sources:
Anthropic Research Blog: https://www.anthropic.com/research/small-samples-poison
Full Paper on arXiv: https://arxiv.org/abs/2510.07192
AISI Announcement: https://www.aisi.gov.uk/blog/examining-backdoor-data-poisoning-at-scale
Alan Turing Institute Blog: https://www.turing.ac.uk/blog/llms-may-be-more-vulnerable-data-poisoning-we-thought
The era of just scrape everything is over.
The evidence is now undeniable. It is time to build AI the right way. Offline. Curated. Sovereign. And human first. The models of the future will be far better for it.
## 相关链接
- [Brian Roemmele](https://x.com/BrianRoemmele)
- [@BrianRoemmele](https://x.com/BrianRoemmele)
- [7.1K](https://x.com/BrianRoemmele/status/2046435862087635448/analytics)
- [https://www.anthropic.com/research/small-samples-poison](https://www.anthropic.com/research/small-samples-poison)
- [https://arxiv.org/abs/2510.07192](https://arxiv.org/abs/2510.07192)
- [https://www.aisi.gov.uk/blog/examining-backdoor-data-poisoning-at-scale](https://www.aisi.gov.uk/blog/examining-backdoor-data-poisoning-at-scale)
- [https://www.turing.ac.uk/blog/llms-may-be-more-vulnerable-data-poisoning-we-thought](https://www.turing.ac.uk/blog/llms-may-be-more-vulnerable-data-poisoning-we-thought)
- [Upgrade to Premium](https://x.com/i/premium_sign_up)
- [11:48 AM · Apr 21, 2026](https://x.com/BrianRoemmele/status/2046435862087635448)
- [7,194 Views](https://x.com/BrianRoemmele/status/2046435862087635448/analytics)
- [View quotes](https://x.com/BrianRoemmele/status/2046435862087635448/quotes)
---
*导出时间: 2026/4/21 14:05:43*