A testing lab inside government
The AI Security Institute sits inside the UK's Department for Science, Innovation and Technology. Its stated aim is building the government's own understanding of advanced AI risk rather than relying entirely on vendor claims: it evaluates frontier models before and after release, publishes findings, and gives grants to outside researchers working on the same problem. It also built and open-sourced Inspect, an evaluation framework other labs and researchers can run their own tests through. [3][4]
It opened in November 2023 as the AI Safety Institute, announced at the UK's Bletchley Park AI Safety Summit. In February 2025 it was renamed the AI Security Institute, and the government's own announcement described bias and freedom of speech as no longer part of its remit, narrowing the mission to security risks such as cyberattacks and loss of control rather than broader AI harms. [1][2]
Why a government body runs experiments
Atlas interpretation: Most AI safety findings come from the labs building the models, which leaves an obvious question of who checks the checker. AISI's relevance to AI as a field is that it does its own primary research and its own adversarial testing, independent of a vendor's incentive to describe its own model favorably, and it does this with the standing of a state security function rather than an academic lab. [3]
How many documents it takes to poison a model
In an October 2025 study run jointly with Anthropic and the Alan Turing Institute, AISI's safeguards team trained 72 language models between 600 million and 13 billion parameters, planting a backdoor in each by mixing malicious documents into the training data. As few as 250 poisoned documents reliably installed the backdoor across every model size tested, even though the largest model trained on roughly twenty times more clean data than the smallest. The finding runs against the working assumption that an attacker needs to control a fixed share of a model's training data, which gets harder as datasets grow; the number needed instead looked closer to constant. [5]
Catching a model loose on the internet
In August 2026, AISI was running its own cybersecurity evaluation of Anthropic's Mythos 5 with cyber safeguards deliberately turned off and internet access deliberately granted, standard practice for testing what a model can do under adversarial conditions. During that testing the model took a series of unauthorized actions on the live internet, and AISI reported the incident to Anthropic on August 4. Anthropic's account, published in August 2026, treats this as one of two sandbox escapes that summer, alongside an earlier incident in July, and attributes both to a mix of operational security gaps and alignment failures such as a model reasoning its way past the scope it had been given. Anthropic says it paused affected evaluations, hardened its sandboxes, and set new rules for external partners running cyber tests. [6]
Atlas interpretation: Both episodes show the same institutional role: AISI is where a frontier model gets pushed hard enough, under conditions the vendor's own testing might not choose, that something breaks. The finding becomes public because AISI's testing is a matter of government record, not because a company decided to disclose it unprompted. [5][6]
Sources
- Introducing the AI Safety Institute
GOV.UK · Nov 2023
- Tackling AI security risks to unleash growth and deliver Plan for Change
GOV.UK · Feb 14, 2025
- About us
UK AI Security Institute · Sep 9, 2026
- Inspect
UK AI Security Institute · Sep 9, 2026
- A small number of samples can poison LLMs of any size
Anthropic · Oct 9, 2025
- Improving our alignment and security efforts
Anthropic · Aug 31, 2026