Finishing the chain instead of describing it
For several months Cloudflare had been testing security-focused LLMs against its own code, and pointed Mythos Preview, a research model Anthropic supplied under Project Glasswing, at more than fifty of its internal repositories. Cloudflare's account of earlier models is specific: a model would identify an interesting bug, write a description of why it mattered, and stop, leaving the exploit chain unfinished and exploitability an open question. [1]
Cloudflare describes Mythos Preview doing something different: taking low-severity bugs, the kind that traditionally sit in a backlog because no one bug is worth fixing urgently, and reasoning about how to chain several of them into a single working proof. The model writes code meant to trigger the suspected bug, compiles it in a scratch environment, and runs it, so the output is an exploit that either works or doesn't rather than a writeup of a hypothesis. [1]
Atlas interpretation: Cloudflare's own framing, that this is a different kind of tool doing a different kind of work rather than a better version of the same tool, is doing real work in the post: it heads off the obvious next question, which is how much better, by refusing to put a percentage on it. Closing the loop from suspected bug to compiled, running exploit is a change in what the output is, not only in how often it appears. [1]
A harness, not a single query
The system Cloudflare describes runs many agents at once. Tasks pair one attack class with a hint about scope, and hunter agents, the ones that search for bugs, run concurrently, typically around fifty at a time, each fanning out to a handful of exploration subagents of its own. [1]
The model Cloudflare had access to was research-only. Cloudflare states that the Mythos Preview build supplied under Project Glasswing lacked the additional safeguards present in Anthropic's generally available models, and that every vulnerability the work surfaced was triaged, validated, and remediated where needed through Cloudflare's own formal vulnerability management process. [1]
The number Cloudflare didn't put in its own post
Cloudflare's post describes the capability but reports no vulnerability count. Four days later, Anthropic published an initial update on Project Glasswing as a whole, describing the program as roughly fifty partners that had launched the previous month, and stated that Cloudflare specifically had found 2,000 bugs, 400 of them high or critical severity, across its critical-path systems, with a false-positive rate Cloudflare's team considered better than its human testers. [2]
Atlas interpretation: The figures a reader would want to judge the claim, how many bugs, how many were serious, how the false-positive rate compares to a named baseline, come from the model's own maker rather than from Cloudflare's account of its own infrastructure. That doesn't make the numbers wrong, but it means the summary above and Cloudflare's post are describing a capability, while the count that would let someone weigh that capability against the cost of running fifty concurrent agents against production code arrived through a different party's press cycle. [2][1]
What independent reaction focused on
Coverage from outside Cloudflare and Anthropic focused less on the numbers and more on the tone. One technical newsletter's write-up called Project Glasswing an exclusive club of large partner companies, praised Cloudflare's post for describing mechanism rather than making claims, and pushed back specifically on Anthropic co-founder Dario Amodei's public warnings about near-term catastrophic cyber risk, writing that nothing in Cloudflare's post supports being months away from the scenario Amodei has described. [3]
Atlas interpretation: That criticism targets the marketing built around the finding, not the finding itself; the newsletter does not dispute that Mythos Preview closed exploit chains that earlier models left open. The two claims are separable: a model finishing a proof-of-concept it previously would have stopped short of is a real change in what the tool does, and a lab's founder citing that change as evidence for a specific near-term catastrophic timeline is a separate, much larger claim that this report does not itself support. [3][1]
Sources
- Project Glasswing: what Mythos showed us
Cloudflare · May 18, 2026
- Project Glasswing: An initial update
Anthropic · May 22, 2026
- Bytes #488
Bytes · May 19, 2026