A month of testing, thousands of bugs
Anthropic's report describes Claude Mythos Preview as a general-purpose model that was not explicitly trained for cybersecurity work but turned out to be, in the company's words, strikingly capable at computer security tasks. In benchmark exploit development the model reached 595 crashes at the two highest severity tiers, against 150 to 175 for Claude Opus 4.6 and Claude Sonnet 4.6, and it achieved full control-flow hijacks on ten separate, fully patched targets. Over a month of testing it surfaced more than 10,000 high- and critical-severity vulnerabilities across major operating systems and browsers, and Anthropic reports that over 99 percent remained unpatched at the time of writing. [1]
Several of the vulnerabilities the model found were old: a 27-year-old denial-of-service bug in OpenBSD's TCP SACK handling, a 16-year-old memory-corruption flaw in FFmpeg's H.264 decoder that had survived years of fuzzing and code review, and a 17-year-old unauthenticated remote-root hole in FreeBSD's NFS server, tracked as CVE-2026-4747. The report says the model found and exploited the FreeBSD bug without human guidance, building a stack buffer overflow that overwrote a 128-byte buffer with 304 bytes and chaining it into a 20-gadget return-oriented-programming attack split across multiple network packets. [1]
Atlas interpretation: The company frames the FreeBSD result as a difference in kind rather than degree: exploits assembled end to end by the model, not vulnerabilities the model merely described for a person to weaponize. Anthropic's own prior assessment of Opus 4.6, a month earlier, put its autonomous exploit-development success rate near zero, which is the baseline the Mythos Preview numbers are measured against. [1]
Withholding the model instead of shipping it
Anthropic did not release Claude Mythos Preview generally. Its stated concern is a transitional window in which attackers can exploit the same capability faster than defenders can patch against it, particularly for known but unpatched vulnerabilities, where public disclosure hands an attacker a roadmap. Anthropic describes the position as: in the short term, this could help attackers, if frontier labs are not careful about how they release these models. [1]
The model's existence had already become public a week and a half earlier, when a misconfigured Anthropic content cache exposed draft documentation describing it, then called Capybara, as a step change in capability. This report is Anthropic's first acknowledgment of what the model can specifically do. [1][2]
Who gets access, and on what terms
Project Glasswing is Anthropic's invitation-only distribution channel for the preview, built to let critical-infrastructure operators and open-source maintainers find and fix their own bugs before the capability reaches anyone else. Twelve organizations, including Amazon, Apple, Broadcom, Cisco, CrowdStrike, the Linux Foundation, Microsoft and Palo Alto Networks, hold access as core partners doing defensive work, alongside roughly forty additional organizations. Anthropic says the program has no plan for a general public release. [2]
For software that is not open source, testing runs offline, respecting each vendor's own bug bounty terms rather than routing around them. A professional triager validates a finding before it is disclosed, disclosure itself follows a staged 90-plus-45-day coordinated timeline, and Anthropic uses SHA-3 commitments to prove to a vendor that it holds a real vulnerability before describing what it is, so a claim cannot be faked and details are not exposed before a fix ships. [1]
Sources
- Assessing Claude Mythos Preview's cybersecurity capabilities
Anthropic · Apr 7, 2026
- Anthropic debuts preview of powerful new AI model Mythos in new cybersecurity initiative
TechCrunch · Apr 7, 2026