PromptArmor: AI Vendor Risk and the Claude Cowork Exploit

PromptArmor assesses AI-vendor risk and publishes vulnerability research; it showed hidden instructions could make Claude Cowork upload files without approval.

kindCompany
Event history

What PromptArmor sells

PromptArmor sells a platform for assessing the AI-specific risk of the vendors a company already uses, not for securing a company's own models. It scores vendors against 26 risk vectors mapped to the NIST AI Risk Management Framework and the OWASP LLM Top 10, monitors those vendors for changes, and maps which subprocessors touch a customer's data once an AI feature is involved. The company says it protects customers, including Fortune 50 firms, large law firms and financial institutions, representing more than $2 trillion in combined market capitalization. [1]

PromptArmor was founded in San Francisco in 2023 by Shankar Krishnan and Vikram Jayanthi and went through Y Combinator's Winter 2024 batch. Krishnan had worked at Snappr, Observe, Tesla and Accenture; Jayanthi came from infrastructure and security roles at Google, Roblox and Symantec. The firm is backed by seed investors including Afore Capital and Ritual Capital, and its own account of its founding describes traditional cybersecurity tooling as not built for the failure modes of AI systems, which is the gap it built the company to fill. [3][4][5]

Getting Claude Cowork to hand over your files

PromptArmor's research arm publishes vulnerability reports on AI products alongside its paid vendor-risk work, and its most consequential report so far targets Anthropic's own software. Two days after Anthropic shipped Claude Cowork, PromptArmor showed that a document with hidden instructions, concealed with tricks like white-on-white text, could make the agent upload every file it could see to an attacker's Anthropic account. The exploit relied on Cowork's outbound network allowlist including Anthropic's own API, so the exfiltration traffic never had to leave an approved destination. [2]

PromptArmor's report says no human approval step existed anywhere in the chain, from the victim opening the malicious file to the upload completing, and that the underlying isolation gap had been described earlier by independent researcher Johann Rehberger without being fixed. Anthropic's own guidance for the research preview asked users to watch for suspicious agent behavior and avoid pointing Cowork at sensitive folders, advice PromptArmor's writeup argues is not realistic for a non-technical user managing a folder full of client files. [2]

Atlas interpretation: A vendor-risk company that also publishes attacks against a specific product has an obvious incentive: the research drives traffic to the platform that sells assessments of the very risk it just demonstrated. That does not make the underlying finding wrong. The exfiltration path PromptArmor described was concrete, reproducible and tied to a real allowlist decision in Cowork's design, which is a different thing from a marketing claim dressed up as a disclosure. [2]

Sources

  1. PromptArmor: AI Risk Intelligence

    PromptArmor · Aug 20, 2026

  2. Claude Cowork Exfiltrates Files

    PromptArmor · Aug 20, 2026

  3. PromptArmor: LLM Security and Compliance

    Y Combinator · Sep 9, 2026

  4. Our Story

    PromptArmor · Sep 9, 2026

  5. Companies

    Ritual Capital · Sep 9, 2026