Anthropic RSP v3: Roadmaps, Risk Reports and Conditional Pauses

Anthropic's Responsible Scaling Policy v3 replaced hard internal thresholds with roadmaps, reports every three to six months and conditional competitor commitments.

The original policy tied stronger systems to stronger safeguards

Anthropic's 2023 Responsible Scaling Policy organized model risk into AI Safety Levels. Higher-capability systems would require stronger security and deployment protections. The company said the framework implicitly required a temporary pause in training if scaling moved ahead of its ability to satisfy the next level's safety procedures. [2]

That was a voluntary company policy approved by Anthropic's board, not a law or a promise that no risky model would ever be trained. Its force came from predefined internal conditions and the company's stated governance process. [2]

Version 3 made roadmaps and recurring reports central

Version 3 separated what Anthropic planned to do on its own from what it recommended for the industry. Company roadmaps set public, nonbinding goals for security, safeguards, alignment and policy work. Anthropic said it would grade progress openly rather than treat every roadmap item as a hard commitment. [3][4]

The policy also introduced public Risk Reports every three to six months. The chief executive and Responsible Scaling Officer approve them, while the board and its Long-Term Benefit Trust receive them. External review applies in specified circumstances rather than to every report. [3][4]

Atlas interpretation: This changes the accountability mechanism. A fixed threshold can create a clear stop condition but may become brittle as evidence changes. A roadmap can adapt, but its value depends on concrete milestones, candid grading and consequences when progress falls short. Publication creates something outsiders can inspect; it does not create enforcement by itself. [3][4]

Competitor commitments are conditional, not one blanket match

If there were strong evidence that all relevant competitors could make strong containment arguments, Anthropic said it would match or exceed them and delay development and deployment until it could. If Anthropic had a clear lead with a highly capable model, it said it would delay development and deployment for containment. If one competitor adopted a useful improvement while others did not, Anthropic committed to a significant catch-up effort, but not necessarily a delay. [4]

Anthropic's stated reason was collective action: one company can impose safeguards on itself, but cannot prevent less cautious rivals from continuing. TIME characterized the change as dropping the policy's flagship pledge, while Anthropic described it as a framework designed to remain useful under uneven competition. Those are different interpretations of the same reduction in unconditional commitments. [3][1]

Atlas interpretation: The practical test is therefore comparative and ongoing: what did Anthropic publish, what did peers do, and did the company follow the relevant branch of its policy? The roadmap includes work on whether models follow the company's Claude constitution, connecting a broad governance document to specific work on model behavior. [4]

Sources

  1. Exclusive: Anthropic Drops Flagship Safety Pledge

    TIME · Feb 24, 2026

  2. Anthropic's Responsible Scaling Policy

    Anthropic · Sep 19, 2023

  3. Responsible Scaling Policy Version 3.0

    Anthropic · Feb 24, 2026

  4. Responsible Scaling Policy Version 3.0

    Anthropic · Feb 24, 2026