Claude Agent Dreaming: Memory, Outcome Grading, Orchestration

Anthropic’s preview replays prior agent sessions into memory, while separate public-beta tools grade results against rubrics and coordinate parallel subagents.

A scheduled review of what an agent already did

Dreaming is a scheduled process, shipped as a research preview, that runs between an agent's sessions rather than during one. It rereads past sessions and the agent's memory store, looking for recurring mistakes, workflows the agent converges on repeatedly, and preferences a team keeps repeating, then writes what it finds back into memory. [2]

Atlas interpretation: Anthropic's own comparison is to sleep consolidation in a brain: replay of recent experience that turns into longer-term memory. The mechanism is narrower than that framing suggests. It is pattern-surfacing over a session log and a memory store, not anything that runs while the agent is otherwise idle, and the improvement it produces depends on the memory store staying high-signal rather than accumulating noise across many sessions. [2]

A grader that can send work back for revision

Outcomes, shipped in public beta, lets a developer define a rubric for what a good result looks like. A grader separate from the agent that did the work checks its output against that rubric, and the agent revises until the output clears the bar rather than returning on the first attempt. [3]

Anthropic reported task-success improvements of up to 10 percentage points in its own testing, with the largest gains on file generation: 8.4 points on Word documents and 10.1 points on PowerPoint decks. Those figures are Anthropic's internal testing, not an independently reproduced benchmark. [2]

One lead agent, several specialists, one filesystem

Multiagent orchestration, also public beta, lets a lead agent delegate parts of a task to specialist subagents that run in parallel on a shared filesystem. Each specialist carries its own model, prompt and tools, and the Claude Console shows how work was delegated and executed across the group. [2][3]

Anthropic named Harvey, Netflix, Spiral and Wisedocs as early users of the combined features, and paired the release with webhooks for outcome-based automation and platform expansion to London and Tokyo. [2]

The gated features from April, opened up and named

Atlas interpretation: Managed Agents launched in April 2026 with self-evaluation and multi-agent coordination already present, but held behind a separate request-access research preview rather than shipped alongside the rest of the runtime.This announcement is where those capabilities got names, a wider release, and a specific new one, dreaming, that the April launch did not include. Anthropic describes the sequence as one product maturing rather than three separate features, which matters for reading the reported improvements: the outcomes grading numbers describe the same runtime that launched a month earlier, not a new baseline. [2]

Sources

  1. Anthropic will let its managed agents dream

    The New Stack · May 6, 2026

  2. New in Claude Managed Agents: dreaming, outcomes, and multiagent orchestration

    Anthropic · May 19, 2026

  3. Code w/ Claude SF 2026: Building on the AI exponential

    Anthropic · May 12, 2026