Claude Opus 4.6: 1M Context, Pricing & Adaptive Reasoning

Anthropic released Opus 4.6 across Claude, APIs and cloud platforms, adding a beta one-million-token context window, adaptive thinking and context compaction.

The flagship moved from one large answer to a longer-running task

Anthropic released Opus 4.6 on Claude, through its API and through major cloud platforms. Developers could call it as claude-opus-4-6. Standard API input and output remained $5 and $25 per million tokens, the same base prices announced for Opus 4.5. [1][2]

The release added Anthropic's first one-million-token context window for an Opus model, initially in beta. That window was an API feature, not a promise that every Claude interface would accept a million tokens. Prompts above 200,000 tokens also used higher long-context prices of $10 per million input tokens and $37.50 per million output tokens. [1]

Reasoning effort and context became controls

Adaptive thinking let the model decide when and how much extended reasoning to use. API users could also set effort to low, medium, high or max. This separated two choices that are often blurred together: which model to call, and how much work to ask it to spend on this request. [1]

A beta context-compaction feature summarized and replaced older conversation material as a task approached its limit. In Claude Code, Anthropic also introduced research-preview agent teams for parallel, largely independent work. Both features address coordination over time, not just the quality of a single completion. [1]

Atlas interpretation: Compaction is a trade: it makes room for more work by replacing detail with a summary. For a long coding or research task, the practical question is whether the retained summary preserves the constraints, decisions and evidence the model will need later. A larger context window delays that trade but does not remove it. [1]

The benchmark gains came from specific settings

Anthropic reported 76 percent on a one-million-token, eight-needle MRCR retrieval test, compared with 18.5 percent for Sonnet 4.5. This tests whether the model can recover scattered information from a very long input; it is not a general score for reasoning over every million-token document. [1]

The launch's Humanity's Last Exam result used tools, web search, code execution, compaction and up to three million total tokens at maximum effort. That is evidence about an agentic research configuration. It should not be read as the result a short, tool-free prompt receives by default. [1]

Sources

  1. Introducing Claude Opus 4.6

    Anthropic · Feb 5, 2026

  2. Introducing Claude Opus 4.5

    Anthropic · Nov 24, 2025