One family, three sizes, priced apart
Anthropic released Claude 3 as three models sharing a name and a training recipe but not a price. Opus was pitched as the most capable, Sonnet as the balance of skill and speed, and Haiku as the fastest and cheapest. Opus listed at $15 per million input tokens and $75 per million output tokens; Sonnet at $3 and $15; Haiku at $0.25 and $1.25. Opus and Sonnet shipped that day on claude.ai and the API, general in 159 countries; Haiku followed shortly after. [1]
All three models could take images as input alongside text, a capability Claude had not had before. The training data cutoff was August 2023. Anthropic described the models as tool-using and less prone to refusing borderline but harmless prompts than the prior generation, a change it framed as reduced over-caution rather than reduced safety review. [1][2]
Ahead of GPT-4 on Anthropic's own scorecard
Anthropic's model card reports Opus scoring 86.8% on MMLU (5-shot) against GPT-4's 86.4%, 50.4% on the GPQA Diamond set against 35.7%, 95.0% on GSM8K grade-school math against 92.0%, and 84.9% on HumanEval coding problems against 67.0%. These are Anthropic's own runs, not an independent lab's, and the GPQA number came with a caveat: Opus fell well short of the 60 to 80 percent range domain experts reach on the same questions. [2]
Atlas interpretation: The margins were real but not lopsided, a few points on most tests rather than a different tier of model. What made the release land as a shift was that a lab other than OpenAI was posting the better numbers at all, on the benchmarks OpenAI's own GPT-4 had defined as the ones that mattered. [2]
Opus noticed it was being tested
Anthropic ran a needle-in-a-haystack evaluation, hiding a target sentence inside a large stack of unrelated documents and asking the model to find it. Opus recalled the hidden sentence with 99.4% average accuracy across context lengths, holding 98.3% at the full 200,000-token window Anthropic shipped in production, even though the model was tested against contexts up to 1 million tokens internally. [2]
In some runs, when the planted sentence was an odd one out, such as a claim about a favorite pizza topping dropped into documents about programming and startups, Opus flagged that it looked inserted to test whether the model was paying attention rather than treating it as a genuine fact to retrieve. Anthropic published the exchange in the model card and cautioned that the artificial setup of the test could itself become a limitation as models improved. [2]
Sources
- Introducing the next generation of Claude
Anthropic · Mar 4, 2024
- The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic · Sep 8, 2026