Claude 2 Launch: Public Access, Context, and Benchmarks

Anthropic opened Claude 2 to consumers in the US and UK with a 100,000-token context window inherited from Claude 1.3 and self-reported benchmark gains.

Anthropic called it a tweak, not a new generation

Sandy Banerjee, Anthropic's head of go-to-market, told TechCrunch that "Claude 2 isn't vastly changed from the last model, it's a product of our continuous iterative approach to model development," and characterized it as a tweaked version of Claude 1.3 rather than a new creation. [3]

The stated change was mostly data: Claude 2 trained on a more recent mix of web text, licensed third party data sets, and voluntarily supplied user data through early 2023, roughly 10% of it non-English. [3]

Atlas interpretation: The name implies a generational jump that Anthropic's own description does not claim. A version number is a marketing decision, and the event's title above inherits it. The more accurate description, in Anthropic's own telling, is a retrained update to an existing model rather than a new architecture. [3]

The 100,000 token context window predates Claude 2

TechCrunch reported that Claude 2's context window, 100,000 tokens, is "the same size of Claude 1.3's," a capability Claude had already carried since a prior update. Claude 2 could theoretically support 200,000 tokens, but Anthropic chose not to enable that at launch. [3]

Reuters described the launch itself as Anthropic having "widened consumer access to its chat program Claude," with businesses able to build on the model and, for the first time, consumers in the US and UK able to sign up and use it online, access Anthropic had previously limited. [2]

Atlas interpretation: The 100,000-token window was briefly a competitive advantage, but it was not new to this release. What was new on July 11 was who could reach it: a public sign-up in two countries replacing a partner list, not a technical leap in how much text the model could hold. [3][2]

What the benchmark gains measured, and what they didn't establish

Anthropic reported Claude 2 scoring 76.5% on the multiple choice section of the bar exam, up from 73.0% for Claude 1.3, 71.2% on the Codex HumanEval coding benchmark versus 56.0%, and 88.0% on the GSM8K grade school math set versus 85.2%. TechCrunch's independent writeup of the same figures added that Claude 2 could also pass the multiple choice portion of the US Medical Licensing Exam. [3]

Reuters set Claude 2's bar exam score against GPT-4's reported 75.7% on the same Multistate Bar Exam multiple choice section, and noted that an Anthropic spokesperson "did not immediately answer whether the result was directly comparable with OpenAI's." [2]

Atlas interpretation: A shared exam name suggests a shared yardstick, but neither company published matching test conditions, and Anthropic would not confirm the two scores measured the same thing. The comparison reads as apples to apples in coverage that quotes both numbers side by side; Reuters' own caveat is the reason to treat it as suggestive rather than settled. [2]

A safety claim with no disclosed definition

Anthropic said Claude 2 was "2x better" at giving harmless responses than Claude 1.3 on an internal evaluation, and separately that it had doubled the model's performance on a safety evaluation. Reuters also reported the model was less likely than its prior version to produce offensive or dangerous responses. [2]

TechCrunch pressed Anthropic on what the "2x better" figure actually measured and reported that the company would not say directly, whether it meant fewer sexist or racist responses, less endorsement of violence or self-harm, or less misinformation. A same-day whitepaper described testing 328 prompts, including public jailbreak prompts, and disclosed that at least one jailbreak still produced a harmful response from Claude 2, at a lower rate than Claude 1.3 but still notable given production-scale prompt volumes. Anthropic separately advised against using Claude 2 in applications involving physical or mental health, or in high stakes situations where a wrong answer would cause harm. [3]

Atlas interpretation: A doubling implies a specific, measured quantity, but the category being doubled was never named, and Anthropic declined to name it when asked directly. That gap, alongside a same-day advisory against high stakes use, is a more honest picture of the safety claim's weight than the headline number alone. [3]

Sources

  1. Claude 2

    Anthropic · Jul 11, 2023

  2. Google-backed Anthropic startup upgrades its Claude AI, better at coding

    Reuters · Jul 11, 2023

  3. Anthropic releases Claude 2, its second-gen AI chatbot

    TechCrunch · Jul 11, 2023