Scaling Monosemanticity: Claude 3 & Golden Gate Claude

Anthropic used sparse autoencoders to extract millions of interpretable features from Claude 3 Sonnet. See its method, safety caveats and 24-hour Golden Gate Claude demo.

Training a bigger dictionary on a production model

The paper trained sparse autoencoders, or SAEs, on the residual stream activations halfway through Claude 3 Sonnet, a production model rather than the small one-layer transformer the same team had used eight months earlier in a predecessor paper, Towards Monosemanticity. An SAE decomposes a layer's activations into a much larger set of sparse, mostly-inactive components, on the premise that each component corresponds to a single interpretable concept rather than being entangled with many others. [1]

The team trained three SAEs of increasing size, roughly 1 million, 4 million and 34 million features, and used a scaling laws analysis to pick the training budget for the largest run. Across all three, fewer than 300 features fired on a given token on average, and the reconstructed activations captured at least 65% of the variance in the model's original activations. [1]

Bridges, brains and features that generalize past text

Among the millions of extracted features, the paper singles out one that activates on descriptions and references to the Golden Gate Bridge, and a related one for tourist landmarks generally. Both fired not only on English text but on the same concept described in Japanese, Chinese, Greek, Vietnamese and Russian, and on images of the bridge, even though the dictionary learning was performed only on text data. [1]

Atlas interpretation: That a feature trained purely on text also fires on a photo of the same landmark is the central claim of the paper: it is evidence the feature tracks a concept rather than a pattern of tokens, which is what makes the later steering result legible as manipulating an idea and not just a string match. [1]

Features for bias, sycophancy and deception, with a caveat

The paper reports finding features it calls safety-relevant: unsafe code and backdoors, overt slurs and subtler bias, sycophancy, lying and power-seeking including what it terms treacherous turns, and dangerous or criminal content such as instructions related to bioweapons. It checked that several of these features were not merely correlated with the topic but causally connected to it, by artificially increasing a feature's activation and observing the model's output shift toward the associated behavior. [1]

The authors caution against reading too much into a feature's existence: finding a feature that activates on deception is not the same as showing the model is currently being deceptive, and they describe the safety work as very preliminary, with further work needed to understand the implications. [1]

Golden Gate Claude, briefly, in public

In the paper, clamping the Golden Gate Bridge feature to ten times its maximum observed activation caused the model to describe itself as the bridge when asked about its physical form. Two days later, on May 23, 2024, Anthropic put a version of Claude tuned this way on claude.ai for the public to try, calling it a research demo and taking it down after 24 hours. [1][2]

Anthropic's own examples of the demo showed it recommending a user spend ten dollars driving across the Golden Gate Bridge to pay the toll, and turning a request for a love story into one about a car that cannot wait to cross its beloved bridge. The announcement described the change as a precise, surgical edit to the model's internal activations rather than a system prompt or fine-tune. [2]

A small slice of a much larger map

The paper is explicit that the millions of features it found are a small subset of all the concepts the model has learned, and estimates that finding a feature for a concept that appears only once in a billion training tokens would require a dictionary on the order of a billion features, a run the paper says would cost more compute than training the underlying model. [1]

Atlas interpretation: The headline result is therefore a proof that the method scales to a production model at all, not a completed inventory of what that model represents internally. The gap between millions of found features and the unmapped remainder is the reason the paper reads as a first detailed map rather than a finished one. [1]

Sources

  1. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

    Anthropic · May 21, 2024

  2. Golden Gate Claude

    Anthropic · May 23, 2024