Scaling Monosemanticity
Sparse autoencoders extract millions of interpretable features from the middle layer of Claude 3 Sonnet, including the Golden Gate Bridge feature.
First detailed look inside a production-grade LLM per Anthropic. Features include scam emails, sycophantic praise, code backdoors and bias. Amplifying the Golden Gate Bridge feature made Claude fixate on the bridge, showing features are causal handles. Authors say a full feature set is cost-prohibitive.
- Date
- Tuesday, 21 May 2024
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Features extracted | millions from the middle layer of Claude 3 Sonnet | company |
'First' is Anthropic's own wording. Understanding what a feature represents does not yet explain how the model uses it, per the authors.
Sources
This record was checked against its sources on 6 October 2026. How we check