AI Research Atlas

Scaling Monosemanticity

Anthropic · 21 May 2024

Sparse autoencoders extract millions of interpretable features from the middle layer of Claude 3 Sonnet, including the Golden Gate Bridge feature.

First detailed look inside a production-grade LLM per Anthropic. Features include scam emails, sycophantic praise, code backdoors and bias. Amplifying the Golden Gate Bridge feature made Claude fixate on the bridge, showing features are causal handles. Authors say a full feature set is cost-prohibitive.

Date
Tuesday, 21 May 2024
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Features extractedmillions
from the middle layer of Claude 3 Sonnet
company

'First' is Anthropic's own wording. Understanding what a feature represents does not yet explain how the model uses it, per the authors.

Sources

  1. www.anthropic.com/research/mapping-mind-language-model
  2. transformer-circuits.pub/

This record was checked against its sources on 6 October 2026. How we check

Related