Towards Monosemanticity
Dictionary learning on a small transformer extracts 4,000+ interpretable features from a 512-neuron layer, a better unit of analysis than single neurons.
Shows that superposed neuron activations can be decomposed with a sparse autoencoder into sparse features that humans can label (DNA sequences, legal language, HTTP requests, Hebrew text). Starts the dictionary-learning line that Scaling Monosemanticity (2024-05) applied to a production Claude model.
- Date
- Thursday, 5 October 2023
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Features from one MLP layer | 4,000+ from a 512-neuron layer of a small one-layer transformer | company |
Date is Anthropic's research-page date; the Transformer Circuits thread lists the paper under Oct 2023 and the exact day was not independently confirmed. Independent concurrent work on sparse autoencoders exists (not reviewed here).
Sources
- www.anthropic.com/research/towards-monosemanticity-decomposing-language-models-with-dictio
- transformer-circuits.pub/
This record was checked against its sources on 6 October 2026. How we check