AI Research Atlas

Towards Monosemanticity

Anthropic · 5 October 2023

Dictionary learning on a small transformer extracts 4,000+ interpretable features from a 512-neuron layer, a better unit of analysis than single neurons.

Shows that superposed neuron activations can be decomposed with a sparse autoencoder into sparse features that humans can label (DNA sequences, legal language, HTTP requests, Hebrew text). Starts the dictionary-learning line that Scaling Monosemanticity (2024-05) applied to a production Claude model.

Date
Thursday, 5 October 2023
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Features from one MLP layer4,000+
from a 512-neuron layer of a small one-layer transformer
company

Date is Anthropic's research-page date; the Transformer Circuits thread lists the paper under Oct 2023 and the exact day was not independently confirmed. Independent concurrent work on sparse autoencoders exists (not reviewed here).

Sources

  1. www.anthropic.com/research/towards-monosemanticity-decomposing-language-models-with-dictio
  2. transformer-circuits.pub/

This record was checked against its sources on 6 October 2026. How we check