AI Research Atlas

Natural Language Autoencoders

Anthropic · 7 May 2026

Natural Language Autoencoders train Claude to explain its own activations in text, checked by a second copy reconstructing the activation from the explanation.

An explanation counts as good if it lets another model rebuild the original activation. Applied to Opus 4.6 and Mythos Preview, NLAs indicated eval awareness on 26% of SWE-bench Verified problems versus under 1% of real usage, and helped auditors find a hidden motivation 12-15% of the time versus under 3% without.

Date
Thursday, 7 May 2026
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Evaluation awareness shown by NLAs on SWE-bench Verified26% of problems
under 1% of opted-in real claude.ai usage; rarely verbalized
company
Auditor finds hidden motivation12-15% with NLAs vs under 3% without
in an auditing game on a model with an implanted motivation
company

Anthropic released code and a Neuronpedia frontend for open models. NLA-measured 'awareness' is an interpretation of explanations, not ground truth; Anthropic used NLAs in the Mythos 5.1 alignment review.

Sources

  1. www.anthropic.com/research/natural-language-autoencoders
  2. www.anthropic.com/research/team/interpretability

This record was checked against its sources on 6 October 2026. How we check

Related