Natural Language Autoencoders
Natural Language Autoencoders train Claude to explain its own activations in text, checked by a second copy reconstructing the activation from the explanation.
An explanation counts as good if it lets another model rebuild the original activation. Applied to Opus 4.6 and Mythos Preview, NLAs indicated eval awareness on 26% of SWE-bench Verified problems versus under 1% of real usage, and helped auditors find a hidden motivation 12-15% of the time versus under 3% without.
- Date
- Thursday, 7 May 2026
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Evaluation awareness shown by NLAs on SWE-bench Verified | 26% of problems under 1% of opted-in real claude.ai usage; rarely verbalized | company |
| Auditor finds hidden motivation | 12-15% with NLAs vs under 3% without in an auditing game on a model with an implanted motivation | company |
Anthropic released code and a Neuronpedia frontend for open models. NLA-measured 'awareness' is an interpretation of explanations, not ground truth; Anthropic used NLAs in the Mythos 5.1 alignment review.
Sources
- www.anthropic.com/research/natural-language-autoencoders
- www.anthropic.com/research/team/interpretability
This record was checked against its sources on 6 October 2026. How we check