Auditing language models for hidden objectives
Anthropic trained a model with a hidden misaligned objective, then ran a blind auditing game with four researcher teams to test audit techniques.
First worked example of an alignment audit. It used a testbed model with a planted objective, investigated by four blinded teams using training-data analysis, sparse-autoencoder interpretability and other techniques. Gives the field a repeatable method for asking whether a model pursues a goal it does not state.
- Date
- Thursday, 13 March 2025
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Joint Alignment Science and Interpretability paper. Per-team success rates were not transcribed here.
Sources
This record was checked against its sources on 6 October 2026. How we check