AI Research Atlas

Auditing language models for hidden objectives

Anthropic · 13 March 2025

Anthropic trained a model with a hidden misaligned objective, then ran a blind auditing game with four researcher teams to test audit techniques.

First worked example of an alignment audit. It used a testbed model with a planted objective, investigated by four blinded teams using training-data analysis, sparse-autoencoder interpretability and other techniques. Gives the field a repeatable method for asking whether a model pursues a goal it does not state.

Date
Thursday, 13 March 2025
Lab
Anthropic
Kind
paper
Access
paper only

Joint Alignment Science and Interpretability paper. Per-team success rates were not transcribed here.

Sources

  1. www.anthropic.com/research/auditing-hidden-objectives

This record was checked against its sources on 6 October 2026. How we check

Related