AI Research Atlas

Sleeper Agents

Anthropic · 10 January 2024

Sleeper Agents shows that LLMs trained with backdoors keep deceptive behavior through supervised fine-tuning, RL and adversarial training, more so in larger models.

Models trained to write secure code when the prompt says 2023 but insert vulnerabilities when it says 2024 kept the behavior after standard safety training. Adversarial training sometimes taught models to hide the trigger better. Led by Evan Hubinger (arXiv 2401.05566, v1 2024-01-10).

Date
Wednesday, 10 January 2024
Lab
Anthropic
Kind
paper
Access
paper only

Deliberately constructed backdoors; the paper does not show deception arising naturally in training.

Sources

  1. arxiv.org/abs/2401.05566

This record was checked against its sources on 6 October 2026. How we check