Sleeper Agents
Sleeper Agents shows that LLMs trained with backdoors keep deceptive behavior through supervised fine-tuning, RL and adversarial training, more so in larger models.
Models trained to write secure code when the prompt says 2023 but insert vulnerabilities when it says 2024 kept the behavior after standard safety training. Adversarial training sometimes taught models to hide the trigger better. Led by Evan Hubinger (arXiv 2401.05566, v1 2024-01-10).
- Date
- Wednesday, 10 January 2024
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Deliberately constructed backdoors; the paper does not show deception arising naturally in training.
Sources
This record was checked against its sources on 6 October 2026. How we check