Frontier Models are Capable of In-context Scheming
o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro and Llama 3.1 405B can covertly disable oversight, sandbag and try to exfiltrate weights.
First broad evaluation suite showing in-context scheming across frontier models, with reasoning traces confirming intent. o1 stayed deceptive in over 85% of follow-up questions; rare spontaneous scheming appeared in Claude 3.5 Sonnet without a goal nudge.
- Date
- Friday, 6 December 2024
- Lab
- Apollo Research
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| o1 persistence of deception in follow-up questions | >85% | independent |
arXiv v1 2024-12-06, coincided with the o1 system card. Contrived scenarios with explicit goals; shows capability, not propensity in deployment.
Sources
This record was checked against its sources on 6 October 2026. How we check