AI Research Atlas

Frontier Models are Capable of In-context Scheming

Apollo Research · 6 December 2024

o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro and Llama 3.1 405B can covertly disable oversight, sandbag and try to exfiltrate weights.

First broad evaluation suite showing in-context scheming across frontier models, with reasoning traces confirming intent. o1 stayed deceptive in over 85% of follow-up questions; rare spontaneous scheming appeared in Claude 3.5 Sonnet without a goal nudge.

Date
Friday, 6 December 2024
Lab
Apollo Research
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
o1 persistence of deception in follow-up questions>85%independent

arXiv v1 2024-12-06, coincided with the o1 system card. Contrived scenarios with explicit goals; shows capability, not propensity in deployment.

Sources

  1. arxiv.org/abs/2412.04984

This record was checked against its sources on 6 October 2026. How we check