Tracing the Thoughts of an LLM (Biology of a Large Language Model)
Attribution-graph circuit tracing in Claude 3.5 Haiku shows shared multilingual concepts, rhyme planning ahead, parallel arithmetic paths and hallucination and jailbreak circuits.
Moves from isolated features (Scaling Monosemanticity) to end-to-end circuits via cross-layer transcoders. Findings include planning several words ahead in poetry, 'motivated reasoning' that fabricates steps, and default-refusal circuits that 'known entity' features override. Tools open-sourced 2025-05-29.
- Date
- Thursday, 27 March 2025
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Blog dated 2025-03-27; two companion papers (methods and case studies) live on transformer-circuits.pub and were not opened. Explains a minority of the model's computation; circuits are for chosen prompts.
Sources
- www.anthropic.com/research/tracing-thoughts-language-model
- www.anthropic.com/research/team/interpretability
This record was checked against its sources on 6 October 2026. How we check