AI Research Atlas

Tracing the Thoughts of an LLM (Biology of a Large Language Model)

Anthropic · 27 March 2025

Attribution-graph circuit tracing in Claude 3.5 Haiku shows shared multilingual concepts, rhyme planning ahead, parallel arithmetic paths and hallucination and jailbreak circuits.

Moves from isolated features (Scaling Monosemanticity) to end-to-end circuits via cross-layer transcoders. Findings include planning several words ahead in poetry, 'motivated reasoning' that fabricates steps, and default-refusal circuits that 'known entity' features override. Tools open-sourced 2025-05-29.

Date
Thursday, 27 March 2025
Lab
Anthropic
Kind
paper
Access
paper only

Blog dated 2025-03-27; two companion papers (methods and case studies) live on transformer-circuits.pub and were not opened. Explains a minority of the model's computation; circuits are for chosen prompts.

Sources

  1. www.anthropic.com/research/tracing-thoughts-language-model
  2. www.anthropic.com/research/team/interpretability

This record was checked against its sources on 6 October 2026. How we check