AI Research Atlas

Tracing the thoughts of an LLM (circuit tracing)

Anthropic · 27 March 2025

Attribution-graph circuit tracing on Claude 3.5 Haiku shows shared multilingual concepts, rhymes planned ahead, and parallel approximate-and-exact arithmetic.

Two companion papers (Circuit Tracing and On the Biology of a Large Language Model) trace computation inside Claude 3.5 Haiku. They find that concepts are shared across English, French and Chinese, the model picks rhyme words before writing a line, arithmetic runs on parallel paths, and explanations are sometimes fabricated. Captures only a fraction of computation.

Date
Thursday, 27 March 2025
Lab
Anthropic
Kind
paper
Access
paper only

Authors note analyzing one prompt takes hours of human effort and observed mechanisms may include artifacts.

Sources

  1. www.anthropic.com/research/tracing-thoughts-language-model
  2. transformer-circuits.pub/

This record was checked against its sources on 6 October 2026. How we check

Related