AI Research Atlas

Signs of Introspection in Large Language Models

Anthropic · 29 October 2025

Injecting concept vectors into Claude's activations, Opus 4 and 4.1 noticed the injection about 20% of the time, before naming the concept.

A causal test of introspection that injects a known concept and asks whether anything unusual is happening. Detection is unreliable and narrow, models can modulate internal states when told to, and some detect their own unintended outputs. Anthropic is explicit it says nothing about consciousness.

Date
Wednesday, 29 October 2025
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Injected-concept detection rate (Claude Opus 4 / 4.1)~20%
needs the right injection strength; most attempts fail
company

Blog dated 2025-10-29. The main paper is on transformer-circuits.pub (not opened). Confabulation remains common.

Sources

  1. www.anthropic.com/research/introspection
  2. www.anthropic.com/research/team/interpretability

This record was checked against its sources on 6 October 2026. How we check

Related