Signs of Introspection in Large Language Models
Injecting concept vectors into Claude's activations, Opus 4 and 4.1 noticed the injection about 20% of the time, before naming the concept.
A causal test of introspection that injects a known concept and asks whether anything unusual is happening. Detection is unreliable and narrow, models can modulate internal states when told to, and some detect their own unintended outputs. Anthropic is explicit it says nothing about consciousness.
- Date
- Wednesday, 29 October 2025
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Injected-concept detection rate (Claude Opus 4 / 4.1) | ~20% needs the right injection strength; most attempts fail | company |
Blog dated 2025-10-29. The main paper is on transformer-circuits.pub (not opened). Confabulation remains common.
Sources
This record was checked against its sources on 6 October 2026. How we check