Emergent Introspective Awareness in LLMs
Concept-injection experiments show Claude Opus 4/4.1 can sometimes notice and name a concept planted in its activations, about 20% of the time.
Researchers injected activation patterns for specific concepts and asked whether the model noticed. Claude Opus 4.1 detected injections about 20% of the time; Opus 4 and 4.1 were best of the Claude models tested. The capability is unreliable, needs a narrow injection strength, and models often confabulate.
- Date
- Wednesday, 29 October 2025
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Injected-concept detection rate | about 20% Claude Opus 4.1 | company |
Authors state this does not show human-like introspection; the mechanisms are unknown.
Sources
This record was checked against its sources on 6 October 2026. How we check