AI Research Atlas

Emergent Introspective Awareness in LLMs

Anthropic · 29 October 2025

Concept-injection experiments show Claude Opus 4/4.1 can sometimes notice and name a concept planted in its activations, about 20% of the time.

Researchers injected activation patterns for specific concepts and asked whether the model noticed. Claude Opus 4.1 detected injections about 20% of the time; Opus 4 and 4.1 were best of the Claude models tested. The capability is unreliable, needs a narrow injection strength, and models often confabulate.

Date
Wednesday, 29 October 2025
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Injected-concept detection rateabout 20%
Claude Opus 4.1
company

Authors state this does not show human-like introspection; the mechanisms are unknown.

Sources

  1. www.anthropic.com/research/introspection
  2. transformer-circuits.pub/

This record was checked against its sources on 6 October 2026. How we check

Related