LLMs show functional introspective awareness of internal states
TL;DR. Research indicates that large language models can introspect on their internal states, noticing and identifying injected concepts within their activations. - Models demonstrated an ability to recall prior internal representations, distinguishing them from raw text inputs and artificial prefills. - Claude Opus 4 and 4.1 generally exhibited the highest degree of introspective awareness in tested scenarios. - This functional capacity remains unreliable and context-dependent but may improve with future model advancements.
- Researchers investigated LLMs' capacity for introspection by injecting known concepts into their activations.
- Models were able to notice and accurately identify the presence of these injected concepts.
- Some models could recall prior intentions to differentiate their own outputs from prefills.
- Claude Opus 4 and 4.1 showed the strongest introspective awareness among tested models.
- The findings suggest current LLMs have some functional introspective awareness, though it is currently unstable.