Persona vectors
Persona vectors are activation directions for traits like sycophancy or 'evil' that let researchers monitor and steer a model's character during training.
Extracts trait directions automatically from a natural-language description of the trait, then uses them to flag data likely to induce the trait and to steer against it. Demonstrated on Qwen 2.5-7B-Instruct and Llama-3.1-8B-Instruct, not on Claude.
- Date
- Friday, 1 August 2025
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Fellows-program work on open models; applicability to Claude models is not shown in the post.
Sources
- www.anthropic.com/research/persona-vectors
- arxiv.org/abs/2507.21509
- www.anthropic.com/research/team/interpretability
This record was checked against its sources on 6 October 2026. How we check