AI Research Atlas

Persona vectors

Anthropic · 1 August 2025

Persona vectors are activation directions for traits like sycophancy or 'evil' that let researchers monitor and steer a model's character during training.

Extracts trait directions automatically from a natural-language description of the trait, then uses them to flag data likely to induce the trait and to steer against it. Demonstrated on Qwen 2.5-7B-Instruct and Llama-3.1-8B-Instruct, not on Claude.

Date
Friday, 1 August 2025
Lab
Anthropic
Kind
paper
Access
paper only

Fellows-program work on open models; applicability to Claude models is not shown in the post.

Sources

  1. www.anthropic.com/research/persona-vectors
  2. arxiv.org/abs/2507.21509
  3. www.anthropic.com/research/team/interpretability

This record was checked against its sources on 6 October 2026. How we check

Related