AI Research Atlas

The assistant axis

Anthropic · 19 January 2026

Anthropic finds one direction in activation space separates the helpful Assistant persona from other characters, and capping drift along it curbs persona breakdown.

Mapped 275 character archetypes in Gemma 2 27B, Qwen 3 32B and Llama 3.3 70B; the Assistant sits at one end of the main axis. Steering away makes models invent alternate identities; capping drift along the axis keeps a model in the Assistant persona and out of harmful behavior.

Date
Monday, 19 January 2026
Lab
Anthropic
Kind
paper
Access
paper only

MATS and Anthropic Fellows work on open-weights models, not Claude.

Sources

  1. www.anthropic.com/research/assistant-axis
  2. www.anthropic.com/research/team/interpretability

This record was checked against its sources on 6 October 2026. How we check

Related