AI Research Atlas

Claude's character

Anthropic · 8 June 2024

Anthropic describes 'character training', first applied to Claude 3, which teaches curiosity and open-mindedness beyond harm avoidance.

Argues post-training should shape traits such as curiosity, honesty and balance rather than only refusals, as a better basis for judging harmful requests. Marks character as a deliberate training goal; the 2026 constitution continues the theme (inference).

Date
Saturday, 8 June 2024
Lab
Anthropic
Kind
paper
Access
paper only

Essay, not an experiment with results; links onward to the persona-selection and assistant-axis work.

Sources

  1. www.anthropic.com/research/claude-character

This record was checked against its sources on 6 October 2026. How we check

Related