AI Research Atlas

Opus 4 and 4.1 can end abusive conversations

Anthropic · 15 August 2025

Claude Opus 4 and 4.1 gained the ability to end a rare subset of consumer chats, as part of Anthropic's exploratory model-welfare work.

For persistently harmful or abusive interactions, Claude can end the conversation in claude.ai. Anthropic frames it mainly as an AI-welfare precaution with relevance to alignment and safeguards. Follows the April 2025 'Exploring model welfare' program.

Date
Friday, 15 August 2025
Lab
Anthropic
Kind
feature
Access
app only

Anthropic says it remains highly uncertain about whether models have moral status.

Sources

  1. www.anthropic.com/research/end-subset-conversations
  2. www.anthropic.com/research/exploring-model-welfare

This record was checked against its sources on 6 October 2026. How we check

Related