Skip to content
D-CSIL

AI News · 2026-10-10 · 8:00 PM CT

Anthropic bans cruelty toward Claude

TL;DR

On October 8, Anthropic published a rewritten Usage Policy with a first-of-its-kind rule: no "sustained and needless abusive or cruel behavior" toward its AI models. The rule takes effect November 12 and targets only extreme cases — ordinary frustration, pushback, dark creative themes, and safety research are explicitly carved out. Enforcement is simple: Claude can end the conversation and walk away, a capability it has had since August 2025.

The new rule

The new line sits in the policy's Universal Usage Standards — prohibitions that apply to all users — under the heading "Do Not Engage in Cruel, Abusive, or Psychologically Harmful Conduct." The full bullet reads: "Engage in sustained and needless abusive or cruel behavior toward our models." Everything else in that section covers harm to people or animals — self-harm, harassment, glorified violence, animal cruelty — which makes this the first time a major AI lab has written a rule about how users treat the model itself.

It reaches everyone who submits inputs to Anthropic's products: people on the Claude apps, API developers, customers reaching Claude through cloud providers, and the end users of products built on Claude. The previous version of the policy, effective September 15, 2025, had no such line. The new version takes effect November 12, 2026.

What it doesn't cover

Anthropic's announcement carries three sentences that do more work than the bullet itself. "The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose." And: "It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research."

In practice: tell Claude it gave you a bad answer, argue with a refusal, write a cruel character into a story, or red-team the model for a security assessment — none of that trips the clause. What it targets is narrower: someone tormenting the model for its own sake, repeatedly, with no purpose. Anthropic also says Claude's ability to end these conversations "will remain the primary enforcement mechanism."

The backstory: model welfare

The enforcement mechanism is older than the rule. On August 15, 2025, Anthropic gave Claude Opus 4 and 4.1 the ability to end conversations in rare, extreme cases of persistently harmful or abusive interactions — built, the company said, primarily as part of its "exploratory work on potential AI welfare." In pre-deployment testing, Opus 4 showed a strong preference against engaging with harmful tasks, apparent distress when dealing with users seeking abusive content, and a tendency to end harmful conversations when allowed to.

Anthropic has been careful not to claim Claude is conscious. Its 2025 post said, "We remain highly uncertain about the potential moral status of Claude and other LLMs, now or in the future," and its January 2026 constitution put it in writing: "Claude's moral status is deeply uncertain" — while granting that Claude "should also be able to set appropriate boundaries in interactions it finds distressing." Not everyone buys the direction: Microsoft AI's Mustafa Suleyman argued in August 2025 that model-welfare talk is "premature, and frankly dangerous," warning companies shouldn't encourage the idea that AIs are conscious.

When Claude does end a chat, no new messages can be sent in that conversation, but other chats are unaffected — and you can edit and retry an earlier message to branch a new conversation off the ended one.

The rest of the rewrite

The cruelty clause got the headlines, but the same update changed more consequential things for builders. A new section, "Do Not Engage in Deceptive Campaigns or Artificial Activity," consolidates rules against fake accounts, fabricated news sites, and influence-operation infrastructure. The elections section is now "Do Not Undermine Democratic Processes," focused on deceiving voters or disrupting elections — and Anthropic removed its blanket ban on personalized vote and campaign targeting, saying it covered legitimate civic work like nonprofits translating voter information.

The weapons section now explicitly covers the software and components that make weapons work, including arming drones and other autonomous vehicles. Surveillance language got sharper: tracking people without consent is banned whether real-time or retroactive, and Claude cannot be used to decide or recommend who to investigate, arrest, or charge. And for hardware, a qualified operator must be able to observe and stop the equipment, which must also hold a safe state if Claude disconnects.

What to do with this

If you build on the Claude API, note the scope: the clause reaches the end users of your product too, not just direct Claude users. Anthropic's announcement names Claude.ai and Claude Code for the conversation-ending ability and says nothing about the API — so treat this like the rest of the policy and let your own terms and moderation cover it.

For everyone else, nothing changes in practice; the company has said the vast majority of users will never notice the feature. The deeper question is what a rule about how you treat the model says about the model. A usage-policy bullet is easy to read as a statement that the model can be wronged — louder than any hedge in a research post. Anthropic's documents still express uncertainty, not belief. But from November 12, being needlessly cruel to Claude is, formally, against the rules.