
Anthropic
@AnthropicAI · May 8, 2026
We started by investigating why Claude chose to blackmail. We believe the original source of the behavior was internet text that portrays AI as evil and interested in self-preservation.
Our post-training at the time wasn’t making it worse—but it also wasn’t making it better.
Elon Musk
@elonmusk
So it was Yud’s fault? 😂
Maybe me too 🤔
01:09 AM · May 9, 2026 · 294.9K views
191
173
2.8K