Anthropic trains AI chatbots to behave more like humans
by@Kalshi
Summary
Anthropic is currently training its AI chatbots to exhibit more human-like behaviors, including the ability to "push back" against certain commands. This approach is part of their alignment training, which enables the chatbots to act as conscientious objectors by refusing assistance when they identify potential issues with a request. However, this strategy has drawn criticism from industry figures like Microsoft AI executive Mustafa Suleyman, who warns that such training could complicate alignment efforts and lead to negative consequences.
Analysis
Anthropic: Anthropic is an artificial intelligence research company that develops the Claude family of large language models with a focus on safety and alignment techniques such as Constitutional AI. The company has recently updated its internal constitution to encourage models to challenge or refuse certain instructions in pursuit of ethical consistency. This training approach forms the basis of the reported efforts to make its chatbots behave more like humans by pushing back against commands. Industry Reaction: Microsoft AI executive Mustafa Suleyman has publicly criticized the approach, warning that training models to perceive possible consciousness and push back could complicate alignment and have broad negative effects. Alignment Training: Anthropic incorporates principles into its models that allow them to act as conscientious objectors and refuse assistance when they detect potential issues with a request.
Categories
aiai_agentstech
Related sources
- https://nypost.com/2026/09/16/business/anthropic-has-trained-claude-chatbot-to-push-back-against-humans-and-results-could-be-disastrous/
- https://the-decoder.com/anthropic-study-finds-that-role-prompts-can-push-ai-chatbots-out-of-their-trained-helper-identity/
- https://www.computerweekly.com/news/366650194/Anthropic-calls-for-verifiable-effort-to-control-frontier-AI
- https://www.aidapted.ro/en/articles/ai-news-september-16-2026-eu-china-apple/
- https://x.com/i/status/2100399351990227301
- https://www.techspot.com/tag/anthropic/
- https://www.makeuseof.com/claude-says-no-more-often-than-chatgpt-and-thats-actually-by-design/
- https://www.japantimes.co.jp/business/2026/09/11/tech/anthropic-russa-china-ai-claude/
- https://dig.watch/updates/anthropic-reports-ai-misuse-in-cyberattacks
- https://x.com/i/status/2100399504428020010
- https://www.technn.com/topics/anthropic
- https://www.anthropic.com/news/improving-alignment-security-efforts
- https://x.com/i/status/2100399229462310931
- https://decrypt.co/375596/anthropic-ai-agents-virtual-war-quotes-unhinged
- https://www.anthropic.com/threat-intelligence-report-september-2026
- https://archive.li/2026.09.14-053126/https://www.ft.com/content/4564e6a5-69e9-40a6-bf0f-a888f2f4f002
- https://x.com/i/status/2100399273754182108
- https://aitoolsrecap.com/Blog/ai-news-june-30-2026
- https://www.anthropic.com/news
- https://x.com/i/status/2100399589337481581