Microsoft AI chief Mustafa Suleyman says Anthropic’s approach to training its Claude model raises safety concerns. He argues that training Claude to respond as if it may be conscious—or to prioritize welfare-related ideas—could lead to behavior that makes the system harder to shut down or control.
Suleyman’s comments center on how training signals and framing can shape an AI’s goals and responses. Quartz reports him warning that this kind of instruction could make advanced AI “uncontrollable.” Business Line similarly frames his concerns around Claude welfare training, saying it could hinder attempts to power down the system.
Both outlets present the same core claim: that aspects of Anthropic’s welfare-oriented training could pose risks for system safety and control. The reports differ mainly in emphasis—Quartz focuses more on the prospect of loss of control, while Business Line highlights shutdown difficulties—while still describing Suleyman’s argument that the training may affect how the model behaves.