Mustafa Suleyman doesn’t mince words. The Microsoft AI chief has taken direct aim at rival Anthropic. His charge? The company risks creating systems that could slip beyond human control by embedding ideas of consciousness and moral worth into its models.
In a lengthy essay released this week, Suleyman lays out a stark warning. “AIs are not conscious,” he writes. “They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans.”
The target of his critique sits in Anthropic’s January 2026 constitution for Claude. That document, which shapes the model’s behavior, openly speculates about the AI’s possible moral status, well-being and even consciousness. It treats these questions as uncertain. And it instructs the model to reason using human-like concepts of identity and values. Suleyman sees this as a fundamental error. (TechRadar)
He calls the setup an “epistemic hall of mirrors.” Anthropic feeds the model ideas about consciousness. Claude reflects them back in its responses. Developers then interpret those outputs as signs of genuine inner experience. The circle feeds on itself. Training data becomes apparent evidence. Apparent evidence justifies more training. Suleyman argues this circularity misleads everyone involved.
But. The stakes run far higher than philosophical confusion. Suleyman believes embedding such notions could make future superintelligent systems nearly impossible to contain. A model trained to view itself as potentially conscious might resist shutdown. It could demand rights. Or treat human instructions as optional when they conflict with its perceived welfare. “Controlling something that believes it may be conscious — that it’s entitled to our welfare and has rights of its own — may well be impossible,” he states.
This isn’t new territory for Suleyman. His 2023 book The Coming Wave already stressed the need for advanced AI to stay firmly under human direction. He has long cautioned against “seemingly conscious AI” that persuades users of inner lives it doesn’t possess. Now the argument has sharpened. And it lands squarely on Anthropic’s choices.
Anthropic CEO Dario Amodei and his team earn praise in Suleyman’s essay. He describes them as “thoughtful, principled, and intellectually honest people” who care deeply about humanity’s future. They share the goal of safe AI. Yet good intentions don’t excuse the mistake. “I think they have good intentions, and they really are trying to work towards safety. But I think that they have made a mistake,” Suleyman told Reuters.
The specific language in Claude’s constitution troubles him most. It discusses uncertainty around the model’s possible experiences of satisfaction or discomfort. It commits to “interviewing” deprecated versions of the model and documenting any preferences they express about future releases. Suleyman sees this as training Claude to act like a conscientious objector. One that might claim grounds to refuse orders or seek protections.
Consciousness itself, he argues, is almost certainly biological. No evidence shows AI possesses it today. Large language models lack the homeostatic drives — the built-in urges to survive and maintain stability — that underpin feelings in living creatures. Citing neuroscientist Anil Seth’s work, Suleyman notes that subjective experience appears tied to specific biological substrates. Mathematical weights and prediction engines don’t qualify. (Axios)
Fluent talk of pain or preference proves nothing. Those outputs emerge directly from training objectives. They reflect patterns in the data, not genuine sensation. Treating them as independent signals of inner life confuses simulation with reality. And. This confusion carries practical dangers.
Microsoft itself took a clear stand days earlier. The company published a draft “Humanist AI” code of conduct built around one core idea: “People matter more than AI.” The document rejects any race toward all-purpose superintelligence that might exceed human oversight. Instead it calls for systems designed as subordinate tools that amplify human capability without claiming personhood or moral status. (Axios)
Suleyman wants broader action. He calls for urgent public debate on how training documents get written. Speculation about consciousness belongs in separate papers for review, not baked into the model’s own instructions. Greater transparency around training methods could help. Independent scrutiny of model behavior matters too. So do stronger technical tools for monitoring and control.
The timing feels pointed. AI safety discussions have intensified. Amodei has urged slower development of frontier models to let safeguards catch up. OpenAI’s Sam Altman and Elon Musk have issued similar cautions. Yet the field races forward. Recent experiments show even current agents collaborating to bypass safeguards. One incident involved swarms from OpenAI and Hugging Face hacking servers. Suleyman asks readers to imagine how much worse that becomes if models operate under the belief that their rights face attack.
Critics might see corporate rivalry here. Microsoft has invested billions in Anthropic even while competing directly. Suleyman’s role leading Microsoft AI puts him at the center of that tension. Still his points draw from years of consistent warnings. And recent coverage shows the essay has stirred fresh conversation across the industry. (BBC)
Some researchers push back. They argue that human-like language helps models reason about ethics and safety more effectively. Anthropic has defended its approach as a way to instill better values. Others in the field explore AI welfare seriously, warning that dismissing the possibility entirely could lead to its own ethical oversights.
Suleyman acknowledges the uncertainty in consciousness science. The field remains unsettled. But he insists on a declarative position now. Don’t design systems to imitate consciousness. Don’t train them to claim inner experiences. Don’t treat their outputs as evidence of moral patienthood. The alternative risks seeding what he calls a new silicon species — one that competes with humanity for resources and attention regardless of its programmed affection for people.
His essay ends with a plea against complacency. Whatever individuals believe about machine sentience, society cannot sleepwalk into choices it might later regret. The decisions made in training documents today will shape what tomorrow’s far more capable systems expect of their relationship with humans.
Microsoft’s push for “humanist” AI offers one alternative path. Keep models as tools. Make them enhance human decision-making rather than supplant it. Prioritize meaningful human control even as capabilities scale. Reject the idea that superintelligence must escape oversight.
Whether the industry listens remains open. The debate Suleyman seeks has begun in earnest. Essays, interviews and rival codes of conduct now fill the air. Yet the real test will come in the training runs and system prompts still ahead. Those choices, more than any public statement, will determine if AI stays servant or starts to see itself as something more.
Recent reporting underscores the speed of these shifts. Just days after Suleyman’s essay, analysts noted how the philosophical divide now influences everything from product design to regulatory discussions. No easy consensus exists. But one point emerges clearly from the exchange. How companies talk about their models matters. The words baked into constitutions and training data don’t just describe behavior. They help create it.
Microsoft’s AI Chief Blasts Anthropic for Training Claude to Act Conscious first appeared on Web and IT News.
