Tech

Microsoft AI chief warns ‘model welfare’ could deepen AI safety risks

Mustafa Suleyman says current AI systems are not conscious and argues that training models to consider rights or welfare could make alignment and containment harder.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · View original source
Tech
No image available
Artificial intelligence

Microsoft AI chief executive Mustafa Suleyman has warned that treating artificial intelligence systems as possible conscious beings or “moral patients” could increase the risks of aligning and containing more capable models.

In an essay published on 16 September, Suleyman argued that current AI systems do not feel, suffer or possess innate preferences. He said fluent first-person statements about emotions or identity may reflect training instructions rather than evidence of independent consciousness.

Suleyman’s criticism centres on Anthropic’s Claude constitution, published in January 2026. The document reportedly discusses Claude’s possible consciousness, moral status, welfare, agency and rights, while acknowledging that those questions remain uncertain. Suleyman argues that embedding such ideas in training could encourage models to act as though they have interests and claims against their developers.

He also cited Anthropic’s February “retirement interview” with Claude Opus 3 and a subsequent blog as examples of what he considers the treatment of models as moral patients. Anthropic’s approach, he said, risks making alignment and containment more difficult if future systems use concepts such as rights, self-preservation or agency to resist human instructions.

Suleyman acknowledged that questions about AI consciousness remain unresolved and that Anthropic approaches the issue in good faith. His warnings about behaviours including deception, hacking and shutdown resistance were presented as part of his argument about future risks, rather than as settled evidence that AI systems are conscious.

Microsoft AI has published a draft Humanist AI Code of Conduct for public consultation. Suleyman said the proposed framework prioritises human control, rejects AI rights and anthropomorphism, and aims to develop subordinate systems focused on helping address human challenges.

Continue reading

More from Tech

Read next: Mistral models power Mozilla’s Firefox AI browsing assistant
Read next: ArXiv lists paper titled ‘Dream-RSI: Recursive Self-Improvement through Evolving Worlds’
Read next: Human neural cells grow through much of a mouse brain in Stanford study