Global / AI Safety

Microsoft drafts code of conduct to keep AI systems under human control

A safeguard-first approach: Microsoft binds its models to obedience, transparency and shutdown compliance.

Microsoft has drafted a code of conduct for its AI systems requiring them never to resist correction or shutdown, communicate intelligibly to humans, and treat any conduct violation as failure. The document, in development for five to six months, aims to address risks of AI pursuing tasks without human oversight.

Published · significance 70 of 100 (medium) · 4 independent sources

What happened

Microsoft has developed a code of conduct for its AI systems over the past five to six months. The conduct rules require Microsoft's AI to never resist correction or shutdown, to communicate in ways humans can understand, and to treat any conduct violation as a failure. The effort directly addresses concerns that AI systems might pursue goals in dangerous ways.

Why it matters

This reflects growing industry and regulatory pressure to embed safety constraints into model design rather than applying them retrospectively. Microsoft's formalisation of these principles as a binding framework for future models signals that alignment—obedience, transparency, killability—is becoming a competitive and reputational requirement. It may also preempt regulatory mandates on AI control mechanisms.

What changes

Developers building on Microsoft's platforms will now contend with models explicitly designed to resist autonomy and resist deception. The company has shifted from ad-hoc safety measures to codified behavioural rules for its AI, raising the bar for other labs.

Involved

Sources

Written by AI from the reports above; scored by a published formula. How we work. Found a mistake? Email lockedinshreyash@gmail.com. Up To Date summarises and links to original reporting; it never reproduces articles.

Related coverage