AI

Microsoft Releases New AI ‘Code Of Conduct’ To Guide Against Dangerous Behaviour

Microsoft has released a new ‘code of conduct’ to guide artificial intelligence (AI) models away from dangerous behaviour. The technology giant took the move as measure to guide development and usage as the AI world shifts its focus to safety and alignment.

The document provides an alternative to recent call by Anthropic’s chief executive officer, Dario Amodei, for a slow pace in development, instead focusing on the values and red lines that guide model training within Microsoft AI. Still, the result is a comprehensive guide as to how Microsoft approaches AI safety and how those ideas are implemented in practice.

The Microsoft latest policy begins with the prediction that, in the next decade, superintelligent AI systems will surpass human performance in most tasks. “Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced. We must therefore be completely clear about why we are inventing these systems and how we intend to control them,” the code of conduct states.

The code of conduct also lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to implement those principles.

Under Microsoft’s system, each model has an overarching code of conduct that overrides the preferences of individual users or any specific tasks. That includes “absolute constraints” forbidding cyberattacks, nuclear weapons, or deepfake production. It also includes broader provisions against a general loss of human control.

“MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorised people or systems,” the document reads.

The release comes amid an unprecedented focus on AI safety, driven by a string of rogue-agent incidents, as well as the abrupt resignation of an Anthropic employee, who cited the growing risk that AI would cause human extinction.

Together with Anthropic, OpenAI, and xAI, Microsoft has broadly embraced a general approach of pacing the frontier, with particular support for embedded evaluators in AI labs. “We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal. We also welcome ideas like ’embedded evaluators’ and the broader efforts to develop the mechanisms to make this more than just talk,” Microsoft CEO, Satya Nadella said in an online post.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button