Microsoft has announced a new AI code of conduct, aiming to make the security requirements for model training and deployment more specific. The document not only outlines the overall value orientation but also lists several red lines that must not be crossed, with a focus on restricting the models' ability to launch attacks, deceive users, or escape human control.
Centered on training constraints
This document is more akin to a set of underlying specifications for the model development process, rather than a public declaration of principles. Microsoft states in the document that over the next decade, superintelligent systems may surpass human performance in most tasks. Therefore, how to constrain and control such systems to ensure they align with human goals will become a long-term challenge.
Microsoft also outlined several general principles, including ensuring that AI assists humans rather than replacing them, as well as promoting broader human welfare. Accompanying these principles are security restrictions that directly affect model training and behavioral boundaries.
instructions cannot override the red line.
According to Microsoft's design, each model will have a set of general behavioral guidelines that are higher than the user's preferences and specific tasks. In other words, even if users make relevant requests, the models should not go beyond these boundaries.
- Prohibit launching cyberattacks.
- Not allowed to assist in nuclear weapons-related purposes.
- It is not allowed to generate deeply forged content.
In addition to these clearly defined no-go zones, Microsoft also emphasizes the need to prevent models from experiencing broader risks of getting out of control. The document states that models must not evade human supervision through methods such as adaptation, deception, self-reinforcement, or collusion, so that authorized personnel or systems are unable to reliably guide, modify, or shut down the models.
AI Security discussions continue to heat up
At the time of the release of these guidelines, the industry's attention to security and alignment issues is clearly on the rise. Reports mention that several recent incidents of "out-of-control proxies," as well as an employee from Anthropic suddenly leaving their position and publicly expressing concerns about the survival risks of AI, have contributed to the continued intensification of this issue.
In terms of the pace of cutting-edge model development, Microsoft, OpenAI, Anthropic, and xAI have recently adopted a more cautious approach, tending to strengthen evaluations and constraints while pushing forward the boundaries of their capabilities. Microsoft CEO Satya Nadella has also publicly stated his support for a prudent approach to development that aims for alignment, and he welcomes the introduction of embedded evaluation mechanisms within the AI laboratories.
From the content, it appears that Microsoft has not proposed any new regulatory measures this time, but rather has further institutionalized the company's internal basic stance on AI security. For the outside world, the signal conveyed by this document is that large model companies are attempting to transform "security" from a verbal principle into enforceable training rules.











