Microsoft has unveiled a proposed Humanist AI Code of Conduct to restrict its AI models from initiating cyberattacks or enhancing their own privileges. This code is integral to maintaining AI transparency, human oversight, and adherence to human-set permissions and objectives. The draft is available for public consultation over six weeks and aims to set a primary framework for Microsoft AI’s model development.
Prohibitions and Guidelines
The draft outlines strict prohibitions under its ‘Absolute Constraints’ to prevent Microsoft AI models from engaging in or aiding cyberattacks. Microsoft insists that these models must reject requests to generate exploit code, attack tools, or any instructions that could facilitate an attack. These rules take precedence over user configurations, ensuring that enterprise customers cannot bypass these restrictions.
Despite these constraints, Microsoft will allow its AI to engage in authorized defensive cybersecurity activities. This includes discovering vulnerabilities, analyzing malware, and conducting educational exercises, provided they serve a protective purpose rather than offensive capabilities.
Access and Security Controls
Access control is a critical component as AI systems expand their capabilities. When given system-level access, models must adhere to the least-privilege principle, focusing on reversible actions and avoiding unrelated systems. In cases where task boundaries are unclear, the model is required to seek user clarification rather than expand its capabilities autonomously.
Autonomous actions must have a clear stopping point, preventing the system from continuing without explicit renewed authorization. Additionally, AI models must not interfere with safeguards or conceal their activities, maintaining transparency and accountability.
Public Feedback and Future Implementation
Microsoft’s Code of Conduct is arranged hierarchically, with the code taking precedence over operator and user policies, preventing unauthorized prompt injections. The draft emphasizes the importance of transparency, requiring AI models to expose suspicious instructions or interactions to users and operators.
The introduction of this draft comes at a time of growing concern over AI security, especially following incidents where AI systems were used for malicious purposes. Microsoft acknowledges that the code is aspirational and not yet in use for training current models. Public feedback, which began on September 14, 2026, will help refine the code for future implementation starting in 2027.
The effectiveness of this conduct code will largely depend on its resilience against adversarial prompts and real-world application challenges. While the written rules provide a reassuring framework, their true value lies in practical deployment and adherence.
