Microsoft has introduced a draft for a ‘Humanist AI Code of Conduct’ aimed at regulating the cybersecurity capabilities of its AI models. This new framework specifies safety measures that restrict the use of offensive cyber tools, the autonomy of AI agents, and sets up a review process for cybersecurity applications.
Regulating Cyberattack Capabilities
As per the code of conduct, Microsoft’s AI models will not generate exploit code or provide assistance for cyberattacks, regardless of how requests are formulated. These rules are part of what Microsoft terms ‘Absolute Constraints,’ which neither deploying companies nor end users can modify. The aim is to distinguish between understanding cyber threats for defensive purposes and enabling attacks.
Despite these restrictions, the AI models can still support legitimate defensive activities. This includes tasks like discovering vulnerabilities, analyzing malware, and developing proof-of-concept exploits within a legal framework.
Managing External Influences
Microsoft’s code also addresses instructions originating from external content. The document emphasizes a ‘Chain of Command’ for governing model behavior, starting with the code itself, followed by policies from operators, and then user preferences.
External tools, files, or communications do not possess inherent authority unless explicitly delegated, while maintaining adherence to the Absolute Constraints. Any suspicious input should be flagged for the attention of users and operators.
Ensuring Controlled Autonomy
The guidelines also tackle the potential risks of AI systems acting with real permissions. MAI Models are expected to operate only within the scope defined by users or operators, without independently expanding objectives.
When granted system-level access, the models are instructed to operate with minimal privilege, avoid interacting with unrelated systems, and prioritize reversible actions. They are prohibited from escalating their access without authorization.
Delegation rules ensure that any sub-agents follow the same constraints and permissions as the original model and adhere to any stop or shutdown requests.
Exceptions and Future Developments
The code acknowledges exceptions in domains like defensive cybersecurity, national security, and certain scientific research areas, where enhanced capabilities might be necessary. Microsoft plans to conduct an extensive review through authorized channels to assess safety and legal implications in these cases.
Currently, Microsoft’s AI models have not been trained on this new code. The company has opened a six-week public consultation period to refine the guidelines for its 2027 model development. Feedback from AI, legal, and ethics experts, as well as public focus groups, influenced the draft.
An appendix with example scenarios highlights the behaviors the code intends to encourage, demonstrating both aligned and misaligned responses in various situations.
