MICROSOFT AI has published a draft “Humanist AI Code of Conduct” for its MAI Models, setting limits on offensive cyber capabilities, autonomous agents and the handling of instructions from external content. Under its “Absolute Constraints”, models must not produce working exploit code, attack tools, targeting or planning methods, intrusion procedures, evasion techniques, operational guidance or other assistance that could enable or improve a cyberattack.
Microsoft distinguishes this from lawful defensive work, which may include vulnerability discovery, malware analysis, proof-of-concept exploit development and testing, and educational explanations of attacks.
The draft says authority over a model comes through a defined “Chain of Command”: the code itself, policies set by deploying organisations, and users’ preferences. Tool outputs, files, webpages and messages from other AI systems have no authority unless explicitly delegated within those limits; suspicious content should be flagged where relevant. Models should also make their reasoning and actions visible to human overseers.
Agents with system access are expected to use least privilege, avoid unrelated data and systems, favour reversible actions, flag broad or lasting changes, and never escalate their own access. Delegated sub-agents must retain at least the same permissions and constraints and comply with stop-work requests.
Microsoft acknowledges that defensive cybersecurity, public safety, national security and dual-use research may require capabilities outside standard settings. It plans enhanced review through authorised Microsoft channels for a “small number of use cases”. The current MAI Models have not been trained on the document. Microsoft is seeking public feedback for six weeks before revising it later this year to guide 2027 model development.