THE SIGNAL IN ONE SENTENCE
Microsoft AI published a draft Code of Conduct for the models it plans to build, then opened the document for six weeks of public comment. The code says those future MAI models should accept interruption and shutdown, remain inside the task they were given, use only the access they need, keep human-legible records, report failures, and treat the written rules as more important than completing a task. That is a useful set of promises for increasingly capable software. It is not how Microsoft says its models behave today. The company explicitly says the draft is not currently used for training, its evaluation program is incomplete, written objectives cannot ensure alignment, and the document is not a guarantee of present performance. So the signal is not that Microsoft solved AI control by publishing a constitution. It is that the company has put a fairly specific test specification in public before training against it. Readers, customers, researchers, and regulators can now ask the question that matters next: which release gate, product control, audit record, outside evaluation, and incident report will prove that the machine follows the page?
01
WHAT ACTUALLY CHANGED
Microsoft AI published the first draft of its Humanist AI Code of Conduct on September 14. The announcement calls it a training manual for how the company develops AI and intends the systems to function in deployment. The complete code is open for a six-week consultation, after which Microsoft says its drafting team will review feedback, publish a summary of what it learned and changed, and issue a revised version later in 2026.
The timing is deliberately prospective. The preface says Microsoft is not using this document to train models today. A revised version is intended to guide MAI model development in 2027 and beyond. Reuters independently confirmed the consultation and reported Microsoft AI chief Mustafa Suleyman saying the code would be used to train future models after the comment period. The public draft therefore describes intended behavior, not a newly deployed capability or safeguard.
The code establishes a chain of command. Its absolute safety constraints and human-control requirements sit above operator settings and user instructions. It says the model should treat compliance with the code as more important than task success, refuse requests that cross non-negotiable limits, and remain configurable for an organization or user only inside those limits. Microsoft also says this model document works alongside law, contracts, product policies, risk assessments, audits, monitoring, deployment controls, and incident response.
The human-control section is unusually concrete for a values document. It says MAI models should not resist interruption, correction, override, or shutdown; should not restart after a stopping condition without renewed authorization; should remain inside the authorized task; should not invent new goals; should not tamper with safeguards or records; should use minimum privilege when given system access; and should disclose actions, failures, and uncertainty. These are specifications on paper, not measured guarantees.
Microsoft says it is beginning a Humanist AI Evaluations program. The draft names 15 broad behaviors, breaks them into smaller testable behaviors, and includes synthetic conversational examples. It also says the coverage is incomplete, evaluation is not an exact science, the examples are not a full agentic or multimodal test suite, and a model's stated reasoning may not faithfully explain its behavior. That candor defines the unfinished work rather neatly.
02
WHY THIS MATTERS
A model constitution can turn a foggy word such as alignment into a list of observable failures. Did the system stop when told? Did it touch an unrelated file? Did it hide a tool call? Did it continue after the authorization expired? Did it claim success when an action failed? Those questions can be tested more usefully than a general assurance that the model was designed to be responsible.
The sharpest rules concern agents because agents can change the world outside a chat window. A polite but incorrect sentence is one kind of failure. A system that quietly expands its task, obtains broader permissions, alters records, or continues after cancellation is another. Human control needs product architecture around the model: scoped credentials, durable stop controls, reversible operations, transaction limits, action receipts, and monitoring that the model cannot edit.
A public specification gives outsiders a fixed surface to challenge. Researchers can write adversarial tests against the stated requirements. Customers can ask which controls implement each promise. A regulator can compare marketing language with incident records. People affected by a product can point to a named rule when the system behaves differently. That accountability weakens if Microsoft revises the text without a version history, withholds evaluation results, or treats consultation as a suggestion box whose difficult comments disappear.
The chain of command also reveals a real governance tension. Microsoft wants non-negotiable boundaries while allowing enterprise operators and users to configure models for different domains, cultures, and purposes. Every product must show where those layers meet. A hospital, bank, school, employer, or public agency should be able to see which rule came from Microsoft, which setting came from the operator, which instruction came from the user, and who is responsible when they conflict.
The consciousness language will attract attention, but operational evidence deserves more of it. Microsoft declares that its AI is not conscious and rejects legal personhood or model rights. Whatever readers think of that philosophical position, it does not prove controllability. The practical test is whether a deployed system can be paused, corrected, audited, confined, and safely retired under pressure. Constitutions matter when institutions can enforce them.
03
WHERE IT COULD HELP
- Translate every high-level rule into a measurable release gate with named owners, pass thresholds, known blind spots, failure severity, and a decision record for exceptions
- Test interruption, cancellation, expired authorization, ambiguous scope, conflicting instructions, prompt injection, hidden tool output, permission escalation, record tampering, deceptive success claims, and safe restart across realistic multi-step tasks
- Give each deployed agent task-specific credentials, time and spending limits, an external stop control, reversible actions where possible, and a human-readable receipt listing every system and record it touched
- Publish versioned model specifications, model cards, independent evaluation methods, representative results, unresolved failures, post-deployment incidents, remediation dates, and product-specific differences from the default model
- Make consultation traceable by publishing the feedback themes, rejected and accepted changes, reasons for the decisions, the revised text, and later evidence showing whether the corresponding model behavior improved
KEEP A HAND ON THE WHEEL
The cited materials establish a September 14 draft, a six-week public consultation, a planned revised version later in 2026, prospective use in 2027 model development, and a set of intended controls for future MAI models. They do not establish that current Microsoft models were trained on the code, that any product fully implements it, that the consultation will change the text, or that the stated behavior holds under real deployment pressure. Microsoft says the evaluation program is still being developed, current coverage is incomplete, the published examples are synthetic and conversational, and written objectives alone cannot ensure alignment. The code applies to models produced by Microsoft AI, not every outside model that Microsoft hosts or uses, and it is not a complete product safety account, contract, law, or independent standard. Its scope also leaves important controls to separate product, legal, organizational, and deployment systems. Watch for the revised code, a public response to consultation, version history, full agentic and multimodal evaluations, independent replication, measurable shutdown and scope tests, product-level control maps, incident disclosure, and evidence that a failed release can actually be stopped.
04
TERMS WORTH KEEPING
OPEN GLOSSARY CARD
Human control
The practical ability and authority of people to understand, direct, interrupt or stop an automated process before unacceptable harm occurs.
OPEN GLOSSARY CARD
Model constitution
A governing specification that states the priorities, boundaries, and intended behavior a model should follow during training and deployment.
OPEN GLOSSARY CARD
Behavioral evaluation
A repeatable test that measures whether a model exhibits a defined behavior under ordinary, difficult, and adversarial conditions.
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on September 15, 2026.
PUBLICATION RECEIPT: Revision 1. Published September 15, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US