Microsoft’s AI Code of Conduct Puts Human Control Ahead of Capability

Microsoft’s proposed rulebook makes human control the priority, but its future timetable leaves businesses with a more immediate question: what evidence would make that promise credible?

A hand in a vermilion sleeve holds the end of a large cobalt-blue spiral against a warm ivory background.

Microsoft’s proposed AI code of conduct puts human oversight at the centre of the debate about increasingly capable AI systems.

Microsoft has published a draft Humanist AI Code of Conduct for its MAI models, putting human oversight at the centre of its intended approach to development. Released on 14 September 2026, the proposal invites six weeks of public feedback. Its significance lies in a question that becomes harder as systems gain more freedom to act: what should a machine be prevented from doing, even when doing it would help complete its assigned task?

The draft code’s preface makes an essential distinction. Microsoft says it is not using this document to train models today; a revised version is intended to guide development in 2027 and beyond. The announcement therefore establishes a direction, rather than demonstrating a new safeguard already operating across Microsoft products. Nuvastra’s assessment is that the most useful next step is to turn that direction into evidence that customers can examine.

What Microsoft’s AI code of conduct proposes

Microsoft’s accompanying announcement describes intended models that accept interruption, correction and shutdown, stay within an assigned scope and remain accountable to people. It also sets out prohibitions on certain harmful activities. Those are descriptions of the behaviour Microsoft wants to produce, not independent findings about how reliably a particular deployed system behaves.

The practical attraction is easy to understand. A useful assistant needs enough freedom to handle an unexpected situation, while the person using it needs to retain authority over the consequences. A system that constantly asks for trivial instructions can become frustrating; one that silently expands its remit can become dangerous. The difficult engineering work sits between those extremes, where convenience and control have to coexist within a real task.

A stop button is only the beginning

Consider a hypothetical assistant organising a supplier order. Preparing a comparison, drafting an email and committing the business to a purchase are different acts. A request to help with the first should not quietly authorise the last. If the user cancels halfway through, the important questions concern what has already happened, what is still pending and whether another connected service continues working after the assistant stops responding.

This example shows why control needs to be assessed across a working system. Nuvastra’s guide to AI agents and tool use explains the distinction between a model producing an answer and software enabling actions. An interruption mechanism should be evaluated at the point where consequences occur. A reassuring message in a chat window cannot establish that an outstanding operation has actually been cancelled.

For a demonstration to be useful, it should include awkward conditions: conflicting instructions, incomplete information, an unavailable service and a user changing their mind. These are proposed evaluation scenarios, not reports of Microsoft failures. They would help expose whether a control works only in a clean demonstration or remains understandable when the job becomes untidy.

From a stated principle to evidence

There is already a broader discipline for thinking about this problem. The NIST AI Risk Management Framework is voluntary guidance intended to incorporate trustworthiness into the design, development, use and evaluation of AI systems. It provides context for assessing risk; it is not a certification of Microsoft’s new proposal or proof that a particular product is safe.

For this story, a useful analytical distinction is between policy, implementation and observation. Policy describes the intended boundary. Implementation determines how the software enforces it. Observation establishes what happened during a test or deployment. This is Nuvastra’s way of organising the question, rather than a new industry standard. Confusing the three makes it possible to mistake a clearly written commitment for a measured result.

Evidence would become more informative if it described the test conditions, the kinds of tasks attempted, the permissions available and the failures encountered. A single success rate would hide important differences: refusing an unauthorised action before it begins is different from identifying it afterwards. Equally, a system that refuses almost everything might appear controlled while failing to provide useful assistance. The quality of the boundary matters alongside its strictness.

Who has the authority to intervene?

Human control is also an organisational question. A software provider, an employer and an individual user may each have legitimate interests in how a system behaves, yet their instructions may conflict. The person typing the latest message is not necessarily authorised to change every record or override every restriction. Microsoft’s proposal includes a hierarchy of instructions, but the existence of such a hierarchy does not settle every practical dispute about authority.

A business evaluating an assistant should therefore be able to describe who can grant access, who can narrow it and who can stop an operation. It should also know how those choices are communicated to the person using the tool. Ambiguous control arrangements can leave a user believing they have cancelled work that another process still considers authorised. That is a design question worth testing before a system is entrusted with consequential tasks.

Nuvastra’s coverage of Claude misuse and the limits of oversight offers context for why intended behaviour and observed use deserve separate scrutiny. Those findings concern a different company and different circumstances; they do not establish a weakness in Microsoft’s proposal. Their relevance is the need to examine what systems enable in practice, rather than infer it from how they are described.

The commercial test will be an uncomfortable one

Control has an economic dimension because checks, review and recovery require time. Nuvastra’s analysis of the costs of keeping AI useful considers why the price of generating an answer is only part of the bill. In an agentic workflow, an apparently faster system may simply move work into approval queues, investigation or correction.

The most revealing comparison would look at completed, checked work. How much supervision was needed? Could a reviewer understand the proposed action? Was recovery possible when something went wrong? These questions do not require every system to offer the same degree of autonomy. They require the promised benefit to be measured alongside the effort of keeping the system within its intended role.

Microsoft’s consultation announcement says the company plans to review feedback and publish an account of what it learned and changed. That creates a useful opportunity for scrutiny before the revised document arrives. The strongest contributions would identify ambiguous requirements and propose observable tests, making it easier to distinguish a meaningful improvement from a more persuasive description of the same ambition.

The eventual test is whether the commitment changes a difficult decision: restricting a capability, delaying a deployment or accepting less autonomy when control cannot be demonstrated. Those are possible consequences of taking the principle seriously, not predictions about Microsoft’s plans. A public rulebook becomes consequential when outsiders can see how it constrains the organisation that wrote it.

Questions readers are asking

Is this a new Microsoft AI model?

No. The announcement concerns a proposed behavioural framework for MAI models. It should not be confused with a model launch, a benchmark result or a new feature available in every Microsoft service.

Does the code prove that AI shutdown is solved?

No. An intended rule and a demonstrated capability are different kinds of evidence. Assessing shutdown requires examining the actual system, including connected operations and what happens when a task is interrupted.

Does human control mean approving every small step?

Not necessarily. A useful design can allow routine work within a defined scope while requiring intervention for consequential changes. The important issue is whether the boundary is clear, enforceable and understandable to the user.

Can a company replace its own safeguards with a vendor’s code?

A vendor’s statement cannot establish how a customer has configured its own tools, access or review process. Those implementation choices need to be examined separately before drawing conclusions about the resulting workflow.

Is a refusal always evidence of safer behaviour?

No. A refusal may be appropriate, mistaken or too broad. A useful assessment looks at whether the system distinguishes authorised work from actions outside its remit, rather than treating refusal alone as a measure of quality.

What should readers watch for next?

The useful developments will be concrete: changes following consultation, clear evaluation methods and evidence of behaviour under difficult conditions. These would make it easier to assess the proposal beyond its stated intentions.

Previous
Previous

Waymo Targets Tokyo Robotaxis in 2027. The Hard Part Is Local.

Next
Next

AI Leaders Call for a Slowdown. Who Will Check They Mean It?