THE SIGNAL IN ONE SENTENCE

A machine that refuses to write an essay can still make a very consequential mistake. Microsoft released Microsoft-Decision-1 on October 9. It is a decision-scoring model built for a narrower job than a general chatbot. Give it text or JSON, ask one or more bounded questions and define the allowed answers. The model returns a typed result and probabilities rather than free-form prose or a written rationale. That sounds modest. Modesty is useful in software. A support system does not need a sonnet about a broken password. It needs to choose account access, billing, bug report or something else. An incident router does not need a paragraph about urgency. It needs a priority band and enough uncertainty to know when a person should look. A model selector does not need to narrate its feelings. It needs to choose among the components the application actually has. Microsoft-Decision-1 gives developers three primitives. A noul question returns a probability for a yes-or-no condition. A choice question selects from named options and returns a probability for each. A score question chooses a value on an ordered scale and also returns probabilities across the levels. Noul is Microsoft's term, not a typo. The unusual name matters less than the boundary it creates: the application decides the question, the evidence and the available moves before the model answers. That is cleaner than asking a large language model to generate a paragraph, parsing the paragraph and hoping the requested label appears in the correct corner. It can also be faster and cheaper because the system is scoring choices in a single pass instead of generating a chain of text. The product is available through Microsoft Foundry and OpenRouter. Microsoft's Foundry catalog describes it as a text-classification model with a 32,768-token context window. The Foundry documentation supports GlobalStandard deployment and DataZoneStandard in selected regions. Microsoft says DataZoneStandard keeps inference processing within the selected data zone, while GlobalStandard may use capacity in any supported Azure region. Under the hood, Microsoft says it post-trained Qwen3.5-9B for this specific style of decision scoring. The company says later versions may be rebased on models from Microsoft AI and OpenAI. The launch claims are punchy. Microsoft reports that Decision-1 had the highest accuracy in its comparison across 36 public and private benchmarks containing nearly 150,000 questions that were kept blind from training. It says the model ran 2.5 times faster than H2O-Lightning-4B v1.1 and about 35 times faster than GPT-6 Sol in its tests. The company also perturbed the same requests in eight ways, including changing formatting, option order and wording. It reports an average decision-flip rate of 1.3 percent, with no flips when option descriptions were paraphrased or when options were reversed or shuffled. Those are useful results. They are also vendor-run results. The reviewed launch page does not publish the full benchmark inputs, deployment configuration, threshold policy, uncertainty intervals or a task-by-task error ledger needed to reproduce the broad comparison. It says the benchmarks were hidden from training, but readers still have to trust Microsoft's construction and scoring of the test. That does not make the results false. It tells us what they are: a strong invitation to evaluate, not a transferable warranty. Microsoft's own documentation is refreshingly direct about the limits. Scores can change when questions or answer options are phrased or ordered differently. Calibration is strongest on familiar task types. The model can rely on outdated knowledge. It does not explain its reasoning. As a safety filter, it may miss subtle harmful material or flag harmless content. It may also carry bias from its base model and training data. The missing rationale is not automatically a flaw. Generated explanations can be persuasive fiction. A model can choose for one reason and produce a neat after-the-fact story for another. For many routing jobs, a concise probability and a preserved input are easier to audit than a paragraph that sounds certain. But no rationale changes what the surrounding system must record. If the model sends a ticket to fraud review, the audit trail should preserve the exact state, instructions, question, answer options, option order, model version, deployment, complete probability vector, threshold, selected action and final human or operational outcome. Without that evidence, the organization cannot tell whether the error came from the model, the wording, the threshold, the source data or the policy that turned a score into an action. Confidence is the most seductive part of the interface. Microsoft explains calibration plainly: if the model gives representative cases a 90 percent probability, roughly nine out of ten should be correct. That relationship is more useful than a raw ranking score because software can use it to decide when to act, ask for another evaluation or send the case to a person. The important word is representative. A model can be well calibrated on Microsoft's test set and poorly calibrated on an insurer's claims, a hospital's messages, a multilingual support queue or a new kind of prompt injection. The score may still contain two decimal places. The decimal places do not know the population changed. Teams need a local reliability curve. Group resolved cases by predicted confidence and compare the stated probability with the observed outcome. Do it again by language, product, customer group, input length, season, source system and consequence. Watch whether a 90 percent bucket really behaves like 90 percent, and whether one subgroup absorbs most of the mistakes. Then price the errors. A false positive that sends a routine ticket to senior support wastes time. A false negative that lets a malicious tool call reach production can expose data or cause damage. A ranking error in a shopping list is not the same thing as a ranking error in a medical queue. One global threshold cannot encode those differences. Microsoft recommends setting thresholds according to the cost of false positives and false negatives. It also recommends including an abstention option such as cannot tell, using clear neutral wording, randomizing option order during testing and keeping meaningful human review for consequential decisions involving credit, employment, housing, healthcare or law. That advice should be treated as part of the product, not a disclaimer beneath it. A forced menu can make an undefined case look decided. If the only options are safe and unsafe, the model must choose one even when the evidence is corrupted or outside the task. Add insufficient evidence, conflicting evidence or send to review when those are honest outcomes. The abstention path is not wasted capacity. It is the part of the interface that admits reality is larger than the schema. The same principle applies to agent controls. Microsoft suggests using Decision-1 to choose whether an agent should continue, stop, retry or hand off. That is a promising use because the model can be cheaper and more predictable than a second generative model. It is not a permission system by itself. A score can recommend that an agent continue. A deterministic policy should still decide whether the agent has authority to send the email, move the money, alter the database or call the production tool. The decision model belongs inside the control plane, not in place of it. Model routing is another sensible application. A bounded model can inspect a request and choose a faster model, a stronger model, a specialist tool or a human queue. But routing creates its own feedback loop. If the router sends difficult cases away, the data used to measure its success may contain mostly easy cases. If downstream models change, yesterday's optimal route may become tomorrow's expensive habit. Record the counterfactual when practical. Periodically send a small, privacy-safe sample through alternative routes and compare quality, latency and cost. Otherwise, the router can keep confirming the world it created. Microsoft reports several internal tests. Xbox Research used Decision-1 to sort more than 10,000 pieces of open-ended feedback into researcher-defined themes. Microsoft says it was competitive with GPT-6 Sol on quality while running more than 14 times faster and costing 200 times less. The company also reports tests in Copilot quality control, incident-response knowledge retrieval and adaptive planning for Microsoft Discovery. Again, these are company accounts, not independent production studies. They do show where the design is likely to earn its keep: high-volume, repeated decisions with a known option set and outcomes that can eventually be checked. The listed price reinforces that shape. Microsoft says input costs $0.042 per million tokens and output tokens are free. A cheap score can be attractive at enormous volume. Cheap mistakes also scale. The rollout should begin in shadow mode. Let the model score live cases without controlling them. Compare its choices and probabilities with resolved outcomes. Freeze the question wording long enough to measure it. Then deliberately perturb wording, order and formatting. Test missing fields, stale facts, adversarial instructions and cases from outside the normal distribution. Publish a decision card for every use case. It should name the allowed action, the evidence supplied, the outcome label, the abstention path, the threshold, the false-positive cost, the false-negative cost, the protected groups checked, the review owner, the appeal route and the condition that shuts automation off. Revalidate after any material change to the model, prompt, option set, source system or population. Calibration is not a certificate attached to the model forever. It is a measured relationship among a model, a task, a threshold and a moment in time. The plain signal is not that Microsoft made a smaller chatbot. It made a probability-shaped component for software that needs to choose. That is a welcome constraint. The harder constraint belongs to the team using it: no precise score gets permission to act until the local evidence says the number deserves the authority.

01

WHAT ACTUALLY CHANGED

Microsoft released Microsoft-Decision-1 on October 9, 2026

The model scores bounded decisions over text or JSON instead of generating free-form prose or a written rationale

It supports yes-or-no noul questions, fixed choices and ordered scores with probabilities

Microsoft says it post-trained Qwen3.5-9B for fast, single-pass decision scoring

The model is available through Microsoft Foundry and OpenRouter

Microsoft Foundry supports GlobalStandard and DataZoneStandard deployment in selected regions

Microsoft reports the top accuracy in its 36-benchmark comparison of nearly 150,000 questions

Microsoft reports an average 1.3 percent decision-flip rate across eight request perturbations

Input pricing is listed at $0.042 per million tokens with no output-token charge

02

WHY THIS MATTERS

A bounded schema removes free-form output parsing and makes decisions easier for software to consume

A probability can support action, abstention or review only when it is calibrated on representative local outcomes

Precise-looking confidence can conceal distribution shift, biased data, poor wording or the wrong option set

The absence of a written rationale increases the importance of preserving every input, option, threshold and outcome

False positives and false negatives carry different costs, so one threshold should not control unrelated decisions

An explicit abstention option prevents the model from being forced to choose when evidence is missing or outside scope

Decision scoring can guide an agent but cannot replace deterministic permissions for consequential external actions

Low price and latency can make an untested policy fail at much larger volume

Vendor benchmark and internal-use claims need independent, task-specific validation before production authority

FIG. 367How a probability earns permission to act
1Define one bounded question and the allowed outcomes→
2Add abstention for missing, conflicting or unfamiliar evidence→
3Collect representative labeled cases and preserve subgroup context→
4Run in shadow mode and record every probability and outcome→
5Measure calibration and price false positives and false negatives→
6Set the action, review and stop thresholds→
7Monitor drift and revalidate after every material change
The model supplies a score. Local outcomes, error costs and review policy decide what that score may control.

03

WHERE IT COULD HELP

  • Route support tickets among a fixed set of queues while sending uncertainty to a person
  • Prioritize incidents with thresholds tied to the cost of delayed or unnecessary escalation
  • Choose a model, tool or agent from an approved list without parsing a generated explanation
  • Grade AI responses against an ordered rubric while preserving the full probability vector
  • Run a proposed agent action through a bounded continue, stop, retry or review recommendation
  • Classify content for review without treating the model as the sole safety control
  • Launch in shadow mode and compare scores with resolved outcomes before enabling automation
  • Measure calibration separately by language, subgroup, source system and consequence
  • Randomize option order and perturb wording to expose brittle configurations
  • Publish an audit card with thresholds, error costs, review owners, appeals and shutdown conditions

KEEP A HAND ON THE WHEEL

The accuracy, latency, perturbation, safety, cost and internal-use results cited in the launch are Microsoft-reported and were not independently replicated in the reviewed materials. Microsoft says calibration is strongest on familiar task types, scores can move with wording or option order, the model may rely on outdated knowledge and it provides no rationale. A probability is not proof of correctness or fairness for a new population. Do not use the model as the sole basis for consequential decisions about people. Validate on representative data, include abstention, set thresholds from error costs, monitor disparities, disclose AI involvement and keep meaningful human review. DataZoneStandard constrains inference processing to a data zone, not necessarily a single region, and should be verified against the workload's residency requirements.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on October 11, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US