← News

Microsoft's official Microsoft-Decision-1 release graphic: the model name in black over a numbered grid of dots on a lime background

10 Oct 2026

Microsoft

Microsoft's new AI model doesn't write anything. It just decides, and it's built on Alibaba's Qwen

Microsoft introduced Microsoft-Decision-1 late on Friday, Oct. 9, 2026, a "decision-scoring" model for routing, classification, prioritization, verification and agent control. Instead of generating text, it scores a fixed set of answers and returns a calibrated probability for each. It is available now in Microsoft Foundry and through OpenRouter, and was built by post-training Alibaba's open-weight Qwen3.5-9B.

Most of what an AI agent does all day isn't writing essays, it's tiny multiple-choice calls: is this a bug or a how-to, approve or escalate, next button or back button. Paying a frontier model to think out loud for each one is like hiring a novelist to sort the mail. Microsoft is betting that a small, cheap model that just returns "82% yes" becomes the plumbing every agent runs through, and it's pricing it like plumbing, with output literally free. The tell is the base: Microsoft, OpenAI's biggest partner, built its new model on Alibaba's open-weight Qwen and says it will swap in its own MAI and OpenAI models later. Coming less than two weeks after OpenAI's Decisions API, and the same week TypeSafe raised $870 million for Jev, "decision models" just went from niche to a category every big platform wants to own. As always, the benchmarks are Microsoft's own until someone else checks them.

Microsoft published “Introducing Microsoft-Decision-1, our model for fast decision-making” on its Command Line blog on 9 October 2026, U.S. time. The author is Achint Srivastava, vice president of software engineering in Microsoft’s Office of the CTO. Those lines are Microsoft’s.

Microsoft calls decision models “an important new category in AI.” Unlike large language models that generate text, they return structured outputs software can act on. Given a fixed set of options — yes or no, multiple choice or ratings, plus rubric-based grading of AI responses and agent actions — Microsoft-Decision-1 returns a calibrated probability for each option through a structured API call. Calibrated, here, means a higher score is supposed to mean the model is actually right more often. An API is a door a program can call. Those lines are Microsoft’s.

Microsoft says it built the model by post-training Qwen3.5-9B for fast, single-pass decision scoring. Post-training means it started from Alibaba’s already-trained open-weight model and trained it further for this narrower job. Microsoft says it will “soon rebase it on other models, including Microsoft AI (MAI) and OpenAI.” Those lines are Microsoft’s.

It is available now in Microsoft Foundry and through OpenRouter. Foundry is Microsoft’s place for a developer to call a model. Pricing: $0.042 per million input tokens; output tokens are free. A token is a small chunk of text the model reads. Those lines are Microsoft’s and OpenRouter’s.

Microsoft’s own claims: highest accuracy in its 36-benchmark comparison spanning nearly 150,000 questions kept blind from training; the fastest model it measured, 2.5 times quicker than runner-up H2O-Lightning-4B v1.1 and, at median (P50) latency, about 35 times quicker than GPT-6 Sol. Latency is the wait. Median, or P50, is the middle result. Microsoft says the post was later updated to add benchmarks for TypeSafe’s Jev. Those comparisons are Microsoft’s. They are not an outside lab’s rerun.

On robustness, across eight kinds of rewording and reordering of the same request, it changed its decision 1.3% of the time on average, with zero flips when options were paraphrased, reversed, or shuffled. On safety, Microsoft tested 5,250 requests across 11 benchmarks covering harmful content, jailbreaks, and prompt injection. A jailbreak is an attempt to talk a model into ignoring its safety rules. Prompt injection is hidden or hostile text meant to hijack the instructions. Those figures are Microsoft’s.

Internal uses Microsoft cites: Xbox Research sorted more than 10,000 pieces of player feedback from surveys, Steam, and X, finding it competitive in quality with GPT-6 Sol while over 14 times faster and 200 times cheaper; the Copilot team found it competitive with GPT5.6 Luna for grading chat and agent responses and 100 times faster; it is also used in on-call incident response and in Microsoft Discovery’s experiment replanning. Those lines are Microsoft’s account of its own teams.

Suggested uses include agent controls (continue, stop, retry, or hand off), model routing, data labeling, AI judging, content filtering, code scanning, and safety screening. Those uses are Microsoft’s. They are not a customer list.

The picture is Microsoft’s official Microsoft-Decision-1 release graphic: the model name in black over a numbered grid of dots on a lime background. It is the company’s launch image. The frame does not print a calendar date.

In plain terms, Microsoft said on Friday that Microsoft-Decision-1 does not write text. It scores a fixed list of answers and sends back a probability for each, at $0.042 per million input tokens with free output. It is in Microsoft Foundry and on OpenRouter now. Microsoft built it by further training Alibaba’s Qwen3.5-9B, and says it will later switch the base to its own MAI models and to OpenAI’s. The speed and accuracy numbers are Microsoft’s own tests, including a later add of TypeSafe’s Jev.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources