TL;DR
On October 9, Microsoft launched Microsoft-Decision-1, a model that scores a fixed set of choices instead of generating text — built for routing, classification, AI judging, and agent controls. It is post-trained from Qwen3.5-9B, available in Microsoft Foundry and on OpenRouter, priced at $0.042 per million input tokens with output tokens free. Microsoft claims 35x the speed of GPT-6 Sol on decision tasks, and its own Xbox Research team labeled more than 10,000 feedback items at GPT-6 Sol quality for 200x less money.

What a decision model actually does
Give it a situation and a fixed set of options — yes or no, a multiple-choice list, a 1-to-5 rating — and it returns a calibrated probability for each option in a single pass. No prose, nothing to parse. It also does rubric-based grading: hand it a definition of good and it scores an AI response or an agent action against it.
To build it, Microsoft post-trained Qwen3.5-9B for single-pass decision scoring, and says it will rebase the system on other models later, including its own MAI models and OpenAI models. The model is live now in Microsoft Foundry and on OpenRouter, with documentation on Microsoft Learn. Announcing it, Microsoft AI VP Achint Srivastava framed decision models as a new category — purpose-built for structured outputs software can act on immediately.
The numbers — and the honest caveat
Microsoft's own numbers: highest accuracy across a 36-benchmark comparison covering nearly 150,000 questions kept blind from training; 2.5x quicker than the runner-up H2O-Lightning-4B v1.1 and 35x quicker than GPT-6 Sol at P50 latency. On robustness, the decision changed on 1.3% of cases when inputs were perturbed in eight ways, with zero flips when option descriptions were paraphrased or options were reversed or shuffled. Safety testing covered 5,250 requests across 11 benchmarks for harmful content, jailbreaks, and prompt injection.
The caveat is one word: vendor-run. Every benchmark above came from Microsoft. An independent test published the next day by Calibrated Agents ran Microsoft-Decision-1 against OpenAI, TypeSafe's Jev, and Cloudflare models on one specific job — checking AI answers against their sources — and concluded no single provider wins everywhere. Treat Microsoft's numbers as the marketing version of a real capability, not gospel.
Microsoft is already using it internally
Xbox Research used it to label more than 10,000 open-ended feedback items and reviews from surveys, Steam, and Twitter/X against a fixed set of researcher-defined themes — competitive with GPT-6 Sol on quality, over 14x faster, and 200x cheaper. The Copilot team found it competitive with GPT-5.6 Luna for measuring chat and agent response quality, at 100x the speed. On-call engineers used it to pull relevant knowledge during live incidents and found it better and faster than an LLM. In Microsoft Discovery, it scored experiments 46x more consistently than an LLM-based score, at three times the speed, and nearly 4x the speed on adaptive replanning.
Those are Microsoft's internal reports, not independent verification — but they show the shape of the work this model is built for: repetitive judgment calls where the options are known and the cost of the big model is the problem.
The practical play: cheap guardrails for your agents
This is the pattern that matters for anyone building with agents. Today, the usual way to check an agent's work is to ask a big LLM to judge it — expensive, slow, and itself prone to flakiness. A decision model flips that: it is the cheap, fast referee standing between the expensive model and the action.
Three places it fits. As a router: score the incoming request and pick the cheapest model that can actually do the job. As a gatekeeper: the agent proposes its next step, the decision model scores continue, stop, retry, or hand off — to a model, a tool, or a human. As a verifier: does the answer meet the rubric — accept, revise, or reject. And because the probability itself is part of the API, you can wire confidence directly into policy: act above 90%, kick to a human below 60%. That is a human-in-the-loop gate you can actually afford to run on every single step.