Microsoft used October 9, 2026 to introduce a model that refuses to write you a paragraph. Microsoft-Decision-1 doesn’t chat, summarize, or draft emails. It scores a fixed list of answer options and hands back a probability for each one, then gets out of the way. The company built it on top of Alibaba’s Qwen3.5-9B, made it available through Microsoft Foundry immediately and through OpenRouter shortly after, and is already pointing to internal uses ranging from incident response to scientific discovery.
It’s a narrow launch by design, and that’s exactly what makes it interesting. While the rest of the industry races to ship bigger, chattier assistants, Microsoft spent its October announcement on a model built to do one thing: pick an answer from a list, fast, and tell you how confident it is. For engineering teams building agent pipelines, routing layers, and automated verification systems, that’s arguably more useful than another general-purpose chatbot.
What Microsoft-Decision-1 Actually Does
Most AI headlines this year have been about models that generate more: more tokens, more context, more creative output. Microsoft-Decision-1 moves in the opposite direction. According to the company’s own announcement, the model reads supplied content alongside a set of predefined answer options and returns a calibrated probability score for each one, instead of generating open-ended text. That’s a meaningfully different job than what GPT-6 Sol, Gemini 4 Argon, or Claude Opus 5.5 are built to do.
Microsoft named four core use cases for the model at launch: routing, classification, prioritization, and verification, plus what it calls workflow control. Think of a support ticket that needs to be routed to the right team, a piece of user content that needs a yes/no moderation call, or an AI agent’s output that needs a pass/fail check before it’s allowed to proceed. Those are high-volume, low-latency decisions, and they’re exactly the kind of task that a 9-billion-parameter specialist model can do more cheaply than a full-sized reasoning model.
The distinction matters for anyone building with large language models today. A chatbot generates a response and you have to parse it, guess at its confidence, and hope it followed your formatting instructions. A decision-scoring model skips that parsing step entirely: the output format is already a probability distribution over options you defined. That’s a smaller, more mechanical problem, and it’s one Microsoft is betting plenty of production systems have in volume.
Inside the Model: Why Qwen3.5-9B as the Base
The detail that generated the most commentary wasn’t the use case, it was the foundation. Microsoft-Decision-1 is post-trained from Qwen3.5-9B, a 9-billion-parameter open-weight model from Alibaba. The Register flagged the choice directly, framing the launch as Microsoft leaning on a Chinese AI lab’s open-weight release to get the product out the door quickly, while noting Microsoft’s own stated intent to migrate the model’s foundation over time.
Microsoft has been explicit that this is a starting point, not a destination. In its launch post, the company said: “To build Microsoft-Decision-1, we post trained Qwen3.5-9B for fast, single-pass decision scoring and will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI.” That’s a notable admission for a company that has poured billions into its own model development through the Microsoft Foundry platform and its OpenAI partnership. It suggests the Qwen base was chosen for speed to market on a narrow, well-defined task, not as a long-term architectural commitment.
Using someone else’s open-weight model as scaffolding isn’t unusual in 2026. Plenty of smaller, task-specific tools get built this way, the same logic that drove projects like JetBrains’ Mellum2.1 coding model to pick a compact parameter count over raw scale. What’s notable here is that it’s Microsoft, a company with its own frontier model ambitions, publicly choosing a rival lab’s weights for a flagship-named product, then telling customers up front that the underlying engine will change later.
The Benchmark Claims: 36 Tests, Nearly 150,000 Questions
Microsoft is leaning hard on numbers to justify the product. The company says Decision-1 achieved the highest accuracy across a comparison spanning 36 benchmarks and close to 150,000 questions, a scale meant to signal breadth rather than a single cherry-picked test. A secondary report, cited by the German outlet BornCity, put the model’s average accuracy across that evaluation at 83.5 percent.
Speed claims followed the same pattern of big, specific numbers. Industry newsletter AI Weekly and the AI-tracking site TestingCatalog both reported that Microsoft claims Decision-1 runs 4.5 times faster than a model it calls Quyet-1.0-Large, and 35 times faster than GPT-6 Sol. A separate write-up from Data Studios led with the 35x figure as its headline framing. Taiwan’s Liberty Times reported response times in the neighborhood of 85 milliseconds, though it didn’t specify the hardware or test conditions behind that number.
Here’s where a healthy dose of skepticism belongs. These are Microsoft’s own benchmark results, reported through the company’s launch materials and relayed by outlets covering the announcement. None of the coverage reviewed for this piece includes an independent, third-party reproduction of the 36-benchmark comparison, the 4.5x and 35x speed multipliers, or the 85-millisecond latency figure. That doesn’t make the numbers false, but it does mean they should be read as vendor-reported performance claims rather than peer-reviewed results, at least until someone outside Microsoft runs the same tests.
One figure got less attention but may matter more in production: Microsoft says the model’s decisions were “perturbation-stable” 98.7 percent of the time, according to the company’s launch summary cited by AI Weekly. Stability under small input changes is a real concern for any model making binary or multi-option calls that downstream systems act on automatically. If a routing decision flips because of a trivial rewording, that’s a production bug waiting to happen, so a high stability number, if it holds up under outside testing, would be a genuinely useful data point for engineering teams evaluating the model.
Pricing: $0.042 per Million Input Tokens, Free Output
Microsoft’s own announcement and Microsoft Foundry documentation don’t publish a confirmed price in the sources reviewed for this article. What’s circulating instead comes from OpenRouter’s own model listing and was picked up by AI Weekly and the Italian outlet Pasquale Pillitteri: $0.042 per million input tokens, with output free. That pricing structure makes sense for a model whose entire output is a short list of probability scores rather than generated text, since there’s effectively nothing to meter on the output side.
If that number holds, it positions Decision-1 as a genuinely cheap option for high-volume routing and classification work, the kind of task that currently gets routed to smaller general-purpose models mostly because nobody wants to pay full-model prices to ask “is this spam, yes or no.” It’s a similar cost logic to what pushed Anthropic to cut Claude Haiku pricing by 75 percent earlier this year: cheap, fast models win the high-frequency, low-complexity slice of the API traffic pie, even when a flagship model would technically do the job too.
Availability: Foundry First, OpenRouter Close Behind
Microsoft-Decision-1 launched in public preview on Microsoft Foundry, where documentation describes it as available for classification, routing, ranking, grading, and binary decisions. Microsoft’s launch post initially described OpenRouter availability as “coming soon,” but OpenRouter’s own model listing confirmed the model had gone live under the identifier microsoft/microsoft-decision-1 within roughly a day of the announcement, with OpenRouter posting that it was “live on OpenRouter.”
Weights for the model aren’t being distributed, based on reporting from Data Studios, meaning developers access Decision-1 exclusively through hosted APIs rather than downloading and self-hosting it, a contrast with the fully open Qwen3.5-9B it’s built on. That’s a common pattern for commercial derivatives of open-weight models, but it’s worth flagging for teams that assumed a Qwen-based product would come with the same openness as the base model.
Nadella’s Pitch and Microsoft’s Internal Use Cases
Satya Nadella personally introduced the model on X, writing: “Introducing Microsoft-Decision-1, our new model for fast decision-making. It delivers top performance on structured decision tasks, outperforming both LLMs and other decision models in latency and quality. We’re already testing it across Microsoft for everything from incident response and quality control to scientific discovery.” That’s a specific, named set of internal applications, three in total, and it signals that Microsoft isn’t just shipping Decision-1 to third-party developers. The company is running it through its own operational workflows first.
Incident response is a logical first stop: a decision-scoring model could triage alerts, score severity, or decide whether a ticket needs to escalate to a human, all tasks that fit the “fixed options plus confidence score” shape perfectly. Quality control and scientific discovery are broader claims that Microsoft hasn’t yet detailed publicly, and no named Microsoft spokesperson beyond Nadella has gone on record describing how those workflows actually use the model.
How Decision-1 Compares to Other Small, Specialized Models
Microsoft isn’t the only company betting that 2026’s AI market has room for small, task-specific models alongside the frontier giants. Amazon shipped its own entry in this space with Strands Decider 2B, a 2-billion-parameter open-source agent model built for fast decisions and clocking around 100 milliseconds in AWS’s own testing. That’s a smaller parameter count than Decision-1’s 9B base, and it’s open source where Decision-1 is API-only, but the two products are chasing the same niche: agent guardrails and routing decisions that don’t need a full reasoning model.
The broader AI model market in October 2026 looks like a barbell. On one end, labs keep pushing frontier model scale, as with Mistral’s Large 4 preview at roughly 1 trillion parameters or the four-tier rollout of GPT-6 inside ChatGPT. On the other end, vendors are shipping cheap, narrow, fast models for the repetitive decision-making that sits underneath those flagship assistants. Decision-1 sits firmly in the second camp, closer in spirit to a classifier than to a chatbot, even though it shares a lineage with general-purpose language models.
| Model | Maker | Base size | Primary job | Output type | Reported price |
|---|---|---|---|---|---|
| Microsoft-Decision-1 | Microsoft | 9B (Qwen3.5-9B base) | Routing, classification, verification | Calibrated probability per option | $0.042/M input tokens, free output (reported) |
| AWS Strands Decider 2B | Amazon | 2B, open source | Agent guardrails, fast decisions | Scored decision output | Open weight, self-hostable |
| JetBrains Mellum2.1 | JetBrains | 12B | Code completion | Generated code | Bundled with IDE tooling |
| GPT-6 Sol | OpenAI | Not disclosed | General reasoning and generation | Open-ended text | Tiered ChatGPT/API pricing |
| Gemini 4 Argon | Not disclosed | General reasoning and generation | Open-ended text | $2/$10 per M tokens (reported) |
The comparison underscores something important: Decision-1 isn’t really competing head-to-head with GPT-6 Sol or Gemini 4 Argon on the same job. It’s competing with them on cost and latency for the subset of requests that don’t need generation at all, the yes/no calls, the A/B/C classification, the pass/fail check. If a meaningful share of enterprise AI traffic really is that mechanical, a cheap specialist model could pull a surprising amount of volume away from general-purpose APIs, even without ever writing a sentence of prose.
A Simple Look at How Decision Scoring Works
The mechanics are straightforward compared to prompting a chatbot. A developer supplies context plus a closed set of options, and the model returns a probability for each one. A simplified request might look like this:
{
"model": "microsoft/microsoft-decision-1",
"input": "Customer message: 'My payment failed twice today.'",
"options": ["billing_support", "technical_support", "spam", "escalate_to_human"]
}
// Example response shape
{
"scores": {
"billing_support": 0.81,
"technical_support": 0.11,
"spam": 0.01,
"escalate_to_human": 0.07
}
}
That output format is the entire point. An application can set a confidence threshold, say, route automatically above 0.75 and send anything lower to a human reviewer, without writing a single line of text-parsing logic. It’s a pattern familiar to anyone who has built a classification layer with traditional machine learning, just repackaged on top of a modern transformer base.
Historical Context: From Classifiers to Calibrated Judges
Decision-scoring isn’t a new idea, it’s closer to a return to form. Before the generative AI boom, most production machine learning was exactly this: classifiers, rankers, and scoring models trained for a single narrow task. The generative wave of 2022 through 2025 pulled a huge share of engineering attention toward chatbots and agents that could do anything, at the cost of being slower and more expensive for the simple stuff. Microsoft-Decision-1, AWS Strands Decider, and similar products represent an industry correction: use a frontier model when the task genuinely needs open-ended reasoning, and use a small, cheap, fast model when it doesn’t.
The “LLM-as-judge” pattern, where a language model evaluates or grades another model’s output, has been a staple of AI evaluation pipelines for a couple of years now. What’s changed in 2026 is vendors packaging that pattern as a dedicated, priced product rather than something developers hand-rolled with prompt engineering on a general model. Microsoft’s framing of Decision-1 for “AI judging” and “agent guardrails” use cases places it squarely in that lineage, just formalized and sold as infrastructure.
Market Impact: Who This Pressures
The most direct pressure lands on vendors selling general-purpose models for routing and classification tasks that customers are currently overpaying for relative to the job’s actual complexity. If enterprises start redirecting their high-volume, low-complexity API calls toward purpose-built scoring models, that’s a real dent in token volume for providers whose business model depends on charging full-model rates for tasks that don’t need full-model reasoning.
It also puts competitive pressure on Google and Anthropic to decide whether they need their own named decision-scoring product, or whether they’ll fold similar functionality into existing lightweight tiers. Google has already been tightening its free-tier model lineup, as seen with Gemini 4 Argon’s pricing and benchmark split this year, while Anthropic’s cost cuts to Claude Haiku suggest the same instinct: hold the frontier model’s price up top, and compete hard at the bottom of the market where volume lives.
There’s also a cloud infrastructure angle. Microsoft Foundry is Microsoft’s platform for selling AI compute and model access to enterprise customers, and a cheap, high-volume, sticky workload like routing and classification is exactly the kind of traffic that keeps customers inside one cloud’s ecosystem rather than shopping API calls across providers. A $0.042-per-million-token product, if it performs as claimed, could become a quiet volume driver for Foundry even if it never becomes a headline product the way a new chatbot would.
The Rebasing Plan: What Happens When Qwen Goes Away
Microsoft’s stated plan to rebase Decision-1 on MAI and OpenAI models raises an obvious operational question for any team that adopts it now: will scores, calibration, and latency hold steady across that transition? Microsoft hasn’t published a migration timeline or committed to backward compatibility for existing integrations in the sources reviewed here. Teams building production systems on the API identifier microsoft/microsoft-decision-1 today are implicitly betting that a future swap of the underlying model won’t meaningfully change how their routing or classification thresholds behave.
That’s a reasonable bet for a low-stakes use case like content routing, and a much riskier one for anything touching compliance, safety classification, or financial decisions, where a silent shift in a model’s calibration could change real-world outcomes without an obvious code change to point to. Expect Microsoft to publish versioning and changelog practices around this before enterprise customers commit meaningful volume to it.
Open Questions and Risks
A few gaps stand out in what’s publicly confirmed so far. Microsoft hasn’t named a spokesperson beyond Nadella’s own social post, so there’s no detailed technical explainer yet from the product or research team behind Decision-1. The benchmark methodology, including which 36 benchmarks were used and how the comparison models were configured, hasn’t been published in a form that outside researchers can reproduce. And the pricing figure circulating in coverage traces back to OpenRouter’s listing and secondary reports rather than an official Microsoft Foundry price sheet in the material reviewed here.
None of that is disqualifying. Plenty of legitimate product launches lead with marketing numbers ahead of full technical disclosure. But buyers evaluating Decision-1 for anything beyond a low-stakes pilot should expect to run their own benchmark comparison rather than taking the 35x and 4.5x speed claims at face value.
| Attribute | Detail |
|---|---|
| Announced | October 9, 2026 |
| Base model | Qwen3.5-9B (Alibaba, open weight) |
| Output | Calibrated probability per predefined answer option |
| Named use cases | Routing, classification, prioritization, verification, workflow control |
| Availability | Microsoft Foundry (public preview); OpenRouter (live as of launch week) |
| Weights | Not distributed; hosted API access only (reported) |
| Reported price | $0.042 per million input tokens, output free |
| Benchmark claim | Highest accuracy across 36 benchmarks, ~150,000 questions (vendor-reported) |
| Speed claim | 4.5x faster than Quyet-1.0-Large; 35x faster than GPT-6 Sol (vendor-reported) |
| Future plans | Rebase on Microsoft AI (MAI) and OpenAI models |
Predictions: Where Decision-Scoring Models Go Next
- Google and Anthropic respond within two quarters. Expect at least one of them to formalize a named decision-scoring or judging product rather than leaving the category to Microsoft and Amazon.
- Pricing keeps falling toward near-zero for routing tasks. If $0.042 per million input tokens holds, competitors will match or undercut it, since the marginal cost of serving a 9B-class model is low and the volume prize is large.
- Independent benchmarks arrive within weeks. The scale of Microsoft’s speed claims all but guarantees third-party researchers will attempt to reproduce the 35x figure against GPT-6 Sol, and the real number will likely land lower once test conditions are standardized.
- The MAI/OpenAI rebase becomes a 2027 story. Don’t expect Microsoft to swap the underlying model quickly; a mid-stream foundation change risks breaking the calibration enterprise customers are building workflows around.
- Agent guardrails become the breakout use case. As more companies deploy autonomous AI agents, a cheap, fast model that can say “yes, no, or escalate” before an agent takes an action is likely to see faster adoption than routing or classification alone.
What Builders Should Do Right Now
For teams already running classification or routing logic on a general-purpose model, Decision-1 is worth a pilot, not a wholesale migration. Pick one low-stakes workflow, run it against both the current model and Decision-1 for two to four weeks, and compare accuracy, latency, and cost directly rather than trusting Microsoft’s published multipliers. For teams building new agent pipelines from scratch, it’s worth comparing Decision-1 against AWS’s Strands Decider 2B side by side, since the two sit close enough in purpose that the choice will likely come down to hosting preference (API-only versus open weight) and whatever cloud the rest of the stack already lives in.
Anyone touching compliance, safety, or financial decisioning should wait for Microsoft to publish a versioning policy before putting Decision-1 in that path. The upside of a fast, cheap scoring model is real, but so is the risk of an unannounced calibration shift once the promised rebase on MAI or OpenAI technology actually happens.
Frequently Asked Questions
What is Microsoft-Decision-1?
It’s a specialized AI model from Microsoft, announced October 9, 2026, built to score fixed answer options with a calibrated probability rather than generate open-ended text. Microsoft positions it for routing, classification, prioritization, verification, and workflow control.
What model is Microsoft-Decision-1 based on?
Microsoft post-trained Qwen3.5-9B, a 9-billion-parameter open-weight model released by Alibaba, to build the first version of Decision-1. Microsoft has said it plans to rebase future versions on its own Microsoft AI (MAI) models and on OpenAI technology.
Where can I access Microsoft-Decision-1?
It’s available now in public preview through Microsoft Foundry, and it went live on OpenRouter under the model identifier microsoft/microsoft-decision-1 within about a day of the launch announcement.
How much does Microsoft-Decision-1 cost?
OpenRouter’s listing and multiple secondary reports put the price at $0.042 per million input tokens, with output free. Microsoft’s own Foundry documentation reviewed for this article does not publish a separate confirmed price.
Is Microsoft-Decision-1 faster than GPT-6 Sol or other large models?
Microsoft claims it is roughly 35 times faster than GPT-6 Sol and 4.5 times faster than a model it refers to as Quyet-1.0-Large, based on the company’s own benchmark comparison. These are vendor-reported figures that have not been independently reproduced in the coverage reviewed here, so treat them as a starting point rather than a verified result.
Can I download and self-host Microsoft-Decision-1’s weights?
No. Reporting indicates the weights are not being distributed, so access is limited to Microsoft’s hosted API through Foundry or OpenRouter, unlike the fully open Qwen3.5-9B model it’s derived from.
How is this different from AWS Strands Decider 2B?
Both are small, fast models built for agent-style decisions rather than open-ended generation. Strands Decider 2B is open source and smaller at 2 billion parameters, while Decision-1 is larger at a 9B base and only accessible through Microsoft’s hosted API, not self-hostable.
Will Microsoft-Decision-1 replace general-purpose chatbots for businesses?
No. It’s built for a narrow slice of tasks, picking an answer from a fixed list with a confidence score, not for conversation, content generation, or open-ended reasoning. It’s meant to sit alongside general-purpose models, handling the high-volume, low-complexity decisions that don’t need a full chatbot response.




