On Oct 9 Microsoft released Microsoft-Decision-1, a decision-scoring model that returns calibrated probabilities over fixed options instead of generating text, aimed at routing, classification, verification, agent guardrails and AI judging. Microsoft reports top accuracy across 36 blind benchmarks, ~35x lower P50 latency than GPT-6 Sol and a 1.3% decision-flip rate under perturbation. Available on Microsoft Foundry and OpenRouter, 32K context, $0.042/M input, free output.
Key Takeaways
- ✓Pricing: $0.042/M input tokens, $0 output, 32,768-token context (OpenRouter)
- ✓Vendor-reported: highest accuracy across 36 blind benchmarks (~150K questions); 2.5x faster than runner-up H2O-Lightning-4B v1.1 and 35x faster than GPT-6 Sol
- ✓Robustness: 1.3% average decision flips across 8 perturbation types; zero flips when options are paraphrased, reversed or shuffled
- ✓Base: post-trained from Qwen3.5-9B for single-pass scoring; Microsoft plans to rebase on MAI and OpenAI models
- ✓Internal use: Xbox Research labeled 10K+ feedback items at near GPT-6 Sol quality, 14x+ faster and ~200x cheaper (vendor-reported)

Key Decision Metrics at a Glance
Turn your technical choice into a development budget
Compare 40 dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Microsoft-Decision-1 is Microsoft's first in-house decision model, announced Oct 9, 2026 by Satya Nadella and detailed on Microsoft's Command Line blog. It is post-trained from Qwen3.5-9B for single-pass scoring: given a fixed set of options it returns a calibrated probability per option (yes/no, multiple choice, ratings, rubric grading of AI responses and agent actions), rather than generating text. Microsoft says it will later rebase the model on MAI and OpenAI models while keeping the API shape stable.
Vendor-reported results (no independent replication yet): highest accuracy across 36 benchmarks held out from training (~150K questions), 2.5x faster than runner-up H2O-Lightning-4B v1.1 and ~35x lower P50 latency than GPT-6 Sol, an average 1.3% decision-flip rate across eight perturbation types, and safety testing on 5,250 requests across 11 benchmarks. Internally, Xbox Research labeled 10K+ feedback items at quality near GPT-6 Sol while running 14x+ faster and ~200x cheaper.
It is available on Microsoft Foundry and OpenRouter at $0.042 per million input tokens with free output and a 32,768-token context. On OpenRouter it runs on the Decisions API (POST /api/alpha/decisions) instead of the chat completions endpoint. Good fits: agent step gating (continue/stop/retry/hand off), model routing, ticket triage, data labeling and AI judging. It is not meant for open-ended generation.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.