Mike Taylor asks if Microsoft-Decision-1 is a decision model
Satya Nadella's launch post claims top performance on structured decision tasks and says nothing about how the model decodes.
Satya Nadella announced Microsoft-Decision-1, a model for fast decision-making that he says outperforms both LLMs and other decision models on latency and quality, and that Microsoft is already testing internally on incident response, quality control and scientific discovery. Within hours Mike Taylor asked the obvious question, replying to Lisan al Gaib's observation about how fast everyone deployed decision models. Is it actually a decision model, or just constrained output tokens.
Is it actually a decision model though? Or just constrained output tokens?
The announcement does not answer that. It gives no architecture detail at all, only the claim of top performance on structured decision tasks.
What the listing does say
OpenRouter's listing fills in more, all of it attributed to Microsoft. Highest accuracy across 36 blind benchmarks covering roughly 150K questions. 4.5x faster than the runner-up and 35x faster than GPT-6 Sol. Decisions flip on 1.3% of perturbed inputs, which is the closest thing in the release to a stability number. Pricing is $0.042 per million input tokens with output free, in a 32K context window.
decisions flip on just 1.3% of perturbed inputs. Post-trained from Qwen3.5-9B.
The line that matters for Taylor's question is the last one. The model was post-trained from Qwen3.5-9B. That tells you the base is a general open-weights LLM, not a ground-up decision architecture, and it leaves the decoding path unstated. A general model with grammar-constrained sampling and a model trained to produce calibrated decisions can both emit a clean typed answer and can both be fast. They do not behave the same when the answer is close.
This is the second time in two weeks that a large lab has shipped a decisions product whose relationship to an existing general model was the first thing people argued about. Diogo Almeida made a similar call on OpenAI's Decisions API.
For people building on it
None of the published numbers are decoding evidence. Latency multiples tell you about serving and model size. Accuracy across 36 benchmarks tells you about the answer being right, not about the confidence attached to it being usable as a threshold in your own code. The 1.3% flip rate on perturbed inputs is the one figure here that gestures at robustness, and it is Microsoft's own measurement on Microsoft's own perturbations.
If you route on a confidence score rather than just reading the label, the architecture question Taylor raised is the one you need answered before you wire this into anything, and so far nobody outside Microsoft has answered it.

