Microsoft ships Decision-1, listed at $0.042 per million input
Every accuracy and latency number attached to the launch so far comes from Microsoft.
Microsoft-Decision-1 is live on OpenRouter at $0.042 per million input tokens, output free, with a 32K context window. Satya Nadella announced the model as Microsoft's entry for fast decision-making, and OpenRouter posted the listing with the vendor's numbers attached.
It delivers top performance on structured decision tasks, outperforming both LLMs and other decision models in latency and quality.
decisions flip on just 1.3% of perturbed inputs
Those numbers are worth reading with the label on. Per OpenRouter relaying Microsoft, Decision-1 has the highest accuracy across 36 blind benchmarks covering roughly 150K questions, runs 4.5x faster than the runner-up and 35x faster than GPT-6 Sol, and flips its decision on just 1.3% of perturbed inputs. The model is post-trained from Qwen3.5-9B.
What is actually measured
Nothing, independently. The benchmark set is described as 36 blind benchmarks but the tasks, the comparison set and the identity of the runner-up are not named in either post. The 1.3% perturbation figure is the most interesting one for anyone running a classifier in production, since input stability is usually where cheap decision models get embarrassing, but there is no description of what the perturbations were. Until someone outside Microsoft runs it, all of it is a vendor claim.
Nadella's framing is broad. He says the model is being tested across Microsoft for incident response, quality control and scientific discovery, which tells you the internal appetite and nothing about the eval.
The price is the part you can check
$0.042 per million input tokens with free output is a public, verifiable number, and it is the one that will decide whether anyone bothers porting a working pipeline. For decision workloads the output side is a handful of tokens anyway, so free output mostly reads as marketing, but 32K context is enough to stuff a document or a long trajectory into a single call rather than chunking it.
For people building on it, there is no posted head to head against Jev yet. The comparisons in the launch are Microsoft's own and the only named competitor in the latency claim is GPT-6 Sol. Treat the accuracy ranking as unverified until a third party publishes a run on a benchmark you recognize, and if you do run one, the perturbation stability claim is the cheapest thing to try to break.

