Jev Daily

Microsoft ships Decision-1 at $0.042 per million input

Microsoft put a decision model of its own on OpenRouter overnight, claiming it is 4.5x faster than the runner-up. Every accuracy and latency figure so far is Microsoft's, and nobody outside has run it against Jev.

Built with it

Harrison Chase says Open SWE routing cut median cost per task 64%

Harrison Chase says Open SWE moved model choice into the harness, sending each task to the cheapest model that still does the job, tested against quality, and median cost per task dropped 64%. He was agreeing with @yuhasbeentaken, who says DeepSeek v4.1 Flash now handles orchestration, repetitive implementation and verification, with Opus or Sol kept for harder edge cases.

Harrison Chase
@hwchase17
X
most orchestration steps don't need a frontier model
Oct 10, 2026 · View on X

Read the full story

Someone measured it

Microsoft ships Decision-1, OpenRouter lists it at $0.042 per million input

Satya Nadella announced Microsoft-Decision-1 and OpenRouter listed it at $0.042 per million input tokens, output free, 32K context. Microsoft's own numbers, relayed by OpenRouter, claim highest accuracy across 36 blind benchmarks of about 150K questions, 4.5x faster than the runner-up, and decisions flipping on 1.3% of perturbed inputs. It is post-trained from Qwen3.5-9B.

Satya Nadella
@satyanadella
X
It delivers top performance on structured decision tasks, outperforming both LLMs and other decision models in latency and quality.
Oct 9, 2026 · View on X

Read the full story

The case against

Mike Taylor asks if Microsoft-Decision-1 is just constrained output tokens

Mike Taylor questioned whether Microsoft-Decision-1 is a decision model at all, replying to Lisan al Gaib's note on how fast everyone deployed decision models. Nadella's announcement says only that it delivers top performance on structured decision tasks, with no architecture detail; OpenRouter's listing says it was post-trained from Qwen3.5-9B.

Mike Taylor
@hammer_mt
X
Is it actually a decision model though? Or just constrained output tokens?
Oct 9, 2026 · View on X

Read the full story

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.