Jev Daily

Harrison Chase says Open SWE routing cut cost per task 64%

Model selection moved out of the agent and into the harness, with the quality bar as the gate.

Harrison Chase says Open SWE now picks the model inside the harness rather than pinning every step to one frontier model. Each task goes to the cheapest model that still does the job, tested against quality, and median cost per task dropped 64%.

Harrison Chase
@hwchase17
X
most orchestration steps don't need a frontier model
Oct 10, 2026 · View on X

That is the whole claim as posted. Chase does not say which models the router chooses between, what the quality test is, how the quality bar is set, or over how many tasks the median was taken. There is no latency figure and no absolute dollar number, only the relative move in the median.

What prompted it

Chase was agreeing with @yuhasbeentaken, who had posted about running DeepSeek v4.1 Flash across more and more of a coding workflow. In that account, the cheaper model coordinates coding tasks, testing, research and subagents, handles repetitive implementation once the plan is clear, and does verification work like reviewing code changes and running checks often enough to be constant. Opus or Sol still get pulled in for harder edge cases. That is one person's experience report, not a benchmark, and no numbers came with it.

Chase's reply framed the general version of it. Most orchestration steps, he says, do not need a frontier model, which is the premise the routing in Open SWE is built on.

For people building on it

The 64% is a median for one coding agent on its own task mix, gated by its own quality test. It is not a number you can expect to transfer, and the part that does most of the work here is the quality check, not the routing. A harness that sends work to the cheapest model without a test that catches the cases where the cheap model is wrong gets the cost reduction and a quieter failure mode. Chase's post says the routing is tested against quality but does not describe the test, so the mechanism that makes the saving safe is the part that is least specified.

Yum⋆₊˚
@yuhasbeentaken
X
i still bring in opus or sol for harder edge cases
Oct 9, 2026 · View on X

Earlier on this story

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.