Jev Daily

Sydney Runkle says a model router cut agent cost 64%

The number comes from one team's own agent, and the eval behind the no quality drop claim is never described.

Sydney Runkle posted that her team built a model router for their coding agent and cut median cost per task by 64%, with what she describes as no measurable drop in quality. The framing is blunt about why it works. Most tasks, she says, do not need top tier intelligence, so the expensive model only gets called when the task earns it.

Sydney Runkle
@sydneyrunkle
X
most tasks don't need top tier intelligence!
Oct 1, 2026 · View on X

Harrison Chase posted the recipe alongside it, in four steps. Understand the tasks. Understand the models. Build the router inside the harness. Track outcomes. The third step is the one doing the argumentative work. Chase's claim is that routing belongs in the agent harness itself, where it can see what the agent is actually about to do, rather than sitting in front of it as a generic classifier over the raw request.

What the number does and does not cover

Median cost per task is a sensible unit for an agent, because the distribution is long tailed and a mean gets dragged around by a handful of runs that spiral. But 64% is a median on one team's own coding agent against their own prior setup, and neither post names the models in the pool, the task mix, or the share of traffic that ended up on the cheap path.

The softer part is the quality claim. No measurable drop implies a measurement, and the posts never describe the eval, the sample size, or what counted as a task passing. On a coding agent, quality is usually where cheap routing goes wrong first, and it tends to go wrong on the hard tail rather than the median, which is exactly the part a median cost figure does not illuminate.

For people building on it, treat 64% as a result from their workload, not a rate you should expect to inherit. Runkle has previously posted open questions on context engineering, and the routing work sits in the same place, a set of choices you have to re-derive against your own traffic. If you copy the recipe, copy step four first and get outcome tracking in place before you start moving traffic off the expensive model.

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.