Kieran Klaassen says Jev still beats OpenAI's Decisions API
One builder's own run, with the accuracy, latency and cost numbers living only in the linked benchmark.
Kieran Klaassen ran a Jev versus Decisions benchmark and said Jev came out ahead. Asked by Ian Nuttall about the Decisions API, Klaassen replied that Jev "is still faster and more accurate in my testing". He posted the benchmark run in a separate reply.
Ian Nuttall, reading the same run, landed in roughly the same place with more hedging. He called the two "very similar" and said Jev is still likely better with many-option questions and nuance, and cheaper on top of that. That is a reader of the benchmark rather than a second independent run, so it is one data set with two people looking at it.
What the posts do not say
No accuracy figure. No latency figure. No price per call. No task list, no sample size, no statement of which Jev configuration or which Decisions setup was used. Everything that would let you decide whether this transfers to your workload sits inside the linked benchmark and not in anything either author wrote on X. Nuttall's qualifier about many-option questions is the closest thing to a scoped claim in the thread, and even that is prefaced with "likely".
Nuttall also asked whether the Decisions API is generally available or a preview that he does not have access to yet. The sources do not contain an answer, so if you are planning a comparison of your own, confirm access before you budget time for it.
For people building on it
This is one builder's run on his own tasks, which is the weakest evidence tier that is still worth reading. It is useful as a signal that someone who works with Jev closely did not immediately switch when an alternative showed up. It is not a head to head you can cite in an architecture doc, and it is not a published eval.
The practical move is the boring one. Open the benchmark, look at what the tasks actually are, and check whether they resemble the decisions you are routing through Jev today. Many-option classification and nuanced judgement calls are the specific places Nuttall flagged an edge, which also means the simple two-way or three-way calls are the place where the gap is most likely to close. If the bulk of your traffic is short binary decisions, this result tells you less than it looks like it does.
Until someone posts accuracy, latency and cost side by side on a task set they describe, the honest summary is that one person prefers Jev and a second person who read his numbers mostly agrees.
jev is still faster and more accurate in my testing
Very similar but Jev still likely better with many-option questions and nuance, and Jev is cheaper too.

