Nilforoshan says Decisions API cost 2x and scored worse
A private benchmark on a product serving 2.5 million users puts the Decisions API behind Jev on both price and score.
Hamed Nilforoshan ran OpenAI's Decisions API against Jev on a live ranking task and posted the result as a thread. The task is predicting how relevant a user query or resume is to a job description, scored on a scale of 1 to 10, for a product he says serves 2.5 million users. His summary is that OpenAI is 2x more expensive and 5 to 10% worse.
That is two numbers, and only one of them is a quality number. The 5 to 10% gap is stated as a range rather than a single figure, and the post does not say in the text we have which metric the range is measured on, what volume was run, or which Jev configuration was on the other side. The cost comparison is a 2x multiple with no absolute price attached. Anyone deciding between the two on cost alone is going to want the per-call figure, and it is not in the post.
One company's data, one task
What makes this worth reading is also what limits it. It is a private benchmark on a real product rather than a public eval, which means the distribution is real traffic and not a scraped test set, and it also means nobody else can rerun it. Diogo Almeida, who has been skeptical of the Decisions API from the start, replied with "yay to domain specific private benchmarks!" That is an endorsement of the method, not a verification of the numbers.
yay to domain specific private benchmarks!
It lands alongside Kieran Klaassen's earlier finding that Jev still beat the Decisions API in his own tests. Two people testing on their own workloads and reaching the same direction is more interesting than one, but it is still two private setups with different tasks, not a head to head, and neither has published numbers a third party can reproduce.
For people building on it, the useful part is the shape of the task rather than the verdict. This is graded relevance scoring on a 1 to 10 scale, which is closer to reranking than to classification, and results on that kind of task do not transfer cleanly to binary routing or extraction. If your workload is a scalar score over pairs of documents, Nilforoshan's thread is the closest public signal you have. If it is anything else, you are back to running it yourself, which is exactly what he did.
OpenAI is 2x more expensive and 5-10% worse

