OpenRouter: Jev 5x faster, same accuracy on 30 labels
OpenRouter ran Jev and four LLMs over the same 200 classification cases and says Jev's slowest requests beat every other model's median, with accuracy inside a handful of cases. Theo, meanwhile, calls ranking reasoning model outputs the worst use he has seen for it.
Someone measured it
OpenRouter clocks Jev over 5x faster than the next model
OpenRouter ran Jev against four popular LLMs on the same 200 cases with Ori Eval, labeling incoming requests as one of 30 task types. Jev was more than 5x faster than the next fastest model, all five landed within a handful of cases of each other on accuracy, and Jev came in second cheapest behind Qwen3.8 Flash.
A decision model was as correct as a full LLM.
LangChain tested Jev as an eval judge against LLM judges
LangChain ran Jev against LLM judges on accuracy, repeatability, latency and cost, to see whether a System One model works for agent evaluation. Sydney Runkle, sharing the guide by Sean and Daniel, said Jev as a judge is cheaper and more precise for online evals. The posts themselves carry no figures, so the comparison lives in the write-up.
jev as a Judge proves to be a cheaper and more precise alternative to LLM as a judge for online evals.
The case against
Theo says Jev cannot validate because it cannot run tools
Theo pushed back on using Jev to validate model output, arguing anything simple enough for Jev to check is something 99% of modern models get right anyway. His second point is structural: Jev cannot run tools or modify its own context, so there is a ceiling on what validation means here.
Ranking outputs from expensive reasoning models might be the worst case I’ve seen for it thus far.


