TypeSafe says Jev beat an LLM judge 250x cheaper on a risk agent
One agent, one workload, and the numbers come from the vendor's summary of it.
TypeSafe AI posted a comparison from Edward Irby, who built an agent on You.com that monitors business risks and swapped an ordinary LLM judge for Jev. According to TypeSafe's summary, report quality came out the same, Jev ran 250x cheaper and 3 to 6x faster, and the LLM judge missed 5 of 11 investigations while Jev missed none.
The detail that does the most work is the instability. TypeSafe says the LLM scored the same threat 0.35, then 0.68, then 0.50. For a monitoring agent that is supposed to decide whether something crosses a threshold, a judge that wanders by 33 points on an unchanged input is not a cost problem, it is a correctness problem.
What the stack actually is
Irby's own write-up prices the whole sweep at $0.02 and lists four pieces. You.com for search, Jev for typed judgments, Qwen for proposals and synthesis, and MCP for integration. Jev is not doing the research here. It is the scoring layer sitting between a search tool and a generation model, which is the same shape as several earlier builds, including Jack Cheng's email triage pipeline.
That shape matters for reading the 250x. The comparison is Jev against an LLM judge for the judging step, not Jev against a frontier model for the whole task. Qwen is still writing the proposals and the synthesis. The $0.02 covers everything, so the judging slice is smaller than that.
How much to trust it
This is the vendor summarizing one builder's result on one agent, and the 11 investigations are Irby's, not a published benchmark. Nobody outside has rerun them, and the source posts do not say which LLM served as the judge or how a miss was defined. A 5 of 11 miss rate against a perfect score is a wide gap, which is the kind of gap that usually gets narrower when someone else sets up the comparison.
For people building on it, the reusable finding is not the multiplier. It is that an LLM asked for a number on the same input three times gave three different numbers, which is a thing you can check on your own workload this afternoon for far less than $0.02 a sweep.
the LLM missed 5 of 11 investigations. Jev missed none.
