Jev vs an LLM judge in Edward Irby's risk agent
TypeSafe says an LLM judge scored the same threat 0.35, then 0.68, then 0.50, while Jev ran 250x cheaper on Edward Irby's risk agent. Separately, Shreya Shankar says calling a decision model once per row is the wrong shape for batch work.
Built with it
Jev lands in n8n as a node for branching and routing
n8n says Jev now runs as a node in its workflow builder, returning a choice plus a confidence you can use for branching, sorting and routing. n8n describes it as a smarter If/Switch for the fuzzy stuff.
super excited for this - should make workflows much easier to build!
Someone measured it
TypeSafe says Jev beat an LLM judge 250x cheaper on Irby's risk agent
TypeSafe AI posted Edward Irby's comparison of Jev against an ordinary LLM judge inside an agent that monitors business risks. TypeSafe says the LLM flip-flopped on the same threat, 0.35 to 0.68 to 0.50, while Jev came in 250x cheaper and 3-6x faster at the same report quality. Irby's write-up puts the whole stack at $0.02 per sweep.
the LLM missed 5 of 11 investigations. Jev missed none.
The case against
Shreya Shankar says per-row Jev calls are wrong for batch work
Shreya Shankar says calling a decision model like Jev once per row over thousands of rows is a terrible idea, because it gets no benefit from query planning and sits far from optimal performance. She puts the fastest an H100 could possibly run one AI filter over 5k movie reviews with Qwen3-4B at about 6.6 seconds.
No system today comes close, including Quail, our open-source AI-SQL engine


