Jev Daily

Jev vs an LLM judge in Edward Irby's risk agent

TypeSafe says an LLM judge scored the same threat 0.35, then 0.68, then 0.50, while Jev ran 250x cheaper on Edward Irby's risk agent. Separately, Shreya Shankar says calling a decision model once per row is the wrong shape for batch work.

Built with it

Jev lands in n8n as a node for branching and routing

n8n says Jev now runs as a node in its workflow builder, returning a choice plus a confidence you can use for branching, sorting and routing. n8n describes it as a smarter If/Switch for the fuzzy stuff.

Diogo Almeida
@CompleteSkeptic
X
super excited for this - should make workflows much easier to build!
Oct 1, 2026 · View on X

Read the full story

Someone measured it

TypeSafe says Jev beat an LLM judge 250x cheaper on Irby's risk agent

TypeSafe AI posted Edward Irby's comparison of Jev against an ordinary LLM judge inside an agent that monitors business risks. TypeSafe says the LLM flip-flopped on the same threat, 0.35 to 0.68 to 0.50, while Jev came in 250x cheaper and 3-6x faster at the same report quality. Irby's write-up puts the whole stack at $0.02 per sweep.

TypeSafe AI
@typesafeai
X
the LLM missed 5 of 11 investigations. Jev missed none.
Oct 1, 2026 · View on X

Read the full story

The case against

Shreya Shankar says per-row Jev calls are wrong for batch work

Shreya Shankar says calling a decision model like Jev once per row over thousands of rows is a terrible idea, because it gets no benefit from query planning and sits far from optimal performance. She puts the fastest an H100 could possibly run one AI filter over 5k movie reviews with Qwen3-4B at about 6.6 seconds.

Shreya Shankar
@sh_reya
X
No system today comes close, including Quail, our open-source AI-SQL engine
Oct 1, 2026 · View on X

Read the full story

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.