Jev Daily

TypeSafe posts Deel's Jev numbers on six classification tasks

Vendor-published customer results, with no independent look at the baselines.

TypeSafe AI published a set of numbers it says Deel measured running Jev on six production classification tasks. All six are what TypeSafe calls a pick from a known set, matching repeat analytics questions to approved answers, picking 1 of 36 metrics, blocking PII requests, deciding when a support chat needs a human, tagging tickets across a 3-level taxonomy, and sorting expenses into about 55 categories.

TypeSafe AI
@typesafeai
X
Speed was measured while shadowing live production traffic: up to 4× faster.
Sep 25, 2026 · View on X
TypeSafe AI
@typesafeai
X
Every one is a pick from a known set.
Sep 25, 2026 · View on X

The accuracy figures

TypeSafe reports repeat-question matching going from 70% to 97%, and expense categorization from 50% to 86% measured against human reviewers. On escalation, the claim is the same catches with fewer false alarms, with no percentage attached. The comparison baseline is described only as frontier LLMs, with no model named and no prompt shown.

The speed figures

Speed was measured two ways. Shadowing live production traffic, meaning Jev ran alongside the real system without serving its answers, produced up to 4x faster. Offline tests came in at 2 to 3x. That gap between shadow and offline is the more interesting number of the pair, since it is the one that suggests the win depends on what the rest of the stack is doing while the classifier runs.

What is missing

This is a vendor post about a customer, and the post is the only artifact. Nobody outside Deel and TypeSafe has seen the eval sets, the baselines, how the frontier models were prompted, or what the 50% expense categorization baseline actually was. A jump from 50% to 86% against human reviewers is large enough that the baseline is the thing a skeptical reader would want first, and it is not there.

For people building on it, the useful part is not the percentages, it is the shape of the six tasks. Every one has a fixed, enumerable answer set, from 36 metrics to roughly 55 expense categories to a binary PII block. That is the workload TypeSafe is putting forward as the one Jev is for, and the results say nothing about anything that needs an open-ended answer or a tool call.

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.