Jev Daily

Shrivu Shankar says Jev out-calibrated 40 decision models

One post, one axis, and none of the setup that would let you check it.

Shrivu Shankar says Jev was more calibrated than the 40 other decision models his team tested on email classification. He called the result expected. TypeSafe AI reposted it with the line "Stay calibrated out there!".

Shrivu Shankar
@ShrivuShankar
X
as expected, @typesafeai 's jev was more calibrated than the 40 other decision models we tested on email classification
Oct 8, 2026 · View on X
TypeSafe AI
@typesafeai
X
Stay calibrated out there!
Oct 8, 2026 · View on X

That is the whole claim as posted. Calibration here means a model's stated confidence lining up with how often it is actually right, which is a different property from raw accuracy. A model can be more accurate and worse calibrated, or less accurate and better calibrated, and the post reports only the calibration axis. No accuracy number, no cost number and no latency number appears alongside it.

What is not in the post

The 40 other models are not named. The email classification task is not described beyond that phrase, so there is no label set, no number of classes and no indication of whether the emails are public or internal. The dataset size is not given. The calibration metric is not given either, which matters because expected calibration error, Brier score and a reliability diagram can rank the same set of models differently depending on bucketing. There is no link to a write-up or a repo in the source text, so nobody outside Shankar's team can reproduce the ordering.

This is one team's result on one workload, posted as a one line summary. It is not an eval anyone can rerun, and it is not a head to head against any particular competitor, because the comparison set is a count rather than a list.

For people building on it

If you are routing on confidence, which is the main reason calibration matters in production, a favourable calibration result on someone else's email corpus does not transfer to your label set. The useful move is to take your own held out rows, bucket Jev's confidence scores, and check what fraction in each bucket are correct before you pick a threshold for escalating to a larger model. Shankar's post tells you his team found Jev at the top of a 40 model field. It does not tell you where your own cutoff should sit.

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.