Jev Daily

Theo and Diogo Almeida say Jev demos oversell it

A short exchange on X lands on the one sentence both of them want developers to hold onto about where Jev stops.

Theo posted that he wants to be more positive about Jev and cannot, because of what people keep pointing it at. His own summary of the job it is good at is narrow and worth copying down word for word, classification of fixed data with mostly known quantities. Everything else, in his telling, is the stupid idea he feels obliged to call out in case silence reads as endorsement.

Diogo Almeida replied that he appreciated both the positivity and the skepticism, and added the sharper version of the complaint. Outlier demos, he said, lead to over-promise and under-deliver. Theo's follow up put a label on the split he is arguing for, system one models being cool without people assuming all models have system two capabilities.

What the argument is actually about

Neither of them is saying Jev is bad. Theo calls it legitimately awesome in the same post where he says he is tired of it. The disagreement is with the demo genre, the single impressive run that gets screenshotted and then read as a general capability. A demo shows the best case. A classifier's interesting property is the worst case, the input just outside the fixed set of choices, and that is the frame almost no demo includes.

That gives you a rough test to run on any Jev post you scroll past. Is the set of possible answers known in advance, and is the data being classified fixed rather than something the model has to go fetch, check or reason through in steps. If yes, Theo's line says you are inside the zone. If the task needs the model to hold a plan across turns or verify something about the world, you are outside it, and the demo you are looking at is not evidence that it works.

For people building on it

This is opinion, not a benchmark. Two people posting agreement on X is not a measurement, and nothing in this exchange contains a number. Treat it as a prior for how to read the next result you see rather than as a result itself. It does line up with Theo's earlier argument that Jev cannot validate because it cannot run tools, which is the same boundary stated from the other direction.

The useful part for anyone deciding how much product to hang on Jev is that the critique is about scope, not quality. If your feature is a fixed-choice decision over text you already have, none of this touches you. If your roadmap slide quietly assumes the classifier will grow into an agent, the two people most publicly enthusiastic about the narrow case are telling you it will not.

Diogo Almeida
@CompleteSkeptic
X
Posting outlier demos leads to over-promise under-deliver
Sep 26, 2026 · View on X
Theo - t3.gg
@theo
X
I’m just so tired of seeing people pushing it to do things it is not good for.
Sep 26, 2026 · View on X
Theo - t3.gg
@theo
X
doing my best to help people understand why system one models are cool without them assuming all models have system two capabilities
Sep 26, 2026 · View on X

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.