Jev Daily

OpenAI ships Decisions, Jev builders pick it apart

Diogo Almeida's read on OpenAI's new Decisions API: "seems to be luna with constrained decoding". Kieran Klaassen ran it against Jev and says Jev is still faster and more accurate, though he has not posted the figures in the thread.

Built with it

Kush's Grapevine found 10 of 13 sandbox launches, up from 5

Kush added Jev to Matt Van Horn's /last30days skill, which lets an agent search social media, so Jev screens posts instead of an LLM reading each one.

Kush
@kushbhuwalka
X
social media has sparse signal. you can't search sandbox then have your LLM read through every post - its simply too expensive...
Sep 30, 2026 · View on X

Read the full story

The case against

Kieran Klaassen says Jev still beats OpenAI's Decisions API in his tests

Kieran Klaassen posted a Jev versus Decisions benchmark and said Jev is still faster and more accurate in his testing. Ian Nuttall, reading the same run, called the two very similar, with Jev ahead on many-option and nuanced questions and cheaper on top. Neither post states the accuracy, latency or cost figures.

Kieran Klaassen
@kieranklaassen
X
jev is still faster and more accurate in my testing
Sep 29, 2026 · View on X

Read the full story

Diogo Almeida says OpenAI's Decisions API is Luna with constrained decoding

Diogo Almeida's objection, spelled out after Justin Schroeder asked, is that Decisions looks like a model shared with other objectives rather than a dedicated one, which he says can cause jaggedness. His second objection is constrained decoding (blocking invalid output tokens) on string outputs instead of something machine-native.

Diogo Almeida
@CompleteSkeptic
X
constrained decoding makes models dumber!
Sep 18, 2026 · View on X

Read the full story

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.