OpenAI ships Decisions, Jev builders pick it apart
Diogo Almeida's read on OpenAI's new Decisions API: "seems to be luna with constrained decoding". Kieran Klaassen ran it against Jev and says Jev is still faster and more accurate, though he has not posted the figures in the thread.
Built with it
Kush's Grapevine found 10 of 13 sandbox launches, up from 5
Kush added Jev to Matt Van Horn's /last30days skill, which lets an agent search social media, so Jev screens posts instead of an LLM reading each one.
social media has sparse signal. you can't search sandbox then have your LLM read through every post - its simply too expensive...
The case against
Kieran Klaassen says Jev still beats OpenAI's Decisions API in his tests
Kieran Klaassen posted a Jev versus Decisions benchmark and said Jev is still faster and more accurate in his testing. Ian Nuttall, reading the same run, called the two very similar, with Jev ahead on many-option and nuanced questions and cheaper on top. Neither post states the accuracy, latency or cost figures.
jev is still faster and more accurate in my testing
Diogo Almeida says OpenAI's Decisions API is Luna with constrained decoding
Diogo Almeida's objection, spelled out after Justin Schroeder asked, is that Decisions looks like a model shared with other objectives rather than a dedicated one, which he says can cause jaggedness. His second objection is constrained decoding (blocking invalid output tokens) on string outputs instead of something machine-native.
constrained decoding makes models dumber!


