Diogo Almeida calls OpenAI's Decisions API Luna with constrained decoding
An architecture complaint, not a benchmark, and nobody has posted accuracy numbers for the new API yet.
Maz posted OpenAI's Decisions API with the line "OpenAI said hey guys we have jev too". Diogo Almeida's first reaction was four words of diagnosis, that it "seems to be luna with constrained decoding", followed by a frowning face. Justin Schroeder asked him to say what he actually meant, and Almeida spelled it out.
seems to be luna with constrained decoding
The two objections
First, he reads Decisions as a model shared with other objectives rather than its own dedicated model, which he says can cause jaggedness. Second, he says they use constrained decoding on top of string outputs instead of making the interface more machine-native. Constrained decoding is the trick of blocking invalid output tokens at sampling time so the result always parses.
Almeida has a standing argument against that trick, posted a few days earlier in reply to another user. His reasoning runs like this. Constrained decoding works by removing probability from invalid tokens. If the model ever assigns any probability to an invalid token, then by definition the model is confused, and removing the visible signs of that confusion pushes it further out of distribution. He promised a longer blog post on it one day, which is to say the long version does not exist yet.
What this is and is not
None of this is a measurement. Almeida has not posted an eval of the Decisions API, and neither has anybody else in these sources. He is inferring an architecture from the shape of the product and objecting to the architecture. That is a different kind of claim from the numbers people have been posting about latency and classification accuracy, and it should be read as one.
It is also not a pile-on. Almeida's other post on the launch was friendlier than the critique suggests. He opened with "begun, the clone war has", immediately walked it back as a joke, and said he loves OpenAI and thinks more competition and validation is great for developers, with the parenthetical hope that the model turns out to be good. He added that he hopes this signals building in a system one compatible way.
For people building on it, the useful thing here is the test to run rather than the verdict. If Almeida is right that constrained decoding on string outputs costs accuracy, that shows up as errors on the hard cases where the model was already uncertain, not on the easy ones. Anybody evaluating Decisions against Jev should sort their eval set by confidence and look at the low end, because that is where the two objections would bite, and right now nobody has published that comparison.
constrained decoding makes models dumber!
