Vogel runs OpenAI's Decisions API against Jev and Clef
The public beta landed and the first side by side posted so far reports speed, not accuracy.
OpenAI moved its Decisions API into public beta for all developers, and vogel ran it head to head with Jev and Clef the same day. His verdict in the post is about one thing, speed. He says it is fast, and links a full benchmark he ran.
OpenAI just released their Decisions API so I put it head to head with Jev & Clef
Let your app choose the right model, tool, or action in near real-time with Decisions API, now available to all developers in public beta.
OpenAI's own pitch is on the same axis. The company says the Decisions API makes decisions up to ten times faster than GPT-6 Luna through the Responses API, and frames the product as letting an app choose the right model, tool, or action in near real time. That is a vendor claim about a vendor baseline, which is to say it is a comparison against OpenAI's own slower path rather than against anything in the Jev family.
What is not in the posts
No accuracy number. No calibration number. No cost per thousand decisions. vogel's post reports the speed impression and points at a benchmark, and OpenAI's post reports the ten times figure against Luna. Nobody in either post has published the three systems lined up on how often they are right, which is the number that decides whether a faster decision is worth taking.
This is the second week running that the Decisions API has been measured in public by people outside OpenAI. Kieran Klaassen earlier said Jev still beat it in his own tests, and Diogo Almeida read the product as Luna with constrained decoding. vogel's run is a separate test by a separate person on a separate workload, so it is not a rebuttal of either, and treating the three posts as one trend would be reading more into them than they contain.
For people building on it
If you are deciding whether to route classification or tool selection through the Decisions API, the only published axis right now is latency, and the only latency figure with a named baseline is OpenAI's own. A routing layer that answers faster and answers differently is a different system, not a drop in replacement, and the posts here do not tell you how differently. Run your own labelled set before you swap anything that is already in production, and look at vogel's linked benchmark for the detail his tweet does not carry.

