Jev Daily

Gregor Zunic made Jev play GTA 5, then called it off

A clip, a follow up captioned "Fun's over boys", and nothing measured in between.

Gregor Zunic posted a video of Jev playing GTA 5. The caption is four words long, "Make Jev play gta 5", and the clip is the whole of the evidence. Hours later he posted a follow up with a second video, captioned "Fun's over boys".

Gregor Zunic
@gregpr07
X
Fun's over boys
Oct 6, 2026 · View on X

That is the entire artifact. There is no repo, no description of how the loop reads the screen, no account of how actions get picked or sent to the game, and no numbers. Nobody has said what the decision latency was, how often the thing did something sensible, or what it cost to run for the length of the clip.

Why it still shows up here

Real time control is the hardest shape of problem to put a classifier in front of, because the budget per decision is set by the frame rate rather than by what you are willing to pay. That is exactly why a video of it is fun, and exactly why a video of it proves very little. Game footage edits well. A loop that picks a plausible action twice and a bad one eight times looks identical to a loop that works, if you choose which seconds to post.

The contrast with the week's other control story is worth a second. When Milind S drove a Mac with Jev, the claim came with a figure, about 90ms per decision, which is a thing you can argue with, reproduce or beat. This has nothing attached to it.

For people building on it, do not read this as a data point about Jev in a real time control loop, in either direction. It is not a benchmark, it is not a demo with a methodology, and the second post reads as the author himself closing the tab. If you are considering Jev for anything with a frame budget, the open questions are still open, starting with what the per decision latency actually is on your hardware and what happens on the frames where the model is wrong.

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.