Jev Daily

Jev aced one browser benchmark, scored 1/20 on another

Gregor Zunic shipped a Jev browser agent that found flights in 7 seconds for $0.0039, then benchmarked the same approach at 1 of 20 on long horizon tasks against 17 of 20 for BrowserCode plus Luna. The speed is settled. The reasoning is not.

Built with it

Kyle Jeong put Jev inside Stagehand and cut Act latency 4.3x

Kyle Jeong added Jev as a decision layer in Stagehand, with an LLM fallback when the choice is uncertain. Early evals across 240 real-site task runs put median Act latency at 0.46s, down from 1.97s, with success going from 85% to 95.4%. Elsewhere Milind S got screenshot-free Mac computer use at ~90ms per decision, and Zachi shipped a jev() Postgres extension judging 129 rows in ~1s for $0.0009.

Milind S
@milindlabs
X
Faster than any LLM computer use I've tried.
Sep 17, 2026 · View on X

Someone measured it

Jev plus WebMCP solved 49 of 49 tasks at 112x lower cost

idan levin ran Jev plus Mercury 2.5 on his open WebMCP benchmark and got 49 of 49 tasks solved, at roughly 112x lower model cost than GPT-6 Astra using computer use with code execution. The same harness without WebMCP, built on Browser Use's Ultrafast, solved 25 of 49. Separately, @fazxes benchmarked Jev against gpt-5.6-luna as the fx auto mode safety classifier and reported 5-18x faster and more accurate.

Pranit
@fazxes
X
~5-18x faster and more accurate than 𝚐𝚙𝚝-𝟻.𝟼-𝚕𝚞𝚗𝚊, our current top choice
Sep 16, 2026 · View on X

Vercel says Jev hit 13% of AI Gateway teams on day one

Vercel posted that Jev was adopted faster than any other model in AI Gateway history, reaching about 13% of teams in the first day, 2x the GPT-5.6 family and 6x Fable 5.1. Guillermo Rauch put that partly down to the "AI is too expensive/slow" zeitgeist rather than the product alone. TypeSafe's Diogo Almeida said 140k people came off the waitlist in under 36 hours.

Guillermo Rauch
@rauchg
X
The data and the anecdata on Jev's adoption are shocking.
Sep 18, 2026 · View on X

Writing criteria

Hassan routed Jev's low confidence calls to Kimi K3 for 96 of 100

Hassan gave Jev 100 emails, half legit and half fraudulent, got them all classified in 1.42 seconds, then routed every prediction under 95% confidence to Kimi K3. Thirty one emails crossed that threshold, and the combined pipeline reached 96% accuracy for about $0.07, of which $0.003 was Jev.

The case against

Gregor Zunic's long horizon benchmark: Jev 1 of 20, BrowserCode 17 of 20

Gregor Zunic, who built the 7 second Jev flight demo, reported the Jev agent scoring 1/20 against 17/20 for BrowserCode plus Luna on his long horizon task benchmarks. Theo separately called Jev-based context compaction a terrible strategy, arguing it drops tool calls without knowing their results and throws away provider reasoning traces.

Gregor Zunic
@gregpr07
X
A model with 0 reasoning ability simply can't do that (yet?).
Sep 19, 2026 · View on X

Also on the timeline

Rob Hallam's 6,000 demo users cost him $1.04

Rob Hallam said the 6,000 people who tried his post scoring demo cost $1.04 in total. Training the classifier on tens of thousands of viral posts ran about $5.

Justin Schroeder says Jev needs to be 10x cheaper

Justin Schroeder said Jev at $0.042/M input is 7x a DeepSeek V4.1 Flash cache read and needs to be roughly 10x cheaper for industrial-scale use. TypeSafe's Diogo Almeida publicly agreed.

Frequently asked questions

How much faster is Jev than other models for computer control tasks?

Kyle Jeong's early tests showed Jev cut median response time from 1.97 seconds to 0.46 seconds for browser automation tasks, about 4.3 times faster. In other tests, Jev achieved decisions at roughly 90 milliseconds per step on Mac computer use without screenshots.

What happens when you combine Jev with other AI models?

When Jev's confidence is low, routing uncertain decisions to other models like Kimi K3 or Luna can improve accuracy. One example showed 96% accuracy on email classification by escalating 31 of 100 uncertain cases to a secondary model, with Jev handling 69 cases for just $0.003.

Where does Jev struggle compared to other AI models?

Jev performs well on single-step decisions but struggles on long-horizon tasks requiring multiple connected steps. One benchmark showed Jev solving 1 of 20 long-horizon tasks versus 17 of 20 for BrowserCode plus Luna.

How much does it cost to run Jev?

Jev costs $0.042 per million input tokens. One user ran 6,000 demo users for $1.04 total, and another solved 49 tasks at roughly 112 times lower cost than GPT-6 Astra, though results depend on whether the website exposes tools.

Did people actually start using Jev after launch?

Vercel reported Jev reached about 13% of AI Gateway teams on the first day, 2 times faster adoption than the GPT-5.6 family. However, these are trial numbers from launch week and adoption may not have stayed at those levels.

Built from 4,009 posts by the builders, researchers and critics we follow on X over 7 days.

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.