Jev aced one browser benchmark, scored 1/20 on another
Gregor Zunic shipped a Jev browser agent that found flights in 7 seconds for $0.0039, then benchmarked the same approach at 1 of 20 on long horizon tasks against 17 of 20 for BrowserCode plus Luna. The speed is settled. The reasoning is not.
Built with it
Kyle Jeong put Jev inside Stagehand and cut Act latency 4.3x
Kyle Jeong added Jev as a decision layer in Stagehand, with an LLM fallback when the choice is uncertain. Early evals across 240 real-site task runs put median Act latency at 0.46s, down from 1.97s, with success going from 85% to 95.4%. Elsewhere Milind S got screenshot-free Mac computer use at ~90ms per decision, and Zachi shipped a jev() Postgres extension judging 129 rows in ~1s for $0.0009.
Faster than any LLM computer use I've tried.
Someone measured it
Jev plus WebMCP solved 49 of 49 tasks at 112x lower cost
idan levin ran Jev plus Mercury 2.5 on his open WebMCP benchmark and got 49 of 49 tasks solved, at roughly 112x lower model cost than GPT-6 Astra using computer use with code execution. The same harness without WebMCP, built on Browser Use's Ultrafast, solved 25 of 49. Separately, @fazxes benchmarked Jev against gpt-5.6-luna as the fx auto mode safety classifier and reported 5-18x faster and more accurate.
~5-18x faster and more accurate than 𝚐𝚙𝚝-𝟻.𝟼-𝚕𝚞𝚗𝚊, our current top choice
Vercel says Jev hit 13% of AI Gateway teams on day one
Vercel posted that Jev was adopted faster than any other model in AI Gateway history, reaching about 13% of teams in the first day, 2x the GPT-5.6 family and 6x Fable 5.1. Guillermo Rauch put that partly down to the "AI is too expensive/slow" zeitgeist rather than the product alone. TypeSafe's Diogo Almeida said 140k people came off the waitlist in under 36 hours.
The data and the anecdata on Jev's adoption are shocking.
Writing criteria
Hassan routed Jev's low confidence calls to Kimi K3 for 96 of 100
Hassan gave Jev 100 emails, half legit and half fraudulent, got them all classified in 1.42 seconds, then routed every prediction under 95% confidence to Kimi K3. Thirty one emails crossed that threshold, and the combined pipeline reached 96% accuracy for about $0.07, of which $0.003 was Jev.
The case against
Gregor Zunic's long horizon benchmark: Jev 1 of 20, BrowserCode 17 of 20
Gregor Zunic, who built the 7 second Jev flight demo, reported the Jev agent scoring 1/20 against 17/20 for BrowserCode plus Luna on his long horizon task benchmarks. Theo separately called Jev-based context compaction a terrible strategy, arguing it drops tool calls without knowing their results and throws away provider reasoning traces.
A model with 0 reasoning ability simply can't do that (yet?).
Also on the timeline
Rob Hallam's 6,000 demo users cost him $1.04
Rob Hallam said the 6,000 people who tried his post scoring demo cost $1.04 in total. Training the classifier on tens of thousands of viral posts ran about $5.
Justin Schroeder says Jev needs to be 10x cheaper
Justin Schroeder said Jev at $0.042/M input is 7x a DeepSeek V4.1 Flash cache read and needs to be roughly 10x cheaper for industrial-scale use. TypeSafe's Diogo Almeida publicly agreed.
Frequently asked questions
How much faster is Jev than other models for computer control tasks?
Kyle Jeong's early tests showed Jev cut median response time from 1.97 seconds to 0.46 seconds for browser automation tasks, about 4.3 times faster. In other tests, Jev achieved decisions at roughly 90 milliseconds per step on Mac computer use without screenshots.
What happens when you combine Jev with other AI models?
When Jev's confidence is low, routing uncertain decisions to other models like Kimi K3 or Luna can improve accuracy. One example showed 96% accuracy on email classification by escalating 31 of 100 uncertain cases to a secondary model, with Jev handling 69 cases for just $0.003.
Where does Jev struggle compared to other AI models?
Jev performs well on single-step decisions but struggles on long-horizon tasks requiring multiple connected steps. One benchmark showed Jev solving 1 of 20 long-horizon tasks versus 17 of 20 for BrowserCode plus Luna.
How much does it cost to run Jev?
Jev costs $0.042 per million input tokens. One user ran 6,000 demo users for $1.04 total, and another solved 49 tasks at roughly 112 times lower cost than GPT-6 Astra, though results depend on whether the website exposes tools.
Did people actually start using Jev after launch?
Vercel reported Jev reached about 13% of AI Gateway teams on the first day, 2 times faster adoption than the GPT-5.6 family. However, these are trial numbers from launch week and adoption may not have stayed at those levels.



