Stories
- tamara and Zachi score rows with Jev instead of prompting
Two builders replaced a generative step with per-item scoring, and Theo says the compaction version will make coding agents dumber.
- Theo calls Jev instant compaction a terrible strategy
A per tool call probability filter is not compaction, and the cache math runs the wrong way.
- Kyle Jeong's Stagehand branch cuts Act latency to 0.46s
Jev picks the element, deterministic code does the clicking, and an LLM catches the uncertain calls.
- idan levin's WebMCP run splits Jev and Mercury 2.5
The 112x cost figure comes from a two model harness, not from Jev driving a browser by itself.
- Ira Bodnar says Jev cut his SEO agent cost by 90%
A vendor post reports speed and price across nine pipeline steps, and says nothing about whether the audits still hold up.
- LangChain tested Jev as a judge against LLM judges
The posts announce the comparison but publish none of the numbers, so the four axes are only as good as the write-up behind them.
- OpenRouter clocks Jev over 5x faster than next model
Jev's slowest requests still beat every other model's median, OpenRouter says.
- Theo says Jev cannot validate because it cannot run tools
The argument is about using Jev as a judge for reasoning model output, not about Jev as a model.