Jev
A decision model from TypeSafe AI that scores options and makes discrete choices. Every Jev Daily story on Jev, newest first. 46 stories since September 2026.
- Chase splits trajectory labeling into typed answers in LangSmith
A shape for agent evals, posted without a cost or accuracy number attached.
- OpenRouter clocks GPT-6 Luna at 180ms, Jev second
The only number OpenRouter disclosed in the post belongs to the model that won.
- Shrivu Shankar says Jev out-calibrated 40 decision models
One post, one axis, and none of the setup that would let you check it.
- Ramp data shows Jev gained a point of adoption in a month
Card spend across Ramp's customer base says a lot about who is trying Jev, and nothing about who kept it.
- Vals AI says Jev matched GPT-6 Astra accuracy at 1/500th the cost
An outside evaluator put a number on the cheap end of the tradeoff, on one narrow task with clean labels.
Know what Jev is good at before you bet on it
One email a day. Under three minutes. What builders found out about Jev yesterday.
- Matthew Berman ran Jev over 193,355 video hooks for $0
A big classification run with a headline price of nothing, and no breakdown of how the price got there.
- Vogel runs OpenAI's Decisions API against Jev and Clef
The public beta landed and the first side by side posted so far reports speed, not accuracy.
- Gregor Zunic made Jev play GTA 5, then called it off
A clip, a follow up captioned "Fun's over boys", and nothing measured in between.
- Harrison Chase says LangChain already ships Breunig's Jev router
The 15 minute sketch has a packaged equivalent, and Chase thinks the interesting version runs the router more than once.
- GPT Researcher now runs on Jev by default
Assaf Elovic says the retrieval path no longer needs embeddings at all, on an eval his own team ran.
- Zachi ends pg-jev over query speed above 50k rows
The Postgres extension that put Jev calls inside queries is done, and the author says the ceiling is scale, not bugs.
- Sydney Runkle says a model router cut agent cost 64%
The number comes from one team's own agent, and the eval behind the no quality drop claim is never described.
- Milind S drives a Mac with Jev at about 90ms per decision
A local CoreML segmenter and on-device OCR feed text to Jev, which returns a probability across the elements on screen.
- OpenRouter benchmarks Jev Router against six other routers
Seven routers, six benchmarks, one weighted index you can re-weight yourself, and no placement numbers in the announcement.
- Vercel added Jev to its AI SDK for Python
Two experiments, one install command, and not a single number attached to either.
- TypeSafe says Jev beat an LLM judge 250x cheaper on a risk agent
One agent, one workload, and the numbers come from the vendor's summary of it.
- Jev lands in n8n as a node for branching and routing
n8n pitches the node as a smarter If/Switch for the fuzzy conditions a regular rule cannot express.
- Shreya Shankar says per-row Jev calls are wrong for batch work
Her argument is about throughput on bulk rows, not whether the answers are any good.
- Jon Kraayenbrink tagged 479 LinkedIn saves for $0.0195
An open source tool runs Jev over your LinkedIn bookmarks and labels topic, hook and format so you can search them.
- Rox says Jev beat GPT-5 Mini on sales reranking by 12%
Three headline numbers from one vendor's internal benchmark, with no dataset or query count attached.
- Kush's Grapevine finds 10 of 13 sandbox launches, up from 5
A Jev screening pass in front of the LLM roughly doubled recall on one social search query, on Kush's own numbers.
- Kieran Klaassen says Jev still beats OpenAI's Decisions API
One builder's own run, with the accuracy, latency and cost numbers living only in the linked benchmark.
- Red team reports 43.5% attack success on Jev
Zhaorun Chen's team says routing Jev's own tool calls through Jev as a gate brings the rate down while keeping most of the utility.
- OpenRouter says Solar Decide beats Jev 1.13 on JevBench
Two rival decision models landed on OpenRouter the same week, both benchmarking themselves against Jev, and neither number is independent.
- Kieran Klaassen uses Jev answers as embedding vectors
Four named dimensions instead of 1,536 anonymous ones, with one Jev call per question per document.
- OpenRouter says Jev took 27% of its classification requests
Platform share reported by the platform, with no accuracy number attached.
- Theo. Jev is for classifying fixed data, not the demos
The loudest Jev critic says he is stuck sounding negative about a tool he likes, and Diogo Almeida agrees the demo stream oversells it.
- Theo and elvis got opposite results from Jev Router
One coding benchmark says the router costs more and runs longer, one small support agent says it halves the bill.
- David shipped jevgrep, claims 40% lower coding agent cost
A CLI that gathers context with Jev before the coding agent starts spending tokens on it.
- TypeSafe posts Deel's Jev numbers on six classification tasks
Vendor-published customer results, with no independent look at the baselines.
- Jack Cheng runs Jev on both ends of email triage
The second classifier takes how well he slept as an input, so the same inbox sorts differently day to day.
- Sydney Runkle posts four open questions on Jev context
A short public checklist of what nobody has settled yet about feeding state and questions to Jev.
- Metaview shipped Jev into every agent on its platform
Shahriar Tajbakhsh reports minutes to seconds on candidate search, with no baseline numbers posted.
- Mike Taylor checked Jev against the probability words chart
An eyeball check against a well known survey chart, with no per-word numbers and no error measure posted.
- Moritz Kremb built a Jev sales copilot that scores closing odds
The demo runs on a recorded call, and no latency, cost or accuracy numbers came with it.
- Jev counts letters at 70%, perfect on character lists
Paolo Rosson found the strawberry question is a coin flip for Jev until you split the word into characters first.
- Hamilton Ulmer shipped prompt_jev() as MotherDuck SQL
Classification moves into the WHERE clause, with speed and cost numbers that are still the vendor's own.
- TypeSafe AI pauses new Jev signups on demand swell
Existing accounts keep working, and nobody has said how long the pause lasts.
- tamara and Zachi score rows with Jev instead of prompting
Two builders replaced a generative step with per-item scoring, and Theo says the compaction version will make coding agents dumber.
- Theo calls Jev instant compaction a terrible strategy
A per tool call probability filter is not compaction, and the cache math runs the wrong way.
- Kyle Jeong's Stagehand branch cuts Act latency to 0.46s
Jev picks the element, deterministic code does the clicking, and an LLM catches the uncertain calls.
- idan levin's WebMCP run splits Jev and Mercury 2.5
The 112x cost figure comes from a two model harness, not from Jev driving a browser by itself.
- Ira Bodnar says Jev cut his SEO agent cost by 90%
A vendor post reports speed and price across nine pipeline steps, and says nothing about whether the audits still hold up.
- LangChain tested Jev as a judge against LLM judges
The posts announce the comparison but publish none of the numbers, so the four axes are only as good as the write-up behind them.
- OpenRouter clocks Jev over 5x faster than next model
Jev's slowest requests still beat every other model's median, OpenRouter says.
- Theo says Jev cannot validate because it cannot run tools
The argument is about using Jev as a judge for reasoning model output, not about Jev as a model.
Get the next one by email
Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.