Jev Daily

Theo bets $10k against BridgeMind's nerf bench

Theo says BridgeMind called a model nerf after running 6 of 30 tests, and is putting $10,000 behind an audit. If you cite that benchmark in a build-or-buy argument, read the terms first.

Built with it

Vercel put Jev in the AI SDK for Python

Vercel added Jev to its AI SDK for Python and shipped two experiments with it: detecting Python versus English as text is typed, and writing Python one decision at a time. Install is `uv add ai`. Guillermo Rauch called it a wonderful writeup; neither post gives accuracy, latency or cost for either experiment.

Read the full story

Someone measured it

OpenRouter benchmarks Jev Router against six other routers

OpenRouter published side by side router benchmarks covering seven routers including Jev Router, Unbiased Pareto, NVIDIA Switchyard and its own Auto Router, across six benchmarks. The Router Index is weighted 60% quality, 20% time per task and 20% cost, with a slider on the page to reweight for your own priorities.

OpenRouter
@OpenRouter
X
Routers aren't always better.
Oct 2, 2026 · View on X

Read the full story

The case against

Theo offers $10,000 to audit BridgeMind's nerf benchmark

Theo listed everything BridgeMind's nerf benchmark post does not disclose: harnesses, tasks, run counts, variance handling, which APIs were used, and how the plus or minus 10% variance band was chosen. He offered to match BridgeMind's $10,000, donating again to charity if his own bench confirms nerfing, and asked for a third party auditor.

Theo - t3.gg
@theo
X
He ran 6 of 30 tests in a benchmark, and when 2 failed he claimed "NERF" because he failed to run the other 24 tests.
Oct 3, 2026 · View on X

Read the full story

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.