Theo bets $10k against BridgeMind's nerf bench
Theo says BridgeMind called a model nerf after running 6 of 30 tests, and is putting $10,000 behind an audit. If you cite that benchmark in a build-or-buy argument, read the terms first.
Built with it
Vercel put Jev in the AI SDK for Python
Vercel added Jev to its AI SDK for Python and shipped two experiments with it: detecting Python versus English as text is typed, and writing Python one decision at a time. Install is `uv add ai`. Guillermo Rauch called it a wonderful writeup; neither post gives accuracy, latency or cost for either experiment.
Someone measured it
OpenRouter benchmarks Jev Router against six other routers
OpenRouter published side by side router benchmarks covering seven routers including Jev Router, Unbiased Pareto, NVIDIA Switchyard and its own Auto Router, across six benchmarks. The Router Index is weighted 60% quality, 20% time per task and 20% cost, with a slider on the page to reweight for your own priorities.
Routers aren't always better.
The case against
Theo offers $10,000 to audit BridgeMind's nerf benchmark
Theo listed everything BridgeMind's nerf benchmark post does not disclose: harnesses, tasks, run counts, variance handling, which APIs were used, and how the plus or minus 10% variance band was chosen. He offered to match BridgeMind's $10,000, donating again to charity if his own bench confirms nerfing, and asked for a third party auditor.
He ran 6 of 30 tests in a benchmark, and when 2 failed he claimed "NERF" because he failed to run the other 24 tests.

