Jev Daily

Theo offers $10,000 to audit BridgeMind's nerf benchmark

No harness, no task list, no run counts, and a $10,000 wager on whether model nerfing is real.

Theo went through BridgeMind's NerfBench post and listed what it does not disclose. The harnesses used, the tasks used, how many times each task is run, what is done to identify variance in daily runs, what analysis is done on bad runs, and which APIs are being used, which he says matters a lot. He also asked how the plus or minus 10% variance band was selected, why tokens and costs are weighted the same when costs are the metric that matters, and how many runs each dot on the chart represents.

Theo - t3.gg
@theo
X
He ran 6 of 30 tests in a benchmark, and when 2 failed he claimed "NERF" because he failed to run the other 24 tests.
Oct 3, 2026 · View on X

He then put money against it. BridgeMind put $10,000 towards benching, so Theo said he will do the same. If his own benchmarks show meaningful model nerfing aligned with BridgeMind's viral post, he donates another $10,000 to a charity of BridgeMind's choice. If they show the opposite, he wants the nerf bench page taken down and replaced with an apology listing the flaws, written by him. He offered to have a qualified third party audit both sides.

The April claim

Mid thread, Theo added that while prepping his own bench he looked at the earlier Opus 4.6 nerf claim from April and said it was wrong. By his account the benchmark had 30 tests, 6 were run, 2 failed, and the nerf call was made without running the other 24.

BridgeMind's public reply did not address any of the methodology questions. It said Theo is obsessed with BridgeMind, that he hates seeing it win, that NerfBench is not hard to understand, and linked to an explanation. Theo also says he sent a DM asking for traces from the runs.

For people building on this, none of it is about Jev. It is useful here as a price list. A chart that claims a model got worse, without a harness, a task list, a run count or a variance method, is an anecdote with error bars drawn on afterwards, and the same standard applies to every Jev number that shows up with a screenshot and no methodology. Theo has not published his alternative bench yet, so at this stage there are questions and a wager, not a result.

BridgeMind
@bridgemindai
X
Theo is obsessed with BridgeMind. He hates seeing me win.
Oct 2, 2026 · View on X

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.