Kush's Grapevine finds 10 of 13 sandbox launches, up from 5
A Jev screening pass in front of the LLM roughly doubled recall on one social search query, on Kush's own numbers.
Kush released Grapevine, a patch that puts Jev inside Matt Van Horn's /last30days, a skill that lets an agent search social media. The problem he was solving is recall. Asked how many sandbox launches happened in the past week, the unpatched skill counted 5 out of 13. With Jev in the loop, Kush says his version finds 10 of the 13.
social media has sparse signal. you can't search sandbox then have your LLM read through every post - its simply too expensive...
Jev can go through a 10 times more posts while keeping costs affordable. our patch finds 10 of the 13 sandboxes.
The mechanism is the one Jev keeps getting used for. Instead of having the LLM read every post that comes back from a search, Jev screens posts first, so the expensive model only sees what survives. Kush frames it as a volume argument rather than an accuracy one.
The cost claim is a description, not a measurement
Kush says Jev "can go through a 10 times more posts while keeping costs affordable". There is no dollar figure, no token count and no latency number in the post, so treat that as his characterisation of the workload rather than a benchmark. Likewise the 10 of 13 is one query about one topic on one week of posts. It is a real before and after on the same skill, which is more than most demos offer, but it is a single run and Kush ran it himself.
The miss rate is worth keeping in view too. Three of the thirteen launches still did not surface. A screening layer in front of search widens the net, it does not make the net exhaustive, and the sparse signal problem Kush describes is a property of social media rather than of the model reading it.
What you need to run it
Kush lists three keys for best results, @GetXAPI for X, Jev from TypeSafe AI, and @Tiny_Fish for web search. In a follow up he adds that they call Jev through Vercel's AI Gateway. He also says it can run keyless, which presumably means degraded coverage rather than the numbers above.
Matt Van Horn, whose skill this builds on, replied with "Cool idea!" and nothing more, so there is no independent run of the patched version yet.
For people building on it, the interesting part is not the specific count. It is that the fix for a retrieval pipeline that misses things was to make the cheap read step cheap enough to read more, rather than to make the search smarter. If your agent is skipping results because reading them costs too much, that is the shape of problem this addresses. If it is skipping them because the search index never returned them, Jev is not in that path.
