Assaf Elovic swapped embeddings for Jev in GPT Researcher
One team's 28 task eval, on their own product, now shipped as the default.
Assaf Elovic replaced embeddings with Jev in GPT Researcher's RAG pipeline and ran both configurations across 28 research tasks drawn from SimpleQA and from open ended research prompts. He reports 73% relevant context with Jev against 46% with embeddings, which he frames as 59% more relevant context, and reports preferred 15 to 3 in blind comparisons. Cost per report came out the same.
GPT Researcher now runs on Jev by default, and no longer needs embeddings at all.
On the back of that, GPT Researcher now runs on Jev by default and, in Elovic's words, no longer needs embeddings at all. His recommended stack is LangChain plus Tavily plus Jev.
Retrieval as a decision, not a distance
The framing is the part that traveled. Harrison Chase picked up the experiment and put it in one line, saying retrieval is a decision problem rather than only a similarity problem, and predicting decision models will show up all over the harness. Embedding retrieval ranks chunks by vector distance and hopes proximity in that space matches usefulness. Swapping in a decision model means asking a classifier whether a given chunk is relevant to this question, which is a different question with a different failure mode.
Sydney Runkle bucketed it further, grouping most of the Jev use cases so far under context optimization and listing retrieval alongside tool and skill selection and context offloading. That is consistent with the open questions she posted about engineering context for Jev.
What the numbers do and do not cover
This is one team's eval of one pipeline, run by the people who ship that pipeline, and nobody outside GPT Researcher has rerun it. Twenty eight tasks is small, and the preference result is 18 comparisons with a stated split of 15 to 3, which leaves no room to argue about margins in either direction. Elovic does not post the relevance grading method in the tweet, and the cost claim is per report at parity rather than a token or latency breakdown.
For people building on it, the interesting number is the one that did not move. Same cost per report means the per chunk decision calls did not blow up the budget relative to embedding and vector search, at least at this corpus size and on this workload. That is the constraint most teams would assume kills the idea before they try it, so it is worth checking against your own document volume before assuming it holds. A pipeline that answers 28 research questions is not a pipeline that retrieves over millions of chunks, and nothing here tests that.
retrieval is a decision problem, not just a similarity problem

