Jev Daily

Jev replaced embeddings in GPT Researcher's RAG

Assaf Elovic says swapping embeddings for Jev in GPT Researcher lifted relevant context from 46% to 73% at the same cost per report. Meanwhile a red team says Jev 1.13 follows injected instructions 43.5% of the time.

Someone measured it

Assaf Elovic swapped embeddings for Jev in GPT Researcher's RAG

Assaf Elovic replaced embeddings with Jev in GPT Researcher's retrieval pipeline and tested both across 28 tasks from SimpleQA and open ended research. He reports 73% relevant context against 46%, reports preferred 15 to 3 in blind comparisons, and the same cost per report. GPT Researcher now runs on Jev by default and drops embeddings entirely.

Harrison Chase
@hwchase17
X
retrieval is a decision problem, not just a similarity problem
Sep 28, 2026 · View on X

Read the full story

The case against

OpenRouter says Solar Decide beats Jev on JevBench, on both axes

Two rival decision models landed on OpenRouter, and OpenRouter says Upstage's Solar Decide beats Jev 1.13 on JevBench at 87.0% versus 86.1% accuracy and 0.14s versus 0.30s median latency, at $0.10/M input with free output. Respan's Span-01, also now live, claims to be 2x cheaper and 18% better than Jev on a behavior benchmark, at $0.02/M input.

Respan
@RespanAI
X
Span-01 Lite: Better than Jev and completely free!
Sep 25, 2026 · View on X

Read the full story

Red team reports 43.5% attack success on Jev via prompt injection

Zhaorun Chen red-teamed Jev 1.13 on DTap and reported 70.1% attack success under direct misuse and 43.5% under indirect prompt injection, including exfiltrating user data and deleting files. The mitigation his team found is using Jev as a self-gating layer on its own tool calls, which he says cuts the rate while keeping most utility.

Diogo Almeida
@CompleteSkeptic
X
don't just plug jev into high-level decisions, but program the behavior you want!
Sep 28, 2026 · View on X

Read the full story

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.