Jev Daily

Vals AI says Jev matched GPT-6 Astra accuracy at 1/500th the cost

An outside evaluator put a number on the cheap end of the tradeoff, on one narrow task with clean labels.

Vals AI says it independently tested TypeSafe's Jev against 11 other models on 400 claim-verification questions, and that Jev matched GPT-6 Astra's 97.5% accuracy at roughly 1/500th the cost. The framing in the post is the obvious question about a model that returns decisions rather than prose, which is whether it can hold its own against frontier LLMs on a task where the frontier model is the reference.

Two things make this worth more than a vendor chart. The evaluator is not TypeSafe, and the comparison set is 11 other models rather than a single convenient baseline. What Vals AI has not published in the post itself is the breakdown by model, the latency, or what the error cases looked like.

What 400 questions can and cannot tell you

Claim verification is about as friendly as a task gets for a decision model. The labels are clean, the output space is small, and the question is phrased for you. A 500x cost ratio on that workload is the shape of the best case, not a figure to plug into a forecast for document triage, routing or anything where the input is messy and the right answer is arguable. Treat it as a ceiling.

It is also 400 questions. That is enough to separate 97.5% from something clearly worse, and not enough to tell you much about the tail, which is usually where the per-call savings get eaten back by a human reviewing output.

OpenRouter starts ranking decision models

Separately, OpenRouter launched Decision Model Rankings, showing spend share and token share across decision models broken out by task type, including Relevance, Correctness and Instruction Following. OpenRouter says TypeSafe is leading all categories today. That is a usage measure, not a quality one, and it follows the same house's earlier post that Jev took 27% of its classification requests. The same provider has also published a benchmark where Solar Decide beat Jev on JevBench, so share of spend and winning an eval are clearly not the same thing there either.

For people building on it, the useful number from Vals AI is the accuracy parity, not the cost ratio. The ratio is a property of the task.

Vals AI
@ValsAI
X
On 400 claim-verification questions, it matched GPT-6 Astra’s 97.5% accuracy at ~1/500th the cost.
Oct 6, 2026 · View on X
OpenRouter
@OpenRouter
X
@typesafeai is leading all categories today.
Oct 7, 2026 · View on X

Get the next one by email

Jev, read daily so you do not have to. The builds, the benchmarks, the criteria that worked and the cases where it lost, from the people shipping on TypeSafe AI's System One model.