OpenRouter says Solar Decide beats Jev 1.13 on JevBench
Two rival decision models landed on OpenRouter the same week, both benchmarking themselves against Jev, and neither number is independent.
OpenRouter put two Upstage models live, Solar Mini 4 and Solar Decide, and published a comparison that has Jev losing on both axes. On JevBench, OpenRouter says solar-decide scores 87.0% accuracy against 86.1% for Jev 1.13, with 0.14s median latency against 0.30s. Both new models are listed at 50% off for a limited time.
Solar Decide is built on Solar Mini 4 and runs on the OpenRouter Decisions API, returning a structured decision rather than free text. Pricing is $0.10 per million input tokens with output free. The base model it sits on is a 35B mixture of experts with 3B active parameters and a 524K context window, priced at $0.05/M input and $0.20/M output, with tool calling, structured outputs and optional reasoning.
The other one is aimed at agent traces
Respan's Span-01 and Span-01 Lite also went live on OpenRouter the same week. They take a span, meaning the system, user, assistant and tool messages of an agent turn, plus a list of behaviors you care about, and return the probability each behavior is present. OpenRouter's examples are questions like whether the user is frustrated or whether a tool call is safe to run, with use cases including gating a risky tool call before it runs and labeling eval traces in bulk. There is no cap on span size. Span-01 is $0.02 per million input tokens with free output, and Lite is free.
Respan's own launch post claims Span-01 is 2x cheaper and 18% better than Jev, and 700x cheaper and 4% better than GPT-6 Luna, and puts it at number one on what it calls the Behavior Benchmark. That claim comes from the vendor, not from a third party run.
For people building on it, neither comparison is independent. The JevBench figures are published by OpenRouter alongside the launch of the model that wins them, and the Behavior Benchmark numbers come from Respan. OpenRouter has previously published latency numbers that went the other way, when it earlier clocked Jev over 5x faster than the next model on its own eval. A 0.9 point accuracy gap on someone else's benchmark is not a reason to swap a production classifier, but the pricing gap, free output on both and free Lite, is cheap enough to run your own eval against your own traces.
On JevBench, solar-decide beats Jev 1.13 on both: 87.0% vs 86.1% accuracy, and 0.14s vs 0.30s median latency.
Span-01 Lite: Better than Jev and completely free!

