Chase splits trajectory labeling into typed answers in LangSmith
A shape for agent evals, posted without a cost or accuracy number attached.
Harrison Chase says Jev as a judge in LangSmith evals does not return a single pass or fail for an agent run. Difficulty and correctness each get their own typed answer, in one pass, on every trace.
trajectory labeling is really several questions, not one pass/fail
He was responding to a post from Vaibhav Tulsyan, who argued that trajectory labeling at scale is key for RSI and that every trajectory needs assessing across several dimensions. Tulsyan listed four. How difficult the task was, whether the agent's solution was correct, whether the agent considered a diverse set of design choices before picking a route, and whether the agent's thinking is directionally correct. Chase called it a nice framing and said that is how the LangSmith setup already works.
Why multiple dimensions, and why cheap ones
Tulsyan's argument for splitting the label is volume. He asks the reader to imagine answering those questions for millions of agents, each of which spawns tens or hundreds of subagents, and says token cost for trajectory labeling needs to be driven down a lot to do that at scale. That is the gap a small decision model is supposed to sit in. A judge that costs real money per trace gets run on a sample. A judge that costs almost nothing gets run on every trace, which is what Chase is claiming here.
The practical difference is in what you get back. A single pass or fail collapses a hard task done well and an easy task done badly into the same bucket. Separate typed fields for difficulty and correctness keep them apart, so you can filter the traces where the agent failed on something easy, which is usually where the bugs are.
For people building on it, neither post carries a cost figure, an accuracy figure, a trace count or a comparison against an LLM judge on the same traces. Chase describes how the feature behaves, not how well it labels. Tulsyan is reacting to the idea, not reporting a run. Treat this as a shape to copy in your own eval harness, and measure the agreement with your human labels yourself before you trust the difficulty field.
Token cost for trajectory labelling needs to be driven down a lot to do this at scale.

