Greg Mushen got 49 of 50 faxes right with 20 questions
A single call underperformed, so he split the decision into a series of questions.
Greg Mushen says he ran Clef-Flash locally over the weekend to classify faxes, and the straightforward version did not work well. His fix was to stop asking for the answer in one shot and instead ask a series of progressive questions, what he describes as a twenty questions style approach. After that change, he says the setup correctly classified 49 out of 50 documents in a hand labeled ground truth set.
He posted the result in reply to Matt Van Horn, who called the original write up excellent and said he needed to dig in.
What is actually in the result
One person, one workload, fifty labeled faxes. That is the whole evidence base. Mushen reports no latency figure, no cost figure, and no breakdown of how many questions each classification took or what the single call baseline actually scored, only that it was not impressive. So 49 of 50 is a number from his bench, not a benchmark, and the honest read is that the pattern worked on his documents rather than that it generalizes.
It is also worth being clear about what the comparison is. He is comparing two ways of prompting the same local model on the same set, not comparing Clef-Flash against another model or against a hosted classifier. There is no head to head here.
For people building on it
The transferable part is the shape, not the score. If a classifier is underperforming on documents with a lot of surface variation, decomposing the label into a chain of narrower questions is cheap to test, and you can run it against fifty of your own hand labeled examples in an afternoon. The cost you take on is more calls per document, which matters if you are paying per call or sitting in a latency budget. Mushen was running locally, so neither constraint shows up in his numbers, and they will show up in yours.
The other thing to copy is the ground truth set. Fifty hand labeled items is small, but it is the reason he could tell the second approach from the first at all. Without it, the first run just feels fine and ships.
The first run out of the gate was not that impressive
it correctly classified 49/50 of the hand labeled ground truth
