Rishi Patidar

The grunt work

Evals come in three levels, and the thing you are optimising is not accuracy.

The last post ended with us finally building the thing we should have built first. This is what that actually looks like.

Start with why normal testing doesn't transfer. In ordinary software one input produces one output and you assert on it. An LLM hands you a different shape every time you ask. A bug isn't a wrong line you can point at; it spreads across inputs and across every stage of the pipeline at once.

Without evals that shows up in three ways. You fix one failure and another appears somewhere else. You have no way to know whether the new prompt is better than the old one. And the prompt grows: ours reached twenty pages of accumulated special cases, several of which contradicted each other, and nobody could tell you which lines still mattered.

The reframing that helped most was realising what you are actually optimising. It isn't accuracy. It's how fast you can tell whether a change helped. If finding out takes days, you stop making changes. If it takes minutes, you make ten before lunch.

Evals come in three levels, and it's worth knowing which one you need before building anything.

Level one is unit tests. Cheapest and fastest. Structure the output so a program can assert on it, then test shape and constraints rather than exact matches. A count should be zero, one, or many. A generated query should never select a UUID field. You aren't checking that the answer is right, only that it isn't obviously wrong.

The cases come from three places: synthetic inputs you ask a model to generate, real traces from production, and bugs. A reported bug isn't fixed until it's a test.

Level two is human review of traces. Someone reads the full input and output at every stage of the pipeline, on one screen, and labels what they see: good or bad, a score, or a correction. The corrections matter most, because each one becomes a new eval.

You can run an LLM judge alongside the human, and the interesting cases are the ones where the two disagree. Once the judge has been tuned on enough of those, a human needs to look less often, though never zero.

Level three is A/B testing with real users. We haven't got there yet. Real people are more creative than any test set we can write, but it takes weeks to produce a signal.

The part that surprised me most was what happens when you chain stages together.

Three stages at 95% each is not 95% end to end. It is about 86%.

Accuracy compounds, which sounds like an argument against splitting the pipeline up. It isn't. Stages don't buy you accuracy, they buy you the ability to find out where the accuracy went. One giant prompt is unmeasurable: when the output is wrong you cannot tell whether extraction failed, or parsing, or normalisation, and the only moves left are adding another line to the prompt or upgrading the model.

So we made every stage emit its intermediate output and evaluated each one separately. Then a failure has an address. Fixing the step that locates where in the document a value sits gave us seventeen points end to end, because everything downstream had been inheriting its mistakes.

None of this needs a platform. Pick one feature that has real users, log around fifty traces a day, read them, and turn them into a list of failure modes. The ones a program can check go into CI; the ones it can't get a judge. A spreadsheet is enough to start.

The timeline is longer than anyone wants to hear. Around twenty working days of focused effort before a beta goes near a customer, and talk to stakeholders in iteration cycles rather than ship dates: a hundred traces, the ten or twenty distinct failures inside them, two days to improve those.

Most of the structure here came out of a session my manager Abhimanyu Khosla ran for our team, and a lot of the underlying thinking is from Hamel Husain's evals FAQ, which is the best single document I have found on this.

Next: provenance. Why the output on its own is never enough.

All writing