Rishi Patidar

Overfitting by hand

Five versions of a document pipeline, and no way to tell whether any of them was better.

We needed to read PDFs. Somebody uploads a document, and the values inside it should end up in the system as structured fields, without a person opening the file and typing them in one at a time.

The first version took an afternoon. We pulled the text out of the PDF, sent the whole thing to the model, and asked for structured output back. It worked on the first document we tried. It worked on the second. By the fourth we had stopped checking carefully and started talking about what to build next.

Then we fed it a document with a table in it, and the numbers came back attached to the wrong fields.

We assumed the model was the problem. We wrote a firmer prompt. We tried a better model. Same mess. Eventually somebody thought to print out the text we were actually sending, and there it was: the table was gone.

Not missing, exactly. Flattened. Every cell in one long run, one after another, with nothing left to say which column any of it came from. A number that meant one thing sat beside a number that meant something else entirely, and nothing in the string told the model which was which.

So it had not misread the document. It had read something that was no longer the document. That is going to be true of any pipeline like this: whether you pull text with OCR or a parsing library, layout is the first thing to go, and layout is carrying half the meaning.

We fixed the extraction so the structure survived. That got us further, and then we hit the next thing.

One prompt was still doing four jobs at once: find the fields, pull the values, normalise the formats, and reconcile anything that showed up twice. On short documents it held. On long ones accuracy fell off across every field at the same time, as if the model ran out of attention partway down the page.

So we split it into steps. Carefully, because splitting is not free. Every boundary you draw is a place where context goes missing, and a step that only sees one section cannot know that something three pages later contradicts it. Cut it badly and you have swapped one vague answer for five confident ones that disagree. What worked was splitting along the seams the document already had, rather than along the fields we wanted out of it, and adding one step at a time until each earned its place.

That worked. Faster, more accurate, and this time we were careful: we ran it against every file we had, and they all came back clean. So we shipped it.

Then someone brought their own PDF, and it broke.

Not on anything exotic. Just a document we had never seen. It took us longer than it should have to say the obvious thing out loud: every file we had tested against was a file we had picked ourselves.

So we did what you do. We fixed it for that document. Someone brought another one, and we fixed it for that. Better prompts, different intermediate steps, new personas, new ways of dividing up the work.

Five versions in, the pattern was hard to miss. Every version fixed the document in front of us and broke on the next unfamiliar one. And if you had asked me whether version five was better than version two, I could not have told you. I had no way to know.

That is not iteration. That is overfitting, by hand, one file at a time.

The uncomfortable part is how easy it is to end up here. Coding harnesses let you stand something up astonishingly fast, and what you get works in the demo, works on the files you happened to have open, and falls over on a stranger's document. We built five of those and called it progress.

What we were missing was never a better prompt. It was a way to tell. Not an opinion about whether the latest output looked better, but a number that moved when we improved something and stayed put when we hadn't.

That is all an eval is. A set of documents you did not choose, the answers you expect from each one, and a score you can compare between versions. It is slow to build, nobody enjoys it, and it is the only thing that turns changing prompts into engineering.

We should have built it before the first version.

Next: what an eval actually is, and how we built ours.

All writing