All articles

The Eval That Never Finishes

Why confidence, not capability, is what deep AI products compete on. A first-person look at two-sided drift, rented models, and the measurement you actually own.

Sep 28, 20267 min readProduct strategy

Anthony Ludwig

Product leader & founder, Product Manager Hub

Writes on product strategy, AI decision quality, and PM leadership—grounded in real operating experience, not generic AI takes.

Key takeaways

  • A grounded take on the eval that never finishes.
  • Structured for product leaders making AI and strategy calls under real constraints.
  • Read the full essay for frameworks, tradeoffs, and practical next steps.

two kinds of AI products

Some products can lean on the last few generations of models and be fine. The questions are general, the answers don't change much from one year to the next, and every time the model improves the product gets proportionally better for free. A light AI layer over simple data is the right call there. I don't buy the line that every wrapper is shallow. If the job is simple, a thin layer is good design, and a heavy evaluation apparatus would be wasted effort.

Other products have to go deep. Specifics matter, being wrong is costly, and the person asking is going to act on what they hear. I work in an industry where the rules change season to season, so I've spent a lot of time on this second kind.

The test I use is a simple one. If the right answer to the same question might be different next year, I'm building the second kind.

Most PMs I talk to haven't asked which kind they're building. It sounds like a small question, but it shapes everything after it…. how much I measure, what I promise users, and whether a model upgrade is a gift or a risk.

two things moving at once

In a deep product, two things keep changing, and neither one asks my permission.

The first is what users ask. Once people trust the answers, they go deeper. They bring harder questions, odd edge cases, and situations I never wrote a test for. Then they stop asking only what the tool knows and start asking what it can do. That's a different kind of question, and it needs a different kind of answer.

The second is what's true. Rules change. Restrictions tighten or loosen. Products get replaced and guidance gets rewritten. An answer that was right in March can be wrong by June without a single line of my code changing.

Here's what that looks like in practice. A product gets pulled from the market over the winter. In the spring, someone asks the tool which option to use, and it recommends the one that's gone, in the same confident tone it used last year. Nothing broke. The model didn't change. The world did.

I've started calling this two-sided drift. The questions drift, and the truth drifts, and wrong answers live in the gap between them. That gap is invisible unless I'm measuring it, which is the whole reason I keep an eval set at all.

It's also why my eval set never gets finished. It unfolds, one new question and one changed rule at a time.

That changed how I build it. The best new items don't come from a spreadsheet I wrote before launch. They come from real user questions that went badly, the ones where someone said "that's not right" or where I realized I was confidently answering yesterday's question. Every one of those is a free lesson in what my users actually need. I wrote about the other side of this in The Thinking Your Users Stopped Doing, where trust is what changes the questions in the first place.

the model is rented

Every AI product in a given space has access to the same frontier models, and those models improve every few months. That's a free upgrade for everyone, which sounds great until I ask what it did to my product.

An upgrade I can't measure is a change I can't explain. Answers can get sharper in one area and quietly worse in another, and if all I have is a general feeling that the new model is smarter, I won't see the second part until a user finds it for me.

Measurement is what lets me adopt new models safely. I run the new model against real questions. I compare area by area rather than as one blended number. I switch when the evidence says it's better for my users, and I wait when it doesn't. That takes the anxiety out of every release announcement.

annnnnd here's the part most teams don't fully reckon with.

The eval set, the expert answers, and the score history are what compound. None of it ships with the next model. Every competitor gets the new model on the same day I do, but nobody gets my record of what the last six models did to my users' questions.

The model is rented. The measurement is yours.

I made a related argument in The Moat Mirage, that competing on model speed loses. This is the other half of it. Speed isn't a moat, but knowing what each model improvement does for my specific users can be.

confidence is the product

The model builders already show their work. Every launch comes with a table of where the model excels and where it's still catching up, category by category and version by version. They say plainly what they're strong at and what they're not.

Most AI products show users nothing like that. They answer everything in the same tone of certainty, whether the answer is a settled fact or a coin flip.

In a deep product, "can it answer this?" isn't the question users care about. They care about "how sure is it?" Someone acting on an answer needs to know how much weight it can carry.

So I'm moving toward telling users how confident the product is by area, what it's best for, and where to double-check. Not one confidence score, which would just be another average hiding how shallow some areas are. A profile. Strong on fact lookups, less sure on situational answers, careful on judgment calls.

I lean toward under-claiming. A confidence level should only go up when the evidence supports it. A user who finds the tool better than I promised trusts it more, and a user who finds it worse stops asking. The first mistake is cheap and the second one isn't.

I also try to measure against human experts rather than reporting a bare percentage. "Matches our experts" is a claim people understand. "92 percent accurate" invites the reasonable question of 92 percent of what.

The question I dodged for a long time was whether I could back a confidence label with evidence. Putting a label on an answer is easy. Standing behind it is the work, and it's exactly what the eval set exists to do.

how I think about the evolution now

I think about this in four stages. It's the path I'm on, not a prescription.

Answer. Does it work at all? Get the basic thing running and see if anyone finds it useful.

Measure. Accuracy by area and by depth, meaning fact lookups, situational answers, and judgment calls, checked against real experts. This is where I find out what I actually have.

Maintain. The evals move as the questions and the source truth move. I'm not writing them once. I'm keeping them alive.

Disclose. Confidence shown to users, not just reported to leadership.

Most teams I see stop at the first stage, and a few reach the second. The third and fourth are where deep products separate from everything else built on the same model, because those stages are the ones a competitor can't copy by switching to the same API.

If I were starting over, I wouldn't begin with a full evaluation platform. I'd start smaller. I'd tag real questions by area and depth, ask an expert to spot-check a sample, and keep the record. That's the start of the system that compounds.

where this leaves the demo

If you have something that answers well on your laptop and you can't yet say how sure it is, or what happens when the rules change under it, that's the stretch between demo and done I spend my time on. Tell me where you're stuck and we can work out which kind of product you're building and what the smallest useful measurement looks like.

What would your product say if a user asked how sure it was? And would you be comfortable showing them the math?

Good luck friends.

If this was useful

Share the essay, or follow along where I post the shorter takes.

Follow

Go deeper

If this maps to a stall you're in—demo that won't ship, vibe-code without a product, or a team that can't take it—join Access for playbooks and community.

More on the same problems—judgment, shipping, and getting from demo to done.