The Meeting That Should Have Been an Eval
Every team shipping an LLM feature eventually holds the same meeting. Someone asks whether the new model — or the new prompt, or the new retrieval tweak — is good enough to ship. The senior engineer who spent the weekend testing says it feels sharper. The PM says a customer complained about exactly this last week. The skeptic on the team pulls up a transcript where the old version was clearly better. Forty-five minutes later, nobody has changed their mind, and the decision gets made by whoever talks last or outranks the room.
That meeting is a symptom. It recurs because the team is trying to settle an empirical question — did this change make the system better or worse? — with anecdotes, and anecdotes don't converge. You can stack them all day. The reason the debate never ends is that there is no shared instrument that everyone agrees to be bound by. The meeting that should have been an eval is the meeting where you discover you're missing one.
