Your Eval Suite Is Overfit to Your Incumbent Model
A new frontier model ships. It's cheaper, faster, and tops every public leaderboard. You run it through your eval suite — the one you've spent eighteen months building — and it scores worse than the model you're running today. So you keep the incumbent, file the challenger under "not ready," and move on.
Here's the uncomfortable part: that result tells you almost nothing about which model is better. It tells you that your eval suite was built by watching your current model fail, one production incident at a time, and then patched to make those specific failures go away. The suite isn't a neutral measurement of quality. It's a catalog of one model's scar tissue. And a challenger that has different weaknesses will always look worse against a test set assembled from the incumbent's particular weaknesses — even when it's better on the traffic you actually serve.
This is incumbent bias, and it's the switching cost nobody prices into the migration decision. It quietly locks you onto a model long after a better option exists, and it does it while wearing the costume of rigorous engineering.
