Eval Selection Bias: Why Your Test Set Goes Blind to the Failures That Drove Users Away
Eval sets refreshed from production traces inherit a survivor bias: the users who hit the worst failures left and stopped generating traces. Scores climb while retention slips. Here is how to break the loop.