The Eval That Converges, Then Quietly Collapses
A plateaued eval score does not always mean a model ceiling. When labelers homogenize, agreement metrics climb and the eval stops measuring what the team thinks it does.
evals
llm-judge
labeling
ai-engineering
+1