The Hidden Ops Org Behind Your Evals: Annotator Economics
Your eval suite reports a number. Behind that number is a label. Behind that label is a person — usually one you have never met, working a two-week gig managed over WhatsApp, paid through a mobile money app, ranking model outputs they were given fifteen seconds to read. The eval score you ship to your VP, the regression gate that blocks your deploy, the leaderboard rank you put in the launch blog — all of it inherits the quality of that person's attention on that afternoon.
We talk about evals as if they were instruments: calibrated, repeatable, objective. They are not. An eval is a measurement device whose sensor is a human labor pipeline, and most teams budget for the harness while treating the people as a free input. That accounting error is why eval scores drift, why your "ground truth" disagrees with itself, and why the most expensive bottleneck in frontier AI is no longer compute.
