Your Users Drift While Your Model Stands Still
Six weeks after launch, your quality dashboard starts sagging. Thumbs-down rates creep up, task completion drifts down, and the on-call channel fills with screenshots of bad responses. The team does what teams do: they diff the prompts (unchanged), check the model version (pinned), audit the retrieval index (fresh), and bisect the deploy history (nothing shipped). Everyone concludes the model provider silently degraded the model. The provider, of course, insists nothing changed.
Everyone is looking in the wrong place. Nothing in the system changed. The users did.
Launch-week metrics assume launch-week users. But people adapt to an AI product within weeks, and they adapt in ways that systematically break the assumptions baked into your prompts, your evals, and your launch benchmarks. Your model is frozen. Your users are not. The gap between them is a form of drift that most teams don't instrument for at all — and it produces the most confusing incident pattern in AI engineering: metric decay with no deploy.
