Model Collapse Starts in Your Own Data Lake
Everyone worries about model collapse as an internet-scale problem: AI slop floods the web, the next generation of foundation models trains on it, and quality decays in a slow civilizational feedback loop. That framing is comforting because it makes collapse someone else's problem — a thing that happens to OpenAI and Anthropic, on a timescale of years, mitigated by armies of data-cleaning PhDs.
Here is the uncomfortable version: the same feedback loop is already running inside your company, and it converges much faster than the internet-scale one. Every logged completion that gets labeled as a "gold example," every model-written document that lands in your RAG corpus, every LLM-judged eval that promotes an LLM-generated answer — each is a small act of training on your own exhaust. You don't need nine generations of recursive pretraining to feel it. In a production system that mines its own logs for few-shot examples and fine-tuning data, generation two ships next quarter.
