The RLAIF Doom Loop: When Your Cheapest Feedback Signal Quietly Poisons Your Fine-Tune
AI-generated preference labels are 100x cheaper than human ones — and they teach your model to prefer the judge's aesthetic, not your users'.
rlaif
rlhf
fine-tuning
llm-as-judge
+1