Preference Dataset Quality Metrics for RLHF
Small, high-quality preference datasets outperform larger, noisier ones in training reward models.
Anjali Raghunathan
Section
7 stories in RLHF Foundations.
Small, high-quality preference datasets outperform larger, noisier ones in training reward models.
DPO works on paper but falters on cyclic preferences and unbounded reward scaling.
A self-correcting loop proves less important than the principles guiding it.
DPO trains faster and simpler, but PPO explores better on hard tasks.
Setting this hyperparameter wrong dooms alignment to either inertia or reward hacking.
The model's four hidden assumptions about human judgment don't hold up in production reward models.
Proxies become liabilities once you optimize hard against them.