Catastrophic Forgetting in RL Fine-Tuned Language Models
Reward optimization in RL fine-tuning systematically erases previously learned capabilities.
Cassiel Oduya
Senior Writer
Cassiel Oduya covers post-training pipelines, reward model engineering and rlhf foundations for The Reward Signal.
6 stories
Reward optimization in RL fine-tuning systematically erases previously learned capabilities.
Choosing the right training objective shapes how well reward models learn from human feedback.
Using multiple reward models to quantify and act on training uncertainty.
DPO works on paper but falters on cyclic preferences and unbounded reward scaling.
Setting this hyperparameter wrong dooms alignment to either inertia or reward hacking.
Proxies become liabilities once you optimize hard against them.