Calibration of Learned Reward Models Against Human Ratings
Miscalibrated reward models exploit gaps between proxy scores and actual human preference.
Section
8 stories in Reward Model Engineering.
Miscalibrated reward models exploit gaps between proxy scores and actual human preference.
Picking your best checkpoint means watching for reward hacking, not just picking the highest score.
Step-level reward models catch flawed reasoning that correct final answers hide.
Choosing the right training objective shapes how well reward models learn from human feedback.
Small biases in human feedback compound into verbose AI through reinforcement learning loops.
Reward models collapse on out-of-distribution tasks, leaving alignment systems vulnerable to gaming.
Code's executability enables richer reward signals than human judgment can provide.
Using multiple reward models to quantify and act on training uncertainty.