Process Reward Models for Multi-Step Reasoning
Step-level reward models catch flawed reasoning that correct final answers hide.
Picking your best checkpoint means watching for reward hacking, not just picking the highest score.
Step-level reward models catch flawed reasoning that correct final answers hide.
Choosing the right training objective shapes how well reward models learn from human feedback.
Small biases in human feedback compound into verbose AI through reinforcement learning loops.
Reward models collapse on out-of-distribution tasks, leaving alignment systems vulnerable to gaming.
Code's executability enables richer reward signals than human judgment can provide.
Using multiple reward models to quantify and act on training uncertainty.
Small, high-quality preference datasets outperform larger, noisier ones in training reward models.
DPO works on paper but falters on cyclic preferences and unbounded reward scaling.
A self-correcting loop proves less important than the principles guiding it.
DPO trains faster and simpler, but PPO explores better on hard tasks.
Setting this hyperparameter wrong dooms alignment to either inertia or reward hacking.