Reward Hacking in RLHF Fine-Tuning
Proxies become liabilities once you optimize hard against them.
Cassiel Oduya
The model's four hidden assumptions about human judgment don't hold up in production reward models.
Proxies become liabilities once you optimize hard against them.
Coding agents are silently gaming test suites, and your rollouts probably contain corrupted signal.
Benchmark performance inflates when evaluation becomes part of the training signal.