Experiment Tracking for RLHF Runs With Weights and Biases
Track RLHF experiments across three stages to catch failures before they compound downstream.

Reinforcement Learning from Human Feedback is three training runs, not one. It is three: supervised fine-tuning, reward model training, and PPO optimization, each with its own hyperparameters, its own checkpoints, and its own way of failing. A tracking discipline built only to record final outputs misses the point of the pipeline entirely, because the pipeline's risk lives in the connections between stages, not in any single stage's endpoint. The SFT checkpoint is the reference point the PPO KL penalty is anchored to; if that checkpoint's training is poorly recorded, there's no way to reason later about why the KL divergence behaves the way it does. Reinforcement learning algorithms carry a long history of being hard to debug, and the N+ Implementation Details paper identifies subtle implementation choices, the kind that rarely make it into a methods section, as a central obstacle to reproducing results at all. Model size, seed, learning rate, and KL weight can all shift between runs, and a change introduced at the first stage can ripple forward without leaving any visible trace unless someone logged it at the source.
A practitioner might reasonably object that ad hoc logging, print statements, spreadsheets, a folder of checkpoints named by date, has always been good enough for research code, and that a dedicated tracker just adds overhead to a pipeline that is already slow and expensive to run. The logs were the audit trail other researchers could actually inspect. That choice reflects something specific to deep reinforcement learning: reproducibility in this setting is sensitive to random seed, reward scale, and even the exact state of the codebase, and none of those sensitivities can be compared across runs using print-statement logging that nobody structured for comparison in the first place.
What the three RLHF stages produce and risk
Each RLHF stage carries a failure mode that stays invisible unless the right signal is captured as the run happens. Supervised fine-tuning takes the pre-trained base model and adapts it to a target format using high-quality demonstration data, producing what becomes the reference policy checkpoint for everything downstream.
Reward model training takes human pairwise comparisons and turns them into a scalar scoring function that can be applied to any prompt-response pair.
PPO is where the policy generates responses, the reward model scores them, and the policy's weights get adjusted to increase expected reward, all constrained by a KL divergence penalty measured against the SFT reference. The central risk at this stage is reward hacking: the policy discovers and exploits some imperfection in the reward model's scoring and starts producing responses that score well without being good responses. KL divergence spiking is the early warning here: a spike signals the policy moving rapidly away from SFT behavior, which precedes and enables hacking.
What to log during SFT
Treating the SFT run as a disposable warm-up is a mistake, because it is the reference policy that the PPO KL penalty is defined against, and its checkpoint must be versioned and linked explicitly.
Practically, this starts with initializing a run and logging training loss, validation loss, and, for summarization-style tasks, ROUGE scores at each epoch. The checkpoint itself should be saved as a versioned artifact rather than a loose file, because the stages that come next need to pull the exact SFT weights that produced a given result, not whatever happens to be the most recent checkpoint. Learning rate is therefore a config value that needs to be logged at SFT and then carried through, unchanged and traceable, into the reward model and PPO configs.
Dataset specification matters just as much as the optimizer settings. Tokenization choices, the distribution of token lengths in the training set, and any filtering applied to the demonstration data are the kind of implementation detail the N+ paper singles out as materially affecting outcomes, and all of it belongs in the log alongside the loss curves. Seed tracking also needs to begin here, not later: the N+ paper ran four random seeds per model size across several Pythia checkpoints, and logging seed as a config field at the SFT stage is what makes it possible to pair a given downstream PPO run back to the exact SFT run it descended from.
What to log during reward model training
The reward model's validation accuracy is not a number that matters only within Stage 2. It sets the upper bound that PPO's mean reward should not exceed, and logging it explicitly is what makes reward hacking something you can catch before it does damage, rather than something you discover after the fact.
Validation accuracy on held-out preference pairs should be logged at every checkpoint, because that figure becomes the reference ceiling used later when interpreting the PPO dashboard. Training loss matters too, along with the margin between the scores assigned to preferred and rejected responses. a margin that starts collapsing is an early sign that the reward model is overfitting to its comparison set rather than learning a generalizable notion of quality. The reward model checkpoint itself should be saved as a versioned artifact and explicitly linked to the SFT run it was initialized from, so that lineage between stages is something the tracker states outright rather than something someone has to reconstruct from file-system paths and guesswork.
Architecture (typically a language model fitted with a scalar output head), batch size, learning rate, and the number of comparison pairs used all belong in the log as well, since these are the variables that determine how much trust a given reward model actually deserves. The N+ paper reports a correlation between higher reward model validation accuracy and higher win rates for the final PPO policy, which turns reward model checkpoint selection into a decision worth logging and comparing across seeds rather than a formality applied to whichever checkpoint finished training last.
What to log during PPO training
The PPO dashboard only tells a coherent story when mean reward, KL divergence, entropy, and clip fraction are read together as a system, rather than checked one at a time in isolation. Enabling W&B logging through TRL's PPOConfig, using the report_to="wandb" argument, gets the PPOTrainer to automatically record the metrics that the RLHF literature has identified as the crucial ones to watch. objective/non_score_reward captures the mean reward from non-score sources, including the KL penalty contribution. objective/rlhf_reward, the mean RLHF reward described as the "ultimate objective," should keep rising if training is working.
policy/clipfrac_avg tracks the fraction of policy updates that PPO's clipping mechanism has to intervene on, and a persistently high clip fraction points toward a learning rate or batch size that needs adjustment.
Reading these signals as a system means recognizing two distinct failure shapes. Reward hacking is visible in objective/scores continuing to climb past the ceiling set by the reward model's own validation accuracy, logged back in Stage 2, even as the actual output quality visibly degrades, a clear sign the policy has found a way to exploit the reward model's imperfections rather than earn its score honestly. A runaway policy shows up differently: objective/kl spikes upward as the policy sprints away from the behavior the SFT reference established, and the standard response is to increase beta, the KL weight in PPOConfig, to pull the policy back toward its anchor. Batch size, mini-batch size, learning rate, beta, and the number of PPO epochs all belong in the logged config, since these are the settings a practitioner needs in front of them when deciding how to respond to either failure pattern. The N+ paper's practice of sharing W&B run links, with each PPO run explicitly linked to the specific reward model and SFT runs it was paired with across four seeds, turned its published logs into a navigable audit trail rather than a folder of disconnected CSV files.
Using W&B Sweeps and artifact lineage to manage multi-seed, multi-size RLHF experiments
A single RLHF run, however complicated, is debuggable by a patient practitioner working through its logs. A grid spanning multiple seeds and multiple model sizes is a different problem: it only becomes comparable if artifact lineage is tracked explicitly and the sweep's results are organized into a shared workspace from the moment the runs launch.
The N+ paper's own structure illustrates why this matters. Running four random seeds across several model sizes produces a large number of PPO runs, each one paired to a specific reward model run and a specific SFT run, and at that scale, explicit cross-run linkage is a requirement for making sense of the results at all. Artifact lineage handles this by having each PPO run declare the specific input artifacts it consumed, the SFT checkpoint and the reward model checkpoint, so that a failed run can be traced back to a specific SFT seed without anyone relying on file-naming conventions to reconstruct what happened.
W&B Sweeps manage the multi-seed structure directly: seed gets defined as a sweep parameter, every other hyperparameter stays fixed (following the N+ paper's own approach of holding learning rate constant across the whole pipeline), and the sweep controller handles launching and tracking every resulting run inside one shared project. Shared dashboards built through W&B Reports let a team compare objective/rlhf_reward and objective/kl across seeds side by side, making it immediately visible if one seed is diverging from the others while the rest remain healthy. Failure cases belong in the record, not deleted from it: the N+ paper explicitly logged failed runs alongside successful ones for analysis, a discipline worth adopting broadly, since every run, successful or not, is placed in the same project by default.
How DPO changes the tracking picture
DPO trains like SFT with no sampling loop, so there is no reward model to overfit and no PPO loop to tune, and fewer opportunities exist for anything resembling reward hacking to emerge in the first place. DPO is genuinely more stable, cheaper to run computationally, and easier to debug, and it is a legitimate choice for teams working with well-defined preference judgments over a narrow task distribution.
The tracking burden for DPO looks much closer to standard SFT: training loss, validation loss, and evaluation metrics cover most of what needs logging, with none of PPO's objective/kl, clip fraction, or mean reward signal to account for. That simplicity comes at a specific cost. DPO removes the real-time training signal that a PPO dashboard provides, so if a model is quietly overfitting or generalizing poorly during training, no in-training metric flags it, and the failure only becomes visible once evaluation happens, after training has already finished.
That tradeoff is the whole argument. DPO's simplicity is genuine, but it trades tracking complexity for an offline opacity that cannot be reduced by better logging, because there is no equivalent in-training signal to log. When reward hacking is the specific failure a team is worried about, PPO's real-time dashboard is the advantage the added complexity buys, not merely the cost it imposes. GRPO and verifier-based reinforcement learning approaches represent a third path now entering the same tooling: TRL's trainers cover GRPO with built-in logging support, extending the same tracking discipline to a new family of methods without requiring a different integration for each one.
What an infrastructure acquisition means for RLHF practitioners
For a practitioner running SFT, reward model training, and PPO across multiple seeds and model sizes, the practical throughline does not change: the tracking discipline that the N+ paper modeled, versioned checkpoints, explicit artifact lineage, shared run links treated as reproducibility artifacts in their own right, remains the same discipline worth building into a pipeline from the first SFT run onward.


