Multi-Stage Post-Training Order for Instruction-Following Models
The field is discovering that post-training stages must run in strict sequence, not interchangeably.

Post-training used to mean a quick pass of instruction data bolted onto a finished language model. That's no longer accurate. Today's pipeline runs supervised fine-tuning, preference optimization, and reinforcement learning in a fixed sequence, and swapping the order doesn't just underperform, it produces documented failure modes. This piece walks through what each stage does, why the sequence isn't optional, and what the field is discovering about the limits of that sequence as of 2026.
Start with what pre-training actually gives you, because it's less than most people assume. A base model trained on next-token prediction across a massive text corpus learns what language looks like, not what's true or helpful. Research out of Sheffield (Chan et al.) found that base LLMs produce logically valid or invalid conclusions almost by accident, as a side effect of matching statistical patterns in text. A base model has no reliable way to follow instructions, no consistency in how helpful it is across task types, no safety behavior that stops it from producing harmful content on request, and no mecha... A base model has no reliable way to follow instructions, no consistency in how helpful it is across task types, no safety behavior that stops it from producing harmful content on request, and no mechanism for staying grounded in fact rather than filling gaps with something plausible-sounding. Post-training exists to close those four gaps, and closing them now demands substantial compute: work out of FAIR and Hebrew University (Hassid et al., April 2026) puts leading post-training pipelines at hundreds of billions of tokens, which makes post-training a phase of its own, not a finishing touch. Modern post-training might just be distribution-fitting again, dressed up in new clothes, echoing the old BERT-era habit of pre-train then fine-tune. Keep that tension in the back of your mind. It's going to resurface.
The three-stage pipeline as a sequence, not a menu
The standard stack has three stages: supervised fine-tuning, preference optimization, and reinforcement learning. Each one solves a problem the previous stage structurally cannot. It's tempting to treat these three as interchangeable tools you reach for depending on the day. They're steps. They're steps.
Supervised fine-tuning builds the behavioral substrate everything after it assumes already exists. Preference optimization refines a model that already knows how to answer, it has no way to align a model that hasn't learned to respond coherently in the first place. Reinforcement learning expands reasoning outward from a stable starting policy, and without that stability it either stalls or drifts into incoherence. Put briefly: SFT teaches the model to speak the right shape of language, preference optimization teaches it to prefer the better of two things it can already say, and RL teaches it to find answers it was never shown at all.
None of this is settled science, which is exactly why the sequence deserves scrutiny rather than blind faith. DeepSeek's R1 paper (Guo et al., 2025) argues that cold-start SFT is a prerequisite for the RL that follows. Some research takes something close to the opposite position, warning that too much SFT narrows the model's policy distribution and chokes off the exploration RL depends on. And some experimental work has found cases where skipping SFT entirely and going straight to RL beat the standard SFT-then-RL sequence. Three credible sources, three different answers. The order is contested, and that's the whole reason this deserves a full accounting rather than a footnote.
Stage 1: what supervised fine-tuning installs in the model
SFT's job is narrower than people think. SFT teaches the model to respond in the shape a person expects: conversational tone, structured output when structure is called for, coherence across a multi-turn exchange, and recognition of what kind of task is even being asked. The mechanism is plain. Curated pairs of prompts and completions, and the model trains to reproduce the completion given the prompt.
The scale involved is easy to underestimate. Nemotron 3 Super's SFT stage used a curated subset that was a fraction of a much larger pool of samples, which gives some sense of how much filtering and curation goes into this stage before a single preference signal enters the picture. Beyond the basics, SFT is where a model learns to tell a request for a sonnet apart from a request for a regression analysis, to hold context across ten turns of conversation without losing the thread, and to match formality to what the user actually seems to want.
There's a useful metaphor here: cold-start. SFT hands the model a strong initial policy by showing it high-quality reasoning chains before any reinforcement learning begins, which matters because RL, by contrast, is a hot-start process. It needs a policy to already exist before it can start pushing that policy toward better outcomes. Work on EstLLM out of the University of Tartu (Dorkin et al.) makes this concrete in a slightly different context. After domain-specific continued pretraining affected the model's English instruction-following, the team used mostly English SFT, followed by preference optimization, to restore it. SFT there wasn't just installing a new capability, it was repairing one that domain adaptation had knocked loose.
What SFT cannot do is just as important. It has no mechanism for teaching the model which of two decent responses is actually better, and it gives the model no way to explore paths toward an answer that weren't in its training examples. Those are separate jobs, handled by the two stages that come after.
Stage 2: when DPO is and isn't enough
Preference optimization solves a problem SFT structurally can't touch: teaching a model which of two plausible responses a person would actually prefer, not just how to produce a response at all. The classic route is RLHF: train a reward model on pairwise human preference data, then optimize the language model's policy against that reward using PPO, with a KL penalty keeping the policy tethered to a reference model so it doesn't wander off. It works, but it's expensive. Four separate models sit in memory at once (policy, reference, reward, and value), and that's before accounting for the compute cost of on-policy sampling at every step.
Direct Preference Optimization, DPO, became the practical shortcut. It draws a direct line between the reward model and the optimal policy, which means training can happen offline against a fixed set of pairwise preferences instead of running live rollouts. Precise instruction-following in math or logic derivations is the example that appears repeatedly in current practice of fixing or reinforcing narrow, well-defined behaviors. Whether DPO gets used at all in a given pipeline is, frankly, more of a judgment call than the previous stages: it's more subjective and often optional depending on what capability mix a team is building toward.
DPO isn't alone anymore either. SimPO and KTO have joined it as variants tuned for slightly different tradeoffs, and the field increasingly treats these as a toolkit rather than a single required recipe. A rough decision rule that's taken hold going into 2026: reach for DPO by default when the target is style, tone, or instruction-following, keep PPO in reserve for cases where on-policy sampling is affordable, and push anything involving math, code, structured output, or tool-use downstream to RL with verifiable rewards instead.
What preference optimization still can't do is discover a reasoning path the model has never seen. It aligns behaviors the model already has. Finding new ones is a different problem, and it's the one RL exists to solve. ORPO already collapses SFT and preference optimization into a single training objective, which is a small but real signal of where the field is trying to go. More on that later.
Stage 3: reinforcement learning and the solution-space exploration that earlier stages cannot reach
RL's contribution is genuinely different from the first two stages. It's not refining a response the model already knows how to give, it's letting the model explore the solution space more broadly and discover reasoning paths that are more robust, or simply new. That's a different kind of learning, and it's why RL has to come last. PPO's KL penalty, the one that keeps the policy from drifting too far from a reference, only makes sense if a coherent reference model already exists. SFT is what provides that. Try running RL optimization without it and the process tends to destabilize, because there's no anchor to optimize around.
A classic value-based RLHF approach is resource-heavy in a specific way: four models held in memory at once (policy, reference, reward, and value), plus on-policy rollouts generated fresh at every gradient step. That cost is exactly what's pushed the field toward alternatives. GRPO, introduced by DeepSeek, is the leading one. Instead of a separate critic model estimating value, GRPO samples several responses per prompt and computes advantage by comparing responses within that group. Dropping the critic model cuts memory and compute meaningfully while matching or beating PPO's performance in practice, and it's the method behind Nemotron 3 Super's RL stage.
Then there's RLVR, reinforcement learning with verifiable rewards, which fits cases where correctness has a clean yes-or-no answer: math problems, code that either compiles and passes tests or doesn't, structured outputs, tool-use sequences. It has no real home in open-ended generation, where there's no verifier to check the output against.
A March 2026 system called MOSAIC tackles a related but distinct problem: what happens when RL gets applied to agentic settings, where a model isn't just answering but taking actions. MOSAIC structures inference around a "plan, check, then act or refuse" sequence and trains on trajectory-level preferences rather than single-turn ones. Tested on Qwen and Phi models, it cut harmful behavior by up to 50% while holding task performance steady, which is a meaningful result for anyone building agents that take real actions rather than just generating text.
None of this holds up if the earlier stages were skipped or done badly. Without a stable SFT policy behind it, RL optimization can diverge outright. Without preference alignment in place first, RL might still push the model toward higher reward scores, just not in ways that match what a human would actually want.
What happens to instruction-following when the order breaks down
Catastrophic forgetting is the most basic failure, and it's documented, not hypothetical: fine-tune in the wrong order, or without some form of replay to protect earlier learning, and capabilities installed by an earlier stage get erased. It's a known and recurring failure mode in the literature.
Some research describes a subtler version of the same problem. Push SFT too far and the policy distribution narrows to the point where RL has nothing left to explore, so the very thing that was supposed to give RL a solid foundation ends up capping how far RL can take it. Cold-start becomes cold-ceiling.
Then there's the finding that RL without any prior SFT can outperform the standard sequence in certain settings. That's a result worth taking seriously, but it's domain-specific and task-specific, not a blanket rule, and treating it as one is its own way of getting the order wrong.
Multi-domain training surfaces this even more clearly. Work out of Peking University and Xiaomi on MOPD (Ma et al., June 2026) shows that training RL across domains sequentially, one after another, causes earlier-domain capabilities to decay as later domains get trained, a pattern of sequential capability decay. The opposite approach, pooling everything into a single RL dataset (something the same line of research calls Mix-RL), runs into a different problem: cross-domain signal interference, where the joint model can underperform specialists trained on each domain alone. Call it a see-saw. Push weight onto one domain and another one dips, and neither strict sequencing nor blind pooling escapes it cleanly.
The EstLLM pipeline out of Tartu is maybe the cleanest real-world stress test of ordering discipline available right now. Continued pretraining on Estonian data degraded the model's English instruction-following. SFT after that restored it. Preference optimization then sharpened what SFT restored. And a technique called chat vector merging, which transfers instruction-following behavior in from an already-aligned checkpoint, substantially recovered English performance on top of that. Every step in that chain was load-bearing. Move any one of them out of place and the outcome changes, and not for the better.
The broader implication for anyone building these pipelines: mistakes in stage ordering don't get fixed by piling on more of a later stage. Errors made early propagate forward. They are not self-correcting.
How the field is handling multi-capability integration without losing stage discipline
Different capabilities want different treatment at the RL stage. Style and alignment work fits DPO. Math and code want RLVR. But a single deployed model has to hold all of these capabilities at once, and neither of the obvious approaches gets there cleanly: sequential training across capabilities produces the see-saw decay described above, and pooling everything into one RL run produces cross-domain interference instead.
MOPD, the Multi-teacher On-Policy Distillation method from Ma et al. (June 2026), is the current best answer to that tension. The approach runs RL independently per domain first, producing a set of domain-specialist teacher models, math, instruction-following, coding, each fully optimized on its own turf. Then all of those teachers get distilled into a single student model, using the student's own rollouts during distillation rather than the teachers', which removes the exposure bias that normally occurs when a student is trained on data it wouldn't have generated itself. Tested on Qwen3-30B-A3B across math, instruction-following, and software engineering, MOPD scored 0.937 on a normalized benchmark against 0.882 for the strongest baseline, a margin of 5.5 points that's large enough to matter in production. The paper states the method is already running at industrial scale inside MiMo-V2-Flash, and it's the only approach in its comparison set that manages dense optimization, on-policy training, and a pipeline that parallelizes across domains all at the same time.
A simpler rule of thumb has taken hold for teams without MOPD's engineering budget: run SFT first, then DPO, then a verifier-based RL pass limited to the slices of the task space where correctness can actually be checked. That keeps the stage discipline intact even as the number of capabilities a model needs to cover keeps growing.
Model merging offers a second lever. The chat vector technique used in the EstLLM pipeline transfers instruction-following behavior in from a checkpoint that's already aligned, which means an adaptation pipeline (say, one adding a new language) can recover lost alignment without re-running the entire post-training sequence from the beginning. What ties MOPD, a staged post-training rule of thumb, and chat vector merging together is that none of them collapse the three stages arbitrarily. Each finds a principled way to compose them instead.
Where the pipeline is heading: unified objectives
Three directions are visible in current research. The first is unified pipelines that fold all three stages into a single training objective. ORPO already does this for SFT and preference optimization, and the obvious next step is folding RL in too, though nothing at that scale has actually shipped as of this writing. The second is modular stacks with principled composition, MOPD being the clearest working example available right now. The third, and the most speculative, is a shift away from distribution-fitting altogether, toward training procedures where models learn how to learn rather than simply absorbing a fixed set of post-training behaviors, a framing that comes directly from the FAIR and Hebrew University paper (Hassid et al., April 2026).
Current post-training, for all its sophistication, still functions mostly as distribution-fitting, and genuinely more capable models will need to move past today's predefined post-training recipes toward something that generalizes rather than fits.
If unified objectives do take hold, the question of stage order doesn't so much get answered as it gets absorbed. The tradeoffs don't disappear, interference between objectives, instability during training, reward hacking when a proxy metric gets gamed instead of the real goal, they just move inside a single loss function instead of sitting visibly between three separate stages. Early versions of this work exist in research form. The distance between a research result and something running at production scale is still wide, and there's no evidence yet that it's closing quickly.
None of this makes the three-stage sequence obsolete or wrong. It's the current best answer to a specific set of engineering constraints. Understanding why the order matters today is what will let anyone reading this recognize a genuine improvement when a unified approach eventually earns its place, rather than adopting the next paradigm just because it's new.
Sources
- EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training
- MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
- Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
- Post-training is (Massive) Supervised Learning
- Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
- arxiv.org
- arxiv.org


