LLM Post-Training with TRL: SFT, Reward Modeling, and DPO
After a base language model is pretrained, it still needs to be shaped into something useful and aligned with human expectations. TRL (Transformer Reinforcement Learning) is Hugging Face's library for that entire post-training stage, from supervised fine-tuning through preference optimization. This tutorial walks through the modern post-training pipeline and shows how to use TRL's trainers in practice, with Direct Preference Optimization (DPO) as the centerpiece.
What TRL Is and Where It Sits
TRL is a library focused on the steps that come after pretraining. Unlike a wrapper that mainly simplifies configuration, TRL provides concrete trainer classes for the full alignment stack: supervised fine-tuning, reward modeling, and several preference-optimization and reinforcement-learning algorithms.
It is built directly on the Hugging Face ecosystem:
- Transformers for model and tokenizer loading.
- PEFT for parameter-efficient adapters such as LoRA, so you can train large models on modest hardware.
- Accelerate for distributed and mixed-precision training.
- Datasets for loading and preprocessing data.
Because TRL reuses these components, anything you already know about AutoModelForCausalLM, tokenizers, or LoRA configs carries over directly.
The Modern LLM Post-Training Pipeline
A typical instruction-following or chat model goes through three stages:
TRL covers stages 2 and 3. Most production alignment work today combines an SFT pass followed by a preference-optimization pass.
Installation
# Core stack
pip install trl peft datasets accelerate
Recommended extras
pip install transformers bitsandbytes
pip install wandb # optional experiment tracking
Verify the install and check that a GPU is visible:
import torch
import trl
print("TRL version:", trl.version)
print("CUDA available:", torch.cuda.isavailable())
TRL trainers run on CPU for tiny experiments, but realistic post-training requires a GPU.
Dataset Formats in TRL
TRL standardizes on a few dataset shapes. Getting these right is most of the work.
- Conversational (chat) format for SFT: a
messagescolumn containing a list of{"role", "content"}dictionaries. - Preference format for reward modeling and DPO: columns named
prompt,chosen, andrejected.
The conversational format looks like this:
example = {
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain what a vector database is."},
{"role": "assistant", "content": "A vector database stores embeddings..."},
]
}
The preference format used throughout the rest of this tutorial:
example = {
"prompt": "Explain what a vector database is.",
"chosen": "A vector database stores embeddings and supports similarity search...",
"rejected": "It's just a normal database.",
}
TRL applies the model's chat template automatically when a dataset is conversational, so you rarely format strings by hand.
Preparing Data with the datasets Library
In practice your raw data rarely arrives in the exact shape a trainer expects. The datasets library makes reshaping cheap, and the transformation runs lazily and in parallel.
from datasets import loaddataset
raw = loaddataset("json", datafiles="mypreferences.jsonl", split="train")
def topreference(example):
return {
"prompt": example["question"],
"chosen": example["goodanswer"],
"rejected": example["badanswer"],
}
dataset = raw.map(topreference, removecolumns=raw.columnnames)
Reserve a held-out split for honest evaluation.
splits = dataset.traintestsplit(testsize=0.05, seed=42)
traindataset, evaldataset = splits["train"], splits["test"]
A few habits that pay off:
- Always keep a held-out split. Preference metrics on the training set are misleading; you want eval numbers the model has not seen.
- Inspect a handful of rows by hand. Print five examples and read them. Label noise is far easier to catch by eye than by aggregate statistics.
- Set a seed. Reproducible splits make experiments comparable across runs.
Supervised Fine-Tuning with SFTTrainer
SFT is the foundation: it teaches the base model the response style and instruction-following behavior you want before any preference tuning. In TRL this is SFTTrainer configured with SFTConfig.
from datasets import loaddataset
from trl import SFTConfig, SFTTrainer
dataset = load
dataset("trl-lib/Capybara", split="train")
config = SFTConfig(
outputdir="sft-model",
numtrainepochs=1,
perdevicetrainbatchsize=2,
gradientaccumulationsteps=8,
learningrate=2e-5,
loggingsteps=10,
packing=True, # pack short samples to use the context window efficiently
maxlength=1024,
bf16=True,
)
trainer = SFTTrainer(
model="Qwen/Qwen2.5-0.5B",
args=config,
traindataset=dataset,
)
trainer.train()
trainer.savemodel("sft-model")
A few points worth noting:
- Chat templates. When the dataset has a
messagescolumn,SFTTrainerapplies the tokenizer's chat template, so special tokens and role markers are added correctly. - Packing. Setting
packing=Trueconcatenates multiple short examples into one sequence up tomaxlength. This improves throughput and reduces wasted padding, at the cost of slightly mixing examples within a sequence. - Passing a model name. You can pass a model ID string and TRL loads it for you, or pass an already-loaded
AutoModelForCausalLMinstance.
Combining SFT with LoRA
For larger models you usually do not want full fine-tuning. TRL accepts a peftconfig, and the trainer wraps the model with PEFT internally. LoRA and PEFT are covered in depth in a separate tutorial, so this is intentionally brief.
from peft import LoraConfig
from trl import SFTConfig, SFTTrainer
peftconfig = LoraConfig(
r=16,
loraalpha=32,
loradropout=0.05,
targetmodules="all-linear",
tasktype="CAUSALLM",
)
trainer = SFTTrainer(
model="Qwen/Qwen2.5-0.5B",
args=SFTConfig(outputdir="sft-lora", bf16=True),
traindataset=dataset,
peftconfig=peftconfig,
)
trainer.train()
The same peftconfig argument works with the other TRL trainers, including DPOTrainer.
Reward Modeling with RewardTrainer
Classic RLHF needs a reward model: a model that takes a prompt and a response and outputs a scalar score reflecting how good that response is. You train it on preference data, teaching it to score chosen higher than rejected.
from datasets import loaddataset
from transformers import AutoModelForSequenceClassification, AutoTokenizer
from trl import RewardConfig, RewardTrainer
modelid = "Qwen/Qwen2.5-0.5B"
tokenizer = AutoTokenizer.frompretrained(modelid)
model = AutoModelForSequenceClassification.frompretrained(modelid, numlabels=1)
dataset = loaddataset("trl-lib/ultrafeedbackbinarized", split="train")
config = RewardConfig(
outputdir="reward-model",
perdevicetrainbatchsize=4,
numtrainepochs=1,
learningrate=1e-5,
loggingsteps=10,
maxlength=1024,
bf16=True,
)
trainer = RewardTrainer(
model=model,
args=config,
traindataset=dataset,
processingclass=tokenizer,
)
trainer.train()
The reward model is loaded as a sequence-classification head with a single output (numlabels=1). The trainer optimizes a pairwise loss so that the score for chosen exceeds the score for rejected. This reward model is what PPOTrainer later uses to provide a training signal.
Direct Preference Optimization (DPO)
DPO is the method most teams reach for first when aligning a model to preferences. It is the centerpiece of this tutorial.
What DPO Is and Why It Helps
Traditional RLHF is a multi-stage process: train a reward model, then run reinforcement learning (PPO) where the policy generates responses, the reward model scores them, and the policy is updated. That loop is powerful but operationally heavy. It requires a separate reward model, online generation during training, and careful tuning to stay stable.
DPO reframes the problem. Instead of training a reward model and then optimizing against it with RL, DPO derives a loss that operates directly on preference pairs. It increases the relative log-probability the policy assigns to the chosen response over the rejected one, while a frozen reference model keeps the policy from drifting too far from its starting point.
The practical consequences:
- No separate reward model to train and serve.
- No online generation loop during training, so it is simpler and more stable.
- It runs as a standard supervised-style training job on a preference dataset.
DPO does still rely on a reference model — typically the SFT model you started from — which it loads automatically when you do not pass one explicitly.
Preference Dataset for DPO
DPO expects the prompt / chosen / rejected format. Many public datasets already ship this way:
from datasets import loaddataset
dataset = loaddataset("trl-lib/ultrafeedbackbinarized", split="train")
print(dataset[0].keys()) # dictkeys(['prompt', 'chosen', 'rejected', ...])
If you build your own data, each row pairs one prompt with a preferred and a dispreferred completion. Quality and consistency of these pairs matter more than raw quantity.
A Complete DPO Example
This example aligns a small instruct model to a preference dataset. In practice you would run SFT first and point DPO at that SFT checkpoint; here we use an already-instruct model for brevity.
from datasets import loaddataset
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl import DPOConfig, DPOTrainer
modelid = "Qwen/Qwen2.5-0.5B-Instruct"
model = AutoModelForCausalLM.frompretrained(modelid)
tokenizer = AutoTokenizer.frompretrained(modelid)
traindataset = loaddataset(
"trl-lib/ultrafeedbackbinarized", split="train"
).select(range(2000)) # small slice for a quick run
config = DPOConfig(
outputdir="dpo-model",
perdevicetrainbatchsize=2,
gradientaccumulationsteps=8,
numtrainepochs=1,
learningrate=5e-6,
beta=0.1, # KL strength: how close to stay to the reference
losstype="sigmoid", # the standard DPO loss
loggingsteps=10,
maxlength=1024,
maxpromptlength=512,
bf16=True,
)
trainer = DPOTrainer(
model=model,
refmodel=None, # TRL creates a frozen copy of model as reference
args=config,
traindataset=traindataset,
processingclass=tokenizer,
)
trainer.train()
trainer.savemodel("dpo-model")
Key DPOConfig knobs:
betacontrols how strongly the policy is held to the reference model. Lower values (around 0.05) allow bigger changes; higher values (0.3+) keep the model conservative. A value near 0.1 is a common starting point.losstypeselects the preference loss."sigmoid"is standard DPO; alternatives such as"ipo"and"hinge"change how the objective treats the preference margin and can reduce overfitting on some datasets.refmodel=Nonetells TRL to make a frozen copy of the policy as the reference. If you train with LoRA, TRL can use the base model with adapters disabled as the reference, avoiding a second model copy in memory.
To run DPO with LoRA instead of full fine-tuning, pass a peftconfig exactly as in the SFT example. This is the common setup for larger models.
How to Read DPO Training Logs
DPO emits metrics that tell you whether the run is healthy, beyond the raw loss. The most informative ones:
rewards/accuracy— the fraction of pairs where the model assigns a higher implicit reward tochosenthan torejected. It should climb above 0.5 and trend upward. If it stays near 0.5, the model is not learning the preference.rewards/margins— the average gap between the chosen and rejected implicit rewards. A growing margin indicates the model is separating the two more confidently.rewards/chosenandrewards/rejected— the implicit rewards themselves. Watch the relationship between them, not the absolute values.
A common warning sign is reward accuracy that climbs quickly to near 1.0 within a few hundred steps on a small dataset. That usually means overfitting rather than genuine alignment, and is a cue to reduce epochs or add more data.
A Tour of the Other Preference and RL Trainers
DPO is not the only option. TRL ships several trainers, and the right one depends on your data and constraints.
PPOTrainer — Classic RLHF
PPOTrainer implements the traditional reward-model-plus-RL approach. The policy generates responses, a reward model scores them, and Proximal Policy Optimization updates the policy. It is the most flexible method and can optimize against arbitrary reward signals, but it is the hardest to tune and the most resource-intensive because generation happens inside the training loop. Reach for PPO when you have a trained reward model and need online optimization that DPO cannot express.
GRPOTrainer — Group Relative Policy Optimization
GRPOTrainer is an RL method that removes the need for a separate value/critic network. For each prompt it samples a group of completions, scores them, and uses the group's relative scores to estimate advantages. It pairs naturally with reward functions rather than a learned reward model, which makes it a strong fit for tasks with verifiable rewards — math, code, or anything you can check programmatically.
from datasets import loaddataset
from trl import GRPOConfig, GRPOTrainer
dataset = loaddataset("trl-lib/tldr", split="train")
A reward function receives the generated completions and returns a score per item.
def lengthreward(completions, kwargs):
# Toy example: prefer concise answers under 200 characters.
return [1.0 if len(c) < 200 else 0.0 for c in completions]
trainer = GRPOTrainer(
model="Qwen/Qwen2.5-0.5B-Instruct",
rewardfuncs=lengthreward,
args=GRPOConfig(outputdir="grpo-model", bf16=True, loggingsteps=10),
traindataset=dataset,
)
trainer.train()
A more realistic verifiable reward checks correctness against a ground-truth answer:
def correctnessreward(completions, groundtruth, kwargs):
scores = []
for completion, answer in zip(completions, ground
truth):
scores.append(1.0 if answer.strip() in completion else 0.0)
return scores
Extra dataset columns (like groundtruth) are forwarded to the reward function as keyword arguments, so you can implement any checkable objective.
ORPO and KTO
Two more methods are worth knowing:
- ORPO (Odds Ratio Preference Optimization) folds preference optimization into the SFT step itself, removing the need for a separate reference model and a separate alignment phase.
- KTO (Kahneman-Tversky Optimization) works with simpler binary feedback — a single completion labeled good or bad — rather than paired chosen/rejected data, which is useful when collecting explicit pairs is hard.
Both have corresponding ORPOTrainer/ORPOConfig and KTOTrainer/KTOConfig classes that follow the same usage pattern as the trainers above.
Logging, Evaluation, and the Hub
TRL trainers integrate with standard Hugging Face tooling.
config = DPOConfig(
outputdir="dpo-model",
evalstrategy="steps",
evalsteps=100,
loggingsteps=10,
reportto="wandb", # or "tensorboard"
pushtohub=True,
hubmodelid="your-username/dpo-model",
)
With an evaluation split passed as evaldataset, DPO logs useful metrics such as reward accuracy (how often chosen scores above rejected) and the reward margin. After training you can push the model:
trainer.pushtohub()
Make sure you are authenticated first with huggingface-cli login. The pushed repository includes the weights, tokenizer, and a generated model card.
Scaling with Accelerate and DeepSpeed
For multi-GPU or memory-constrained training, launch with Accelerate rather than python:
accelerate config # one-time interactive setup
accelerate launch traindpo.py
For large models that do not fit on a single GPU, DeepSpeed ZeRO shards optimizer state, gradients, and parameters across devices:
accelerate launch --configfile deepspeedzero3.yaml traindpo.py
Because TRL is built on Accelerate, your training script does not change — only the launch command and config do. TRL's repository ships example Accelerate and DeepSpeed config files you can adapt.
Best Practices and Common Pitfalls
- Do SFT before DPO. Preference optimization adjusts an already-capable model. Running DPO on a base model that has not learned the chat format usually produces poor results. Treat SFT as a prerequisite, not an optional step.
- Tune
betadeliberately. Too low and the model drifts from the reference, degrading general capability while chasing the preference signal. Too high and it barely moves. Start at 0.1 and adjust based on reward accuracy and qualitative checks. - Mind the reference model. DPO needs a reference. With full fine-tuning that means a second frozen copy in memory; with LoRA, disabling adapters provides the reference for free. Plan your memory budget accordingly.
- Prioritize dataset quality. Preference data is noisier than it looks. Inconsistent or contradictory
chosen/rejectedpairs teach the model conflicting signals. A smaller, clean, consistently-labeled dataset beats a large noisy one. - Watch for overfitting. Preference optimization can overfit quickly, especially on small datasets. Keep epochs low (often one is enough), use a held-out eval split, and monitor reward accuracy alongside generation quality rather than trusting a single number.
- Match the chat template. SFT, the reference model, and DPO must use the same chat template and tokenizer. Mismatches silently corrupt the formatting the model expects and quietly hurt results.
- Keep learning rates small. Post-training uses much lower learning rates than pretraining — typically
1e-6to5e-6for DPO. Higher rates tend to destabilize alignment. - Evaluate beyond the metrics. Reward accuracy can look healthy while generations drift toward verbosity or refusal. Always sample real completions on held-out prompts and read them before declaring a run successful.
Conclusion and Key Takeaways
TRL gives you a single, consistent toolkit for the entire LLM post-training pipeline, built on the Transformers, PEFT, and Accelerate stack you already use.
- Post-training has three stages: pretrain, SFT, then preference optimization. TRL covers the last two.
SFTTrainerhandles supervised fine-tuning, applies chat templates, supports packing, and combines cleanly with LoRA viapeftconfig.RewardTrainertrains a scalar reward model on chosen/rejected pairs for classic RLHF.DPOTraineris the practical default for preference alignment: it optimizes directly on preference pairs, avoids a separate reward model and RL loop, and is tuned mainly throughbetaandloss_type.PPOTrainer,GRPOTrainer,ORPO, andKTOcover the rest of the spectrum, from full RLHF to verifiable-reward RL to binary-feedback methods.- Success depends less on the algorithm and more on doing SFT first, curating clean preference data, tuning
beta, and guarding against overfitting.
Start with an SFT pass, move to DPO on a clean preference dataset, evaluate honestly, and only graduate to PPO or GRPO when your task genuinely needs online reinforcement learning.