Fine-Tuning LLMs Efficiently with Unsloth
Fine-tuning large language models used to demand expensive multi-GPU servers and hours of patience. Unsloth changes that equation by rewriting the hot paths of the training loop, letting you fine-tune models like Llama, Mistral, Qwen, Gemma, and Phi roughly twice as fast and with substantially less VRAM. In this tutorial we walk through a complete, realistic workflow: installing Unsloth, loading a 4-bit model, attaching LoRA adapters, preparing a dataset, training with Hugging Face's SFTTrainer, running inference, and exporting the result for production use.
What Is Unsloth and Why It Is Faster
Unsloth is an open-source library that accelerates the supervised fine-tuning (SFT) of transformer language models. It is not a new training algorithm and it does not change the math of your model. Instead, it replaces several performance-critical operations with hand-written implementations that do the same work more efficiently.
The speedups come from a few concrete engineering decisions:
- Custom Triton kernels. Operations such as RoPE embeddings, RMSNorm, cross-entropy loss, and the LoRA matrix multiplications are reimplemented as fused Triton kernels. Fusing several steps into one kernel avoids repeated reads and writes to GPU memory, which is usually the real bottleneck.
- Manual autograd. Rather than relying entirely on PyTorch's automatic differentiation graph, Unsloth provides hand-derived backward passes for the operations it owns. This removes intermediate tensors that PyTorch would otherwise keep around, lowering peak memory.
- Memory-aware design. Unsloth integrates 4-bit quantization (QLoRA) and a custom gradient checkpointing strategy so that long sequences and larger models fit on a single consumer GPU.
A point worth stressing: these are exact reimplementations, not approximations. Unsloth does not trade accuracy for speed. The loss curve you get should match a standard Hugging Face plus PEFT run, just produced faster and with a smaller memory footprint.
What Unsloth Supports
Unsloth targets LoRA and QLoRA fine-tuning of the popular open-weight families:
- Llama (3, 3.1, 3.2 and derivatives)
- Mistral and Mixtral
- Qwen (2 and 2.5)
- Gemma (1 and 2)
- Phi (3 and 3.5)
It works on a single NVIDIA GPU, including the free T4 offered in Google Colab. Full fine-tuning of every parameter is generally outside its sweet spot; the library is built around parameter-efficient methods.
Installation
For most recent setups, a single pip command is enough:
pip install unsloth
Unsloth depends on a CUDA-enabled PyTorch build, transformers, trl, peft, accelerate, and bitsandbytes. If you are on a clean machine, install a PyTorch build matching your CUDA version first, then install Unsloth:
# Example for CUDA 12.1 (adjust for your environment)
pip install torch --index-url https://download.pytorch.org/whl/cu121
pip install unsloth
To pull the latest fixes directly from the repository:
pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"
You can confirm the install and inspect your GPU with:
import torch
from unsloth import FastLanguageModel
print("CUDA available:", torch.cuda.isavailable())
print("Device:", torch.cuda.getdevicename(0) if torch.cuda.isavailable() else "CPU")
Loading a 4-Bit Model
The entry point for almost everything is FastLanguageModel.frompretrained. Loading in 4-bit (QLoRA) is what keeps VRAM low. Here we load a small instruct model that fits comfortably on a single GPU.
from unsloth import FastLanguageModel
import torch
maxseqlength = 2048 # Context length you intend to train on
dtype = None # None lets Unsloth auto-detect (bf16 on Ampere+, fp16 otherwise)
loadin4bit = True # 4-bit QLoRA; set False for 16-bit LoRA if you have the VRAM
model, tokenizer = FastLanguageModel.frompretrained(
modelname="unsloth/llama-3.2-3b-instruct-bnb-4bit",
maxseqlength=maxseqlength,
dtype=dtype,
loadin4bit=loadin4bit,
)
A few notes on these arguments:
maxseqlengthcontrols how long your training examples can be. Unsloth supports long contexts through its kernels, but longer sequences cost more memory, so pick the smallest value that covers your data.- The
unsloth/...-bnb-4bitmodel names are pre-quantized mirrors hosted by the Unsloth team. They download faster and skip an on-the-fly quantization step. You can also pass any standard Hugging Face model ID. - Leaving
dtype=Noneis the safe default; Unsloth selects bfloat16 on GPUs that support it.
Adding LoRA Adapters
We do not update all of the model's weights. Instead we attach small trainable LoRA matrices and freeze the rest. getpeftmodel wires these in.
model = FastLanguageModel.getpeftmodel(
model,
r=16, # LoRA rank: higher = more capacity, more memory
target
modules=[
"qproj", "kproj", "vproj", "oproj",
"gateproj", "upproj", "downproj",
],
loraalpha=16, # Scaling factor; a common choice is alpha == r
loradropout=0, # 0 is optimized and works well in practice
bias="none", # "none" is the optimized, recommended setting
usegradientcheckpointing="unsloth", # See the memory tips below
randomstate=3407,
userslora=False, # Rank-stabilized LoRA, optional
loftqconfig=None,
)
Key choices explained:
r(rank) is the main capacity knob. Values of 8, 16, 32, and 64 are common. For most instruction-following tasks 16 is a sound starting point.targetmoduleslists which linear layers receive adapters. Including both the attention projections and the MLP projections (as above) is the typical, well-performing configuration.usegradientcheckpointing="unsloth"activates Unsloth's own checkpointing implementation, which is more memory-efficient than the generic PyTorch version and is the single most important flag for fitting long sequences.
Preparing the Dataset
Suppose we want a model that answers questions about our company's data-consulting services. We will use a small custom Q&A dataset in the Alpaca instruction format, then convert it with the model's chat template so the special tokens are correct.
Alpaca-Style Formatting
from datasets import loaddataset
A tiny illustrative dataset. In practice load your own JSON/CSV/Hub dataset.
raw = loaddataset("yahma/alpaca-cleaned", split="train[:2000]")
alpacaprompt = """Below is an instruction that describes a task, paired with an input that provides further context. Write a response that appropriately completes the request.
Instruction:
{}
Input:
{}
Response:
{}"""
EOSTOKEN = tokenizer.eostoken # Critical: without EOS the model never learns to stop
def formatalpaca(examples):
texts = []
for instruction, inp, output in zip(
examples["instruction"], examples["input"], examples["output"]
):
text = alpacaprompt.format(instruction, inp, output) + EOSTOKEN
texts.append(text)
return {"text": texts}
dataset = raw.map(formatalpaca, batched=True)
The most common dataset bug is forgetting to append EOSTOKEN. Without it the model keeps generating past the answer and never produces a stop token.
Chat-Template Formatting
If your data is conversational, the cleaner path is to use a ShareGPT-style structure and apply the model's chat template. Unsloth ships helpers for this.
from unsloth.chattemplates import getchattemplate
tokenizer = getchattemplate(
tokenizer,
chattemplate="llama-3.1", # Match the template to your base model
)
def formatchat(examples):
convos = examples["conversations"] # list of {"role", "content"} turns
texts = [
tokenizer.applychattemplate(c, tokenize=False, addgenerationprompt=False)
for c in convos
]
return {"text": texts}
dataset = chatdataset.map(formatchat, batched=True)
For data that arrives as separate instruction/output columns, tosharegpt can merge them into the conversational structure before you apply the template. The principle is the same: produce a single text field per example that contains the fully formatted, tokenizer-correct string.
Training with SFTTrainer
Unsloth integrates directly with the Hugging Face TRL SFTTrainer. There is nothing exotic here; you pass the Unsloth model and the formatted dataset to the standard trainer.
from trl import SFTTrainer
from transformers import TrainingArguments
from unsloth import isbfloat16supported
trainer = SFTTrainer(
model=model,
tokenizer=tokenizer,
traindataset=dataset,
datasettextfield="text",
maxseqlength=maxseqlength,
datasetnumproc=2,
packing=False, # Set True to pack short sequences for extra throughput
args=TrainingArguments(
perdevicetrainbatchsize=2,
gradientaccumulationsteps=4, # Effective batch size = 2 4 = 8
warmupsteps=5,
maxsteps=60, # Use numtrainepochs for full runs
learningrate=2e-4,
fp16=not isbfloat16supported(),
bf16=isbfloat16supported(),
loggingsteps=1,
optim="adamw8bit", # 8-bit optimizer saves memory
weightdecay=0.01,
lrschedulertype="linear",
seed=3407,
outputdir="outputs",
reportto="none",
),
)
trainerstats = trainer.train()
Notes on the configuration:
gradientaccumulationstepslets you simulate a larger batch size without the memory cost. The effective batch size isperdevicetrainbatchsize gradientaccumulationsteps.optim="adamw8bit"uses the bitsandbytes 8-bit optimizer, which roughly halves optimizer-state memory with negligible quality impact.- Use
maxstepsfor quick experiments and switch tonumtrainepochsonce you are training for real. - A learning rate of
2e-4is a reasonable default for LoRA; lower it to1e-4if the loss is unstable.
Monitoring Memory
Unsloth prints a memory and speed summary during training. You can also check it manually:
gpustats = torch.cuda.getdeviceproperties(0)
startmem = round(torch.cuda.maxmemoryreserved() / 1024 / 1024 / 1024, 2)
print(f"GPU: {gpustats.name}, total memory: {round(gpustats.totalmemory/1024/1024/1024,2)} GB")
print(f"Reserved memory before/after training: {startmem} GB")
Memory and VRAM Tips
Getting a run to fit on a modest GPU is mostly about a handful of levers:
- Use 4-bit (
loadin4bit=True). This is the largest single saving and is the default reason to reach for Unsloth. - Set
usegradientcheckpointing="unsloth". This trades a little compute for a large reduction in activation memory and outperforms the generic checkpointing. - Keep
maxseqlengthtight. Memory scales with sequence length. If 90 percent of your examples are under 1024 tokens, do not set the limit to 4096. - Lower the batch size, raise gradient accumulation. This keeps the effective batch size constant while reducing peak memory.
- Prefer
adamw8bit. The 8-bit optimizer noticeably reduces optimizer-state memory. - Enable
packing=Truefor many short examples. Packing concatenates short sequences to fill the context window, improving throughput, though it changes loss accounting slightly.
On a free Colab T4 (16 GB) you can comfortably fine-tune 3B-class models in 4-bit and even some 7B/8B models with a conservative maxseqlength and batch size.
Running Inference
After training, switch the model into Unsloth's optimized inference mode, which enables a native fast generation path.
FastLanguageModel.forinference(model) # ~2x faster generation
messages = [
{"role": "user", "content": "What services does a data consulting firm typically offer?"},
]
inputs = tokenizer.apply
chattemplate(
messages,
tokenize=True,
add
generationprompt=True,
return
tensors="pt",
).to("cuda")
outputs = model.generate(
inputids=inputs,
maxnewtokens=256,
usecache=True,
temperature=0.7,
topp=0.9,
)
print(tokenizer.batchdecode(outputs, skipspecialtokens=True)[0])
For streaming output token by token, use a TextStreamer:
from transformers import TextStreamer
streamer = TextStreamer(tokenizer, skipprompt=True)
= model.generate(inputids=inputs, streamer=streamer, maxnewtokens=256)
Saving and Exporting the Model
Unsloth supports three practical save targets depending on how you intend to deploy.
LoRA Adapters Only
The smallest artifact. It stores just the trained adapter weights and must be loaded on top of the base model later.
model.savepretrained("loramodel")
tokenizer.save
pretrained("loramodel")
To push to the Hugging Face Hub:
model.push
tohub("your-username/loramodel", token="hf...")
Merged 16-Bit Model
Merges the LoRA weights back into the base model and saves a standalone 16-bit checkpoint. This is the format most serving stacks (vLLM, TGI, plain Transformers) expect.
model.savepretrainedmerged(
"merged
16bitmodel",
tokenizer,
save
method="merged16bit",
)
You can also save a merged 4-bit model with savemethod="merged4bit" when disk and download size matter more than precision.
GGUF for llama.cpp and Ollama
To run the model in llama.cpp, Ollama, or LM Studio, export to GGUF with the quantization of your choice.
# Single quantization (q4km is a strong size/quality balance)
model.save
pretrainedgguf("ggufmodel", tokenizer, quantizationmethod="q4km")
Multiple quantizations at once
model.save
pretrainedgguf(
"gguf
model",
tokenizer,
quantizationmethod=["q4km", "q80", "f16"],
)
Once exported, you can register the GGUF file with Ollama using a simple Modelfile:
# Modelfile
FROM ./ggufmodel/unsloth.Q4KM.gguf
Then build and run:
ollama create my-finetune -f Modelfile
ollama run my-finetune
Using Unsloth on Free Colab
Unsloth is popular partly because it runs on the free Colab T4. The workflow is identical to the one above. Two practical reminders:
- Pin versions at the top of the notebook so a runtime restart does not pull an incompatible dependency.
- Save your adapters to Google Drive or the Hub before the session times out; free Colab sessions are not persistent.
from google.colab import drive
drive.mount("/content/drive")
model.savepretrained("/content/drive/MyDrive/loramodel")
Best Practices
- Start small and iterate. Run 60 steps first to confirm the pipeline works end to end before launching a multi-hour job.
- Match the chat template to the base model. A Llama template on a Qwen model produces malformed prompts and poor results.
- Always append the EOS token in Alpaca-style formatting so the model learns to stop.
- Track the loss curve. A healthy LoRA run shows a steadily decreasing loss; a flat or exploding loss usually means a learning rate or formatting problem.
- Validate qualitatively. Quantitative loss is not enough; sample a handful of prompts and read the outputs before deploying.
- Keep a small held-out set to detect overfitting, especially with small datasets and high rank.
Common Pitfalls
- Forgetting
EOSTOKEN. The single most frequent mistake; the model rambles indefinitely at inference time. - Too high a learning rate. LoRA tolerates higher rates than full fine-tuning, but
1e-3and above often destabilize training. Stay near2e-4. - Oversized
maxseqlength. Setting it far above your actual data length wastes memory and can cause out-of-memory errors for no benefit. - Mismatched dtype or CUDA. A PyTorch build that does not match your CUDA driver is the usual cause of install-time failures.
- Expecting full fine-tuning behavior. Unsloth is optimized for LoRA/QLoRA. If you need to update every parameter on a large cluster, it is not the right tool.
- Skipping inference mode. Forgetting
FastLanguageModel.forinference(model)leaves generation on the slower training path.
Conclusion and Key Takeaways
Unsloth makes LLM fine-tuning accessible on hardware that was previously too small for the job, without asking you to compromise on output quality. By reimplementing the expensive parts of the training loop as fused Triton kernels with hand-written backward passes, it delivers roughly 2x faster training and meaningfully lower VRAM while producing the same results as a standard PEFT run.
The essentials to remember:
- Unsloth accelerates LoRA and QLoRA fine-tuning for Llama, Mistral, Qwen, Gemma, and Phi families on a single GPU.
- Load a 4-bit model with
FastLanguageModel.frompretrained, attach adapters withgetpeftmodel, and train with the familiar Hugging FaceSFTTrainer. - The biggest memory wins come from 4-bit loading,
usegradientcheckpointing="unsloth", a tightmaxseqlength, and the 8-bit optimizer. - Save what fits your deployment: LoRA adapters for portability, a merged 16-bit model for standard serving, or GGUF for llama.cpp and Ollama.
- It runs on free Colab, which makes it an excellent way to learn fine-tuning before committing to larger infrastructure.
With this workflow you can take a small instruct model, adapt it to a custom Q&A dataset, and ship a deployable artifact in a single afternoon.