Unsloth Tutorial: Fast and Memory-Efficient LLM Fine-Tuning

# Fine-Tuning LLM Secara Efisien dengan Unsloth Dahulu, melakukan fine-tuning model bahasa besar membutuhkan server multi-GPU yang mahal dan waktu tunggu berjam-jam. Unsloth mengubah persamaan itu de...

By Ruby Abdullah · · tutorial
UnslothLLMFine-TuningLoRAQLoRAPython

Fine-Tuning LLMs Efficiently with Unsloth

Fine-tuning large language models used to demand expensive multi-GPU servers and hours of patience. Unsloth changes that equation by rewriting the hot paths of the training loop, letting you fine-tune models like Llama, Mistral, Qwen, Gemma, and Phi roughly twice as fast and with substantially less VRAM. In this tutorial we walk through a complete, realistic workflow: installing Unsloth, loading a 4-bit model, attaching LoRA adapters, preparing a dataset, training with Hugging Face's SFTTrainer, running inference, and exporting the result for production use.

What Is Unsloth and Why It Is Faster

Unsloth is an open-source library that accelerates the supervised fine-tuning (SFT) of transformer language models. It is not a new training algorithm and it does not change the math of your model. Instead, it replaces several performance-critical operations with hand-written implementations that do the same work more efficiently.

The speedups come from a few concrete engineering decisions:

  • Custom Triton kernels. Operations such as RoPE embeddings, RMSNorm, cross-entropy loss, and the LoRA matrix multiplications are reimplemented as fused Triton kernels. Fusing several steps into one kernel avoids repeated reads and writes to GPU memory, which is usually the real bottleneck.
  • Manual autograd. Rather than relying entirely on PyTorch's automatic differentiation graph, Unsloth provides hand-derived backward passes for the operations it owns. This removes intermediate tensors that PyTorch would otherwise keep around, lowering peak memory.
  • Memory-aware design. Unsloth integrates 4-bit quantization (QLoRA) and a custom gradient checkpointing strategy so that long sequences and larger models fit on a single consumer GPU.

A point worth stressing: these are exact reimplementations, not approximations. Unsloth does not trade accuracy for speed. The loss curve you get should match a standard Hugging Face plus PEFT run, just produced faster and with a smaller memory footprint.

What Unsloth Supports

Unsloth targets LoRA and QLoRA fine-tuning of the popular open-weight families:

  • Llama (3, 3.1, 3.2 and derivatives)
  • Mistral and Mixtral
  • Qwen (2 and 2.5)
  • Gemma (1 and 2)
  • Phi (3 and 3.5)

It works on a single NVIDIA GPU, including the free T4 offered in Google Colab. Full fine-tuning of every parameter is generally outside its sweet spot; the library is built around parameter-efficient methods.

Installation

For most recent setups, a single pip command is enough:

pip install unsloth

Unsloth depends on a CUDA-enabled PyTorch build, transformers, trl, peft, accelerate, and bitsandbytes. If you are on a clean machine, install a PyTorch build matching your CUDA version first, then install Unsloth:

# Example for CUDA 12.1 (adjust for your environment)

pip install torch --index-url https://download.pytorch.org/whl/cu121

pip install unsloth

To pull the latest fixes directly from the repository:

pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"

You can confirm the install and inspect your GPU with:

import torch

from unsloth import FastLanguageModel

print("CUDA available:", torch.cuda.isavailable())

print("Device:", torch.cuda.getdevicename(0) if torch.cuda.isavailable() else "CPU")

Loading a 4-Bit Model

The entry point for almost everything is FastLanguageModel.frompretrained. Loading in 4-bit (QLoRA) is what keeps VRAM low. Here we load a small instruct model that fits comfortably on a single GPU.

from unsloth import FastLanguageModel

import torch

maxseqlength = 2048 # Context length you intend to train on

dtype = None # None lets Unsloth auto-detect (bf16 on Ampere+, fp16 otherwise)

loadin4bit = True # 4-bit QLoRA; set False for 16-bit LoRA if you have the VRAM

model, tokenizer = FastLanguageModel.frompretrained(

modelname="unsloth/llama-3.2-3b-instruct-bnb-4bit",

maxseqlength=maxseqlength,

dtype=dtype,

loadin4bit=loadin4bit,

)

A few notes on these arguments:

  • maxseqlength controls how long your training examples can be. Unsloth supports long contexts through its kernels, but longer sequences cost more memory, so pick the smallest value that covers your data.
  • The unsloth/...-bnb-4bit model names are pre-quantized mirrors hosted by the Unsloth team. They download faster and skip an on-the-fly quantization step. You can also pass any standard Hugging Face model ID.
  • Leaving dtype=None is the safe default; Unsloth selects bfloat16 on GPUs that support it.

Adding LoRA Adapters

We do not update all of the model's weights. Instead we attach small trainable LoRA matrices and freeze the rest. getpeftmodel wires these in.

model = FastLanguageModel.getpeftmodel(

model,

r=16, # LoRA rank: higher = more capacity, more memory

targetmodules=[

"qproj", "kproj", "vproj", "oproj",

"gateproj", "upproj", "downproj",

],

loraalpha=16, # Scaling factor; a common choice is alpha == r

loradropout=0, # 0 is optimized and works well in practice

bias="none", # "none" is the optimized, recommended setting

usegradientcheckpointing="unsloth", # See the memory tips below

randomstate=3407,

userslora=False, # Rank-stabilized LoRA, optional

loftqconfig=None,

)

Key choices explained:

  • r (rank) is the main capacity knob. Values of 8, 16, 32, and 64 are common. For most instruction-following tasks 16 is a sound starting point.
  • targetmodules lists which linear layers receive adapters. Including both the attention projections and the MLP projections (as above) is the typical, well-performing configuration.
  • usegradientcheckpointing="unsloth" activates Unsloth's own checkpointing implementation, which is more memory-efficient than the generic PyTorch version and is the single most important flag for fitting long sequences.

Preparing the Dataset

Suppose we want a model that answers questions about our company's data-consulting services. We will use a small custom Q&A dataset in the Alpaca instruction format, then convert it with the model's chat template so the special tokens are correct.

Alpaca-Style Formatting

from datasets import loaddataset

A tiny illustrative dataset. In practice load your own JSON/CSV/Hub dataset.

raw = loaddataset("yahma/alpaca-cleaned", split="train[:2000]")

alpacaprompt = """Below is an instruction that describes a task, paired with an input that provides further context. Write a response that appropriately completes the request.

Instruction:

{}

Input:

{}

Response:

{}"""

EOSTOKEN = tokenizer.eostoken # Critical: without EOS the model never learns to stop

def formatalpaca(examples):

texts = []

for instruction, inp, output in zip(

examples["instruction"], examples["input"], examples["output"]

):

text = alpacaprompt.format(instruction, inp, output) + EOSTOKEN

texts.append(text)

return {"text": texts}

dataset = raw.map(formatalpaca, batched=True)

The most common dataset bug is forgetting to append EOSTOKEN. Without it the model keeps generating past the answer and never produces a stop token.

Chat-Template Formatting

If your data is conversational, the cleaner path is to use a ShareGPT-style structure and apply the model's chat template. Unsloth ships helpers for this.

from unsloth.chattemplates import getchattemplate

tokenizer = getchattemplate(

tokenizer,

chattemplate="llama-3.1", # Match the template to your base model

)

def formatchat(examples):

convos = examples["conversations"] # list of {"role", "content"} turns

texts = [

tokenizer.applychattemplate(c, tokenize=False, addgenerationprompt=False)

for c in convos

]

return {"text": texts}

dataset = chatdataset.map(formatchat, batched=True)

For data that arrives as separate instruction/output columns, tosharegpt can merge them into the conversational structure before you apply the template. The principle is the same: produce a single text field per example that contains the fully formatted, tokenizer-correct string.

Training with SFTTrainer

Unsloth integrates directly with the Hugging Face TRL SFTTrainer. There is nothing exotic here; you pass the Unsloth model and the formatted dataset to the standard trainer.

from trl import SFTTrainer

from transformers import TrainingArguments

from unsloth import isbfloat16supported

trainer = SFTTrainer(

model=model,

tokenizer=tokenizer,

traindataset=dataset,

datasettextfield="text",

maxseqlength=maxseqlength,

datasetnumproc=2,

packing=False, # Set True to pack short sequences for extra throughput

args=TrainingArguments(

perdevicetrainbatchsize=2,

gradientaccumulationsteps=4, # Effective batch size = 2 4 = 8

warmupsteps=5,

maxsteps=60, # Use numtrainepochs for full runs

learningrate=2e-4,

fp16=not isbfloat16supported(),

bf16=isbfloat16supported(),

loggingsteps=1,

optim="adamw8bit", # 8-bit optimizer saves memory

weightdecay=0.01,

lrschedulertype="linear",

seed=3407,

outputdir="outputs",

reportto="none",

),

)

trainerstats = trainer.train()

Notes on the configuration:

  • gradientaccumulationsteps lets you simulate a larger batch size without the memory cost. The effective batch size is perdevicetrainbatchsize gradientaccumulationsteps.
  • optim="adamw8bit" uses the bitsandbytes 8-bit optimizer, which roughly halves optimizer-state memory with negligible quality impact.
  • Use maxsteps for quick experiments and switch to numtrainepochs once you are training for real.
  • A learning rate of 2e-4 is a reasonable default for LoRA; lower it to 1e-4 if the loss is unstable.

Monitoring Memory

Unsloth prints a memory and speed summary during training. You can also check it manually:

gpustats = torch.cuda.getdeviceproperties(0)

startmem = round(torch.cuda.maxmemoryreserved() / 1024 / 1024 / 1024, 2)

print(f"GPU: {gpustats.name}, total memory: {round(gpustats.totalmemory/1024/1024/1024,2)} GB")

print(f"Reserved memory before/after training: {startmem} GB")

Memory and VRAM Tips

Getting a run to fit on a modest GPU is mostly about a handful of levers:

  • Use 4-bit (loadin4bit=True). This is the largest single saving and is the default reason to reach for Unsloth.
  • Set usegradientcheckpointing="unsloth". This trades a little compute for a large reduction in activation memory and outperforms the generic checkpointing.
  • Keep maxseqlength tight. Memory scales with sequence length. If 90 percent of your examples are under 1024 tokens, do not set the limit to 4096.
  • Lower the batch size, raise gradient accumulation. This keeps the effective batch size constant while reducing peak memory.
  • Prefer adamw8bit. The 8-bit optimizer noticeably reduces optimizer-state memory.
  • Enable packing=True for many short examples. Packing concatenates short sequences to fill the context window, improving throughput, though it changes loss accounting slightly.

On a free Colab T4 (16 GB) you can comfortably fine-tune 3B-class models in 4-bit and even some 7B/8B models with a conservative maxseqlength and batch size.

Running Inference

After training, switch the model into Unsloth's optimized inference mode, which enables a native fast generation path.

FastLanguageModel.forinference(model)  # ~2x faster generation

messages = [

{"role": "user", "content": "What services does a data consulting firm typically offer?"},

]

inputs = tokenizer.applychattemplate(

messages,

tokenize=True,

addgenerationprompt=True,

returntensors="pt",

).to("cuda")

outputs = model.generate(

inputids=inputs,

maxnewtokens=256,

usecache=True,

temperature=0.7,

topp=0.9,

)

print(tokenizer.batchdecode(outputs, skipspecialtokens=True)[0])

For streaming output token by token, use a TextStreamer:

from transformers import TextStreamer

streamer = TextStreamer(tokenizer, skipprompt=True)

= model.generate(inputids=inputs, streamer=streamer, maxnewtokens=256)

Saving and Exporting the Model

Unsloth supports three practical save targets depending on how you intend to deploy.

LoRA Adapters Only

The smallest artifact. It stores just the trained adapter weights and must be loaded on top of the base model later.

model.savepretrained("loramodel")

tokenizer.savepretrained("loramodel")

To push to the Hugging Face Hub:

model.pushtohub("your-username/loramodel", token="hf...")

Merged 16-Bit Model

Merges the LoRA weights back into the base model and saves a standalone 16-bit checkpoint. This is the format most serving stacks (vLLM, TGI, plain Transformers) expect.

model.savepretrainedmerged(

"merged16bitmodel",

tokenizer,

savemethod="merged16bit",

)

You can also save a merged 4-bit model with savemethod="merged4bit" when disk and download size matter more than precision.

GGUF for llama.cpp and Ollama

To run the model in llama.cpp, Ollama, or LM Studio, export to GGUF with the quantization of your choice.

# Single quantization (q4km is a strong size/quality balance)

model.savepretrainedgguf("ggufmodel", tokenizer, quantizationmethod="q4km")

Multiple quantizations at once

model.savepretrainedgguf(

"ggufmodel",

tokenizer,

quantizationmethod=["q4km", "q80", "f16"],

)

Once exported, you can register the GGUF file with Ollama using a simple Modelfile:

# Modelfile

FROM ./ggufmodel/unsloth.Q4KM.gguf

Then build and run:

ollama create my-finetune -f Modelfile

ollama run my-finetune

Using Unsloth on Free Colab

Unsloth is popular partly because it runs on the free Colab T4. The workflow is identical to the one above. Two practical reminders:

  • Pin versions at the top of the notebook so a runtime restart does not pull an incompatible dependency.
  • Save your adapters to Google Drive or the Hub before the session times out; free Colab sessions are not persistent.

from google.colab import drive

drive.mount("/content/drive")

model.savepretrained("/content/drive/MyDrive/loramodel")

Best Practices

  • Start small and iterate. Run 60 steps first to confirm the pipeline works end to end before launching a multi-hour job.
  • Match the chat template to the base model. A Llama template on a Qwen model produces malformed prompts and poor results.
  • Always append the EOS token in Alpaca-style formatting so the model learns to stop.
  • Track the loss curve. A healthy LoRA run shows a steadily decreasing loss; a flat or exploding loss usually means a learning rate or formatting problem.
  • Validate qualitatively. Quantitative loss is not enough; sample a handful of prompts and read the outputs before deploying.
  • Keep a small held-out set to detect overfitting, especially with small datasets and high rank.

Common Pitfalls

  • Forgetting EOSTOKEN. The single most frequent mistake; the model rambles indefinitely at inference time.
  • Too high a learning rate. LoRA tolerates higher rates than full fine-tuning, but 1e-3 and above often destabilize training. Stay near 2e-4.
  • Oversized maxseqlength. Setting it far above your actual data length wastes memory and can cause out-of-memory errors for no benefit.
  • Mismatched dtype or CUDA. A PyTorch build that does not match your CUDA driver is the usual cause of install-time failures.
  • Expecting full fine-tuning behavior. Unsloth is optimized for LoRA/QLoRA. If you need to update every parameter on a large cluster, it is not the right tool.
  • Skipping inference mode. Forgetting FastLanguageModel.forinference(model) leaves generation on the slower training path.

Conclusion and Key Takeaways

Unsloth makes LLM fine-tuning accessible on hardware that was previously too small for the job, without asking you to compromise on output quality. By reimplementing the expensive parts of the training loop as fused Triton kernels with hand-written backward passes, it delivers roughly 2x faster training and meaningfully lower VRAM while producing the same results as a standard PEFT run.

The essentials to remember:

  • Unsloth accelerates LoRA and QLoRA fine-tuning for Llama, Mistral, Qwen, Gemma, and Phi families on a single GPU.
  • Load a 4-bit model with FastLanguageModel.frompretrained, attach adapters with getpeftmodel, and train with the familiar Hugging Face SFTTrainer.
  • The biggest memory wins come from 4-bit loading, usegradientcheckpointing="unsloth", a tight maxseqlength, and the 8-bit optimizer.
  • Save what fits your deployment: LoRA adapters for portability, a merged 16-bit model for standard serving, or GGUF for llama.cpp and Ollama.
  • It runs on free Colab, which makes it an excellent way to learn fine-tuning before committing to larger infrastructure.

With this workflow you can take a small instruct model, adapt it to a custom Q&A dataset, and ship a deployable artifact in a single afternoon.

Related Articles

Axolotl Tutorial: Configuration-Driven LLM Fine-Tuning

Fine-Tuning LLM Berbasis Konfigurasi dengan Axolotl Kebanyakan proyek fine-tuning dimulai dengan cara yang sama: seseora...

Fireworks AI: A Super Fast Inference Platform for Open LLMs and Multimodal Models

Fireworks AI: Platform Inference Super Cepat buat Open LLM dan Model Multimodal Halo temen-temen, di tutorial kali ini a...

Together AI: A Complete Guide to Inference and Fine-Tuning Open Source Models with One API

Together AI: Panduan Lengkap Inference dan Fine-Tuning Model Open Source dengan Satu API Halo temen-temen! Kali ini aku ...

TRL Tutorial: LLM Post-Training with SFT, DPO, and Reward Modeling

Post-Training LLM dengan TRL: SFT, Reward Modeling, dan DPO Setelah sebuah base language model selesai dipretraining, mo...