Configuration-Driven LLM Fine-Tuning with Axolotl
Most fine-tuning projects start the same way: someone copies a training script, edits a dozen hard-coded values, and hopes the next person can reproduce the run. Axolotl takes a different stance. It treats the entire fine-tuning job as a single YAML file, so the model, the dataset format, the LoRA settings, and the multi-GPU strategy all live in one versionable artifact. This tutorial walks through Axolotl end to end: what it actually wraps, how to install it, how to read and write the config field by field, and how to run a realistic QLoRA fine-tune of an 8B instruct model on one or two GPUs.
What Axolotl Is
Axolotl is not a new training framework. It is an opinionated wrapper that sits on top of the Hugging Face ecosystem and orchestrates the pieces you would otherwise wire together by hand:
- Transformers for model and tokenizer loading.
- PEFT for LoRA and QLoRA adapters.
- TRL for the supervised fine-tuning (SFT) and preference-optimization trainers.
- bitsandbytes for 4-bit and 8-bit quantization.
- Accelerate, DeepSpeed, and FSDP for distributed and multi-GPU training.
The central idea is that you describe what you want rather than writing the glue code that does it. A single YAML config declares the base model, how your dataset should be parsed, which adapter to attach, the optimizer schedule, and the distributed backend. Axolotl reads that file, builds the right objects, and launches the run. The same file is the thing you commit to git, hand to a colleague, or attach to an experiment in Weights & Biases.
When to Choose Axolotl
Axolotl is a good fit when one or more of these is true:
- You have more than one GPU. Axolotl's integration with DeepSpeed ZeRO and FSDP is first-class, and switching from one GPU to eight is mostly a launcher change, not a code rewrite.
- You work across many model families. Llama, Mistral, Mixtral, Qwen, Gemma, Phi, and others are supported through the same config surface, so you are not relearning a new API per model.
- Reproducibility matters. Because the run is fully described by a config file, "rerun exactly what we did last month" becomes a
git checkoutplus one command. - You want recipes, not scripts. Axolotl ships dozens of example configs you can copy and adjust.
If your situation is the opposite — a single consumer GPU where raw training speed and minimal VRAM are the priority — a kernel-optimized library such as Unsloth may train faster. Axolotl's strength is breadth, distributed scaling, and reproducibility rather than squeezing the last token-per-second out of one card. The two are not competitors so much as tools for different jobs.
Installation
Axolotl depends on a CUDA-enabled PyTorch build, and the most common source of trouble is a mismatch between PyTorch, the CUDA toolkit, and flash-attn. Install PyTorch first, matched to your driver, then install Axolotl.
# 1. Create an isolated environment
python -m venv .venv
source .venv/bin/activate
2. Install a CUDA-matched PyTorch (example: CUDA 12.1)
pip install torch==2.3.1 --index-url https://download.pytorch.org/whl/cu121
3. Install Axolotl with the flash-attention and deepspeed extras
pip install "axolotl[flash-attn,deepspeed]"
If your environment is fragile or you simply want something that works on the first try, the official Docker image bundles a known-good combination of CUDA, PyTorch, and the optimized kernels:
docker run --gpus all --rm -it \
-v "$(pwd)":/workspace \
axolotlai/axolotl:main-latest \
bash
After installation, verify the CLI is on your path:
axolotl --help
The Role of Accelerate
Axolotl uses Hugging Face Accelerate to abstract the device placement and distributed launch. For a single GPU you usually do not need to touch it, but for multi-GPU runs Accelerate decides how processes are spawned. You can generate a default config once with accelerate config, but in practice most Axolotl users let the DeepSpeed or FSDP settings in the YAML drive the behavior and launch with accelerate launch. We return to this in the training section.
A Sensible Project Layout
Because the config is the artifact, it helps to keep a tidy, predictable directory structure so a run is easy to reproduce and review:
my-finetune/
├── configs/
│ └── qlora-8b.yml # the run definition
├── data/
│ └── supportinstructions.jsonl
├── outputs/ # adapters and checkpoints (gitignored)
├── deepspeedconfigs/ # copied from the Axolotl examples
│ └── zero2.json
└── requirements.txt # pinned versions for reproducibility
Commit configs/, data/ (or a pointer to it), and requirements.txt; gitignore outputs/. With that convention, anyone can clone the repository and rerun the exact experiment from the config alone.
The YAML Config Is the Central Artifact
Everything below revolves around one file. Create qlora-8b.yml and we will build it up, then explain each block. Here is the complete config for our running example — a QLoRA fine-tune of an 8B instruct model on a custom instruction dataset.
# qlora-8b.yml
--- Base model ---
basemodel: meta-llama/Meta-Llama-3.1-8B-Instruct
modeltype: LlamaForCausalLM
tokenizertype: AutoTokenizer
--- Quantization ---
loadin4bit: true
loadin8bit: false
--- Adapter ---
adapter: qlora
lorar: 32
loraalpha: 16
loradropout: 0.05
loratargetlinear: true
Or target specific modules instead of loratargetlinear:
loratargetmodules:
- qproj
- kproj
- vproj
- oproj
--- Datasets ---
datasets:
- path: ./data/supportinstructions.jsonl
type: chattemplate
chattemplate: llama3
fieldmessages: messages
chattemplate: llama3
valsetsize: 0.05
--- Sequence handling ---
sequencelen: 4096
samplepacking: true
padtosequencelen: true
--- Batch and schedule ---
microbatchsize: 2
gradientaccumulationsteps: 8
numepochs: 3
optimizer: adamwbnb8bit
lrscheduler: cosine
learningrate: 0.0002
warmupsteps: 20
--- Precision and memory ---
bf16: auto
gradientcheckpointing: true
flashattention: true
--- Logging and checkpoints ---
wandbproject: support-finetune
wandbname: llama31-8b-qlora-run1
outputdir: ./outputs/llama31-8b-qlora
savesteps: 100
evalsteps: 100
loggingsteps: 5
Base Model, modeltype, and tokenizertype
basemodel: meta-llama/Meta-Llama-3.1-8B-Instruct
model
type: LlamaForCausalLM
tokenizertype: AutoTokenizer
basemodel is a Hugging Face repo ID or a local path. modeltype names the Transformers architecture class; for most popular models Axolotl can infer it, but stating it explicitly avoids ambiguity. tokenizertype is almost always AutoTokenizer, which lets Transformers pick the correct fast tokenizer for the model.
Quantization: loadin4bit and loadin8bit
loadin4bit: true
loadin8bit: false
These flags control bitsandbytes quantization of the base weights. loadin4bit: true is the foundation of QLoRA — it loads the frozen base model in 4-bit, dramatically cutting VRAM so an 8B model trains comfortably on a single 24 GB card. loadin8bit is a middle ground: less aggressive compression, slightly higher quality, more memory. For full fine-tuning you set both to false.
Adapter: lora vs qlora
adapter: qlora
lorar: 32
loraalpha: 16
loradropout: 0.05
loratargetlinear: true
adapter selects the training method. qlora means "LoRA adapters on top of a 4-bit base"; lora means LoRA on a full-precision (or bf16) base; omitting adapter entirely means full fine-tuning of every parameter.
The LoRA hyperparameters control the low-rank update:
loraris the rank — the size of the injected low-rank matrices. Higher rank means more trainable capacity and more memory. 16–64 is a sensible range for instruction tuning.loraalphascales the update. A common heuristic is to set it equal to or roughly half oflorar, then tune.loradropoutregularizes the adapter; 0.05 is a safe default.loratargetlinear: truetells Axolotl to attach LoRA to all linear projection layers automatically. This is the easiest correct choice. If you prefer manual control, comment it out and listloratargetmodulesexplicitly (for exampleqproj,kproj,vproj,oproj, plus the MLP projections).
Datasets and the type Field
datasets:
- path: ./data/supportinstructions.jsonl
type: chattemplate
chat
template: llama3
fieldmessages: messages
chattemplate: llama3
valsetsize: 0.05
This is the block people get wrong most often, so it deserves attention. datasets is a list — you can mix several sources. Each entry has a path and a type. The type tells Axolotl how to parse each row and turn it into training tokens. The common types are:
alpaca— rows withinstruction, optionalinput, andoutputfields. Axolotl applies the classic Alpaca prompt template.chattemplate(the modern replacement for the oldersharegpttype) — rows that contain a list of role/content messages. You pointfieldmessagesat the list field and setchattemplateto the model's template (herellama3). Axolotl then renders the conversation using that template, including the correct special tokens and turn boundaries.completion— raw text continuation, no instruction structure. Useful for domain-adaptive pretraining on plain documents.
The top-level chattemplate: llama3 sets the template used for rendering and, importantly, for inference later. valsetsize: 0.05 carves a 5% validation split from the training data automatically; alternatively you can supply a separate validation dataset.
Sequence Length and Sample Packing
sequencelen: 4096
sample
packing: true
padtosequencelen: true
sequencelen is the maximum token length per example; sequences longer than this are truncated. samplepacking: true is one of Axolotl's most useful features — it concatenates multiple short examples into one full-length sequence so the GPU is not wasting compute on padding. With proper attention masking this does not let examples bleed into each other, and it can multiply throughput on datasets with many short rows. padtosequencelen pads packed batches to a uniform length, which keeps CUDA graphs and flash-attention happy.
Batch Size, Gradient Accumulation, and Schedule
microbatchsize: 2
gradientaccumulationsteps: 8
numepochs: 3
optimizer: adamwbnb8bit
lrscheduler: cosine
learningrate: 0.0002
warmupsteps: 20
The effective batch size is microbatchsize × gradientaccumulationsteps × numberofgpus. Here on a single GPU that is 2 × 8 × 1 = 16. Keep microbatchsize as large as your VRAM allows and make up the rest with accumulation. optimizer: adamwbnb8bit is the 8-bit Adam from bitsandbytes, which cuts optimizer-state memory substantially and pairs naturally with QLoRA. lrscheduler: cosine with a learning rate around 2e-4 is a solid starting point for LoRA; full fine-tuning uses much smaller rates (often 1e-5 to 2e-5). warmupsteps ramps the learning rate up to avoid an unstable first few steps.
Precision, Checkpointing, and Attention
bf16: auto
gradientcheckpointing: true
flashattention: true
bf16: auto uses bfloat16 mixed precision when the hardware supports it (Ampere and newer). gradientcheckpointing: true trades a bit of compute for a large memory saving by recomputing activations during the backward pass — usually essential for fitting longer sequences. flashattention: true enables the memory-efficient fused attention kernel; it speeds things up and lowers memory, and is the main reason flash-attn appears in the install extras.
Logging, Checkpoints, and Evaluation
wandbproject: support-finetune
wandb
name: llama31-8b-qlora-run1
outputdir: ./outputs/llama31-8b-qlora
savesteps: 100
evalsteps: 100
loggingsteps: 5
The wandb* fields connect the run to Weights & Biases; set WANDBAPIKEY in your environment, or remove these fields to disable tracking. outputdir is where adapters and checkpoints land. savesteps and evalsteps control how often checkpoints are written and validation is run; matching them is convenient so every saved checkpoint has an eval score. loggingsteps controls how often metrics print to the console.
Preparing Your Data
Axolotl reads JSONL where each line is one example. For our chattemplate dataset, each row holds a messages list, exactly like a chat completion request:
import json
examples = [
{
"messages": [
{"role": "system", "content": "You are a concise support assistant for an e-commerce platform."},
{"role": "user", "content": "How do I track my order?"},
{"role": "assistant", "content": "Open Orders, pick the order, and tap Track Shipment to see live status."},
]
},
{
"messages": [
{"role": "user", "content": "Can I change my delivery address after checkout?"},
{"role": "assistant", "content": "Yes, within 30 minutes of placing the order. Go to Orders > Edit Address."},
]
},
]
with open("data/supportinstructions.jsonl", "w", encoding="utf-8") as f:
for ex in examples:
f.write(json.dumps(ex, ensureascii=False) + "\n")
If your data is in the older Alpaca shape instead, each line would look like {"instruction": "...", "input": "...", "output": "..."} and you would set type: alpaca. The principle is the same: the JSONL field names must match what the chosen type expects.
Preprocessing
Tokenization and packing can be slow, so Axolotl separates it into a preprocess step that caches the result. Running it once before training avoids re-tokenizing on every launch and catches data problems early:
axolotl preprocess qlora-8b.yml
This validates the config, renders every example through the chat template, packs sequences, and writes a tokenized dataset to a cache directory. If your chat template is wrong or a field name is mismatched, you find out here rather than ten minutes into a multi-GPU job.
Running Training
On a single GPU, training is one command:
axolotl train qlora-8b.yml
Axolotl loads the cached dataset, builds the quantized model with QLoRA adapters, and runs the TRL trainer according to your schedule. Checkpoints and the final adapter appear under outputdir.
Overriding Config Values on the Command Line
You do not have to edit the YAML for every small change. Any field can be overridden on the command line, which is handy for sweeps or quick experiments without touching the committed config:
# Try a different learning rate and a shorter run without editing the file
axolotl train qlora-8b.yml \
--learning-rate 0.0001 \
--num-epochs 1 \
--output-dir ./outputs/lr-1e4-sweep
The override leaves your base config untouched, so the committed file stays the canonical recipe while you explore variations. For systematic sweeps, prefer generating one config file per variant so each run remains fully self-describing.
Multi-GPU with Accelerate and DeepSpeed
For two or more GPUs, launch through Accelerate so processes are spawned correctly, and point Axolotl at a DeepSpeed ZeRO config. Axolotl ships ready-made ZeRO configs:
accelerate launch -m axolotl.cli.train qlora-8b.yml \
--deepspeed deepspeedconfigs/zero2.json
Equivalently, you can put the DeepSpeed path directly in the YAML:
deepspeed: deepspeedconfigs/zero2.json
DeepSpeed ZeRO shards optimizer state and gradients across GPUs to reduce per-device memory. ZeRO stage 2 shards optimizer state and gradients and is the common default for LoRA/QLoRA on a couple of GPUs. ZeRO stage 3 additionally shards the model parameters themselves and is what you reach for when a model does not fit on one device even quantized — at the cost of more inter-GPU communication.
As an alternative to DeepSpeed, Axolotl also supports FSDP (PyTorch Fully Sharded Data Parallel) through an fsdp block in the YAML. FSDP and ZeRO-3 solve the same problem with different implementations; for most LoRA workloads ZeRO-2 is simpler and sufficient, and you only move to ZeRO-3 or FSDP when memory forces you to shard parameters.
Reading the Training Output
While the run proceeds, Axolotl prints metrics every loggingsteps and runs validation every evalsteps. Two numbers matter most:
- Training loss should fall steadily and then flatten. A loss that drops to near zero within a fraction of an epoch usually means the dataset is too small or repetitive, and the model is memorizing rather than generalizing.
- Validation loss is the honest signal. If it bottoms out and then starts rising while training loss keeps falling, you are overfitting — stop at the checkpoint with the lowest validation loss rather than the last one.
Because valsetsize: 0.05 produced a held-out split, every saved checkpoint has an associated eval score. When you point inference or merge at a directory, prefer the checkpoint with the best validation loss, not necessarily the final step. If you set wandbproject, these curves are plotted automatically; otherwise watch the console logs.
Inference and Chatting with the Adapter
Once training finishes, you can talk to the fine-tuned model directly through the same config — Axolotl loads the base model plus your trained adapter:
axolotl inference qlora-8b.yml --lora-model-dir="./outputs/llama31-8b-qlora"
This drops you into an interactive prompt using the chat template you trained with, which is the quickest way to sanity-check that the model learned the behavior you intended before doing any formal evaluation.
Merging the Adapter and Exporting
A LoRA/QLoRA run produces a small adapter, not a full model. For deployment you often want a single merged model — the base weights with the adapter folded in — so downstream serving (vLLM, TGI, llama.cpp) does not need to know about adapters:
axolotl merge-lora qlora-8b.yml --lora-model-dir="./outputs/llama31-8b-qlora"
This writes a merged model directory (by default under outputdir) containing standard safetensors weights and the tokenizer. From there you can push it to the Hugging Face Hub or convert it to GGUF for local inference. One caution: merging a QLoRA adapter dequantizes the base to compute the merge, so the merged model is full or half precision, not 4-bit — quantize it separately afterward if you need a 4-bit deployment artifact.
Sizing the Hardware
Choosing the method and the GPU count is mostly a memory question. The rough guidance below is for an 8B model with sequencelen: 4096, gradient checkpointing on, and flash-attention enabled; treat it as a starting point, not a guarantee, since exact usage depends on rank, packing, and batch size.
- QLoRA, single 24 GB GPU (e.g. RTX 4090 / A10): comfortable with
microbatchsize: 2and accumulation. The most common real-world setup for an 8B fine-tune. - LoRA in bf16, single 48 GB GPU (e.g. A6000): the base in 16-bit plus adapters fits, with headroom for a larger micro batch.
- Full fine-tuning, 2–8 × 40/80 GB GPUs (A100/H100): needs ZeRO-3 or FSDP to shard optimizer state and parameters; an 8B full fine-tune is impractical on a single 24 GB card.
When memory is tight, the levers in order of impact are usually: enable 4-bit (QLoRA), lower sequencelen, lower microbatchsize, ensure gradient checkpointing and flash-attention are on, then shard with ZeRO-3.
Full Fine-Tuning vs LoRA vs QLoRA
Axolotl supports all three with the same config surface, so the choice is about resources and goals rather than tooling:
- Full fine-tuning updates every parameter. It gives the most capacity to change the model but needs the most VRAM (often multiple high-memory GPUs for an 8B model) and risks catastrophic forgetting if your dataset is narrow. Set
adapterto nothing,loadin4bit: false, and use a small learning rate. This is usually combined with ZeRO-3 or FSDP. - LoRA freezes the base in bf16 and trains small low-rank adapters. Far cheaper than full fine-tuning, with quality close to it for most instruction-tuning tasks. Use it when you have enough VRAM to hold the base in 16-bit.
- QLoRA is LoRA on a 4-bit base. It is the most memory-efficient option and the right default for single-GPU or modest multi-GPU setups. The 4-bit quantization introduces a small quality cost that is negligible for most fine-tuning, which is why our example uses it.
Preference Optimization: DPO, ORPO, and Reward Modeling
Axolotl is not limited to supervised fine-tuning. Reinforcement-learning-style alignment methods are configured through an rl: field, again reusing the same YAML structure and TRL under the hood:
rl: dpo
datasets:
- path: ./data/preferences.jsonl
type: chatml.intel # a prompt/chosen/rejected format
Setting rl: dpo switches the trainer to Direct Preference Optimization, which learns from pairs of preferred and rejected responses. rl: orpo selects ORPO, which combines the SFT and preference objectives in a single stage and so does not need a separate reference model. Axolotl also supports KTO and reward-model training through the same mechanism. The dataset format changes to carry prompt, chosen, and rejected fields instead of a single target, but the rest of the config — model, adapter, schedule, distributed backend — looks just like the SFT case.
Best Practices and Common Pitfalls
A few issues account for most failed Axolotl runs:
- Out-of-memory (OOM). If training crashes with a CUDA OOM, reduce
microbatchsizefirst and raisegradientaccumulationstepsto keep the effective batch size constant. Then confirmgradientcheckpointing: trueandflashattention: true. Loweringsequencelenhelps a lot because attention memory grows with sequence length. On multiple GPUs, move from ZeRO-2 to ZeRO-3. - Chat template mismatch. The single most common cause of a model that trains cleanly but behaves oddly at inference is using a different chat template (or special tokens) for training and serving. Match the
chattemplatein your config to the model family, and use the same template when you deploy. If you fine-tune an instruct model, do not invent your own prompt format. - Sample packing surprises.
samplepackingis excellent for throughput, but it requires correct attention masking, which depends on flash-attention being active. If you see examples appearing to influence each other, verify flash-attention is enabled and your model supports packed attention. For very small datasets the throughput gain is marginal and you can leave it off. - Preprocess before you scale. Always run
axolotl preprocesson a single process before launching an expensive multi-GPU job. It surfaces data and template errors cheaply. - Pin your versions. Because Axolotl sits on a fast-moving stack (Transformers, PEFT, TRL, flash-attn), record the exact versions alongside your YAML, or use the Docker image, so a reproducible config stays reproducible months later.
- Watch the effective batch size when scaling GPUs. Adding GPUs multiplies the effective batch size. If you go from one to four GPUs without adjusting accumulation, your effective batch quadruples and your learning-rate schedule no longer matches — retune or rescale.
Conclusion and Key Takeaways
Axolotl reframes fine-tuning as configuration rather than code. By pushing the base model, dataset parsing, adapter choice, optimization schedule, and distributed strategy into one YAML file, it makes runs reproducible, easy to share, and straightforward to scale from one GPU to many.
The points worth carrying forward:
- The YAML config is the artifact — version it, review it, and treat it as the source of truth for a run.
- Choose Axolotl for breadth, multi-GPU scaling, and reproducibility; reach for a kernel-optimized library when single-GPU speed on one model is the only goal.
- QLoRA (
adapter: qlora+loadin4bit: true) is the practical default; step up to LoRA or full fine-tuning only when you have the memory and a reason. - Preprocess first, then train, then merge — and keep the chat template consistent between training and serving.
- Scaling to multiple GPUs is mostly a launcher and DeepSpeed/FSDP decision, not a code change.
Start from one of the shipped example configs, adapt it to your dataset, run preprocess, and iterate. The discipline of keeping everything in the config pays off the first time someone has to reproduce your work.