Full fine-tuning updates most or all model weights and has the highest compute and storage cost. LoRA keeps the base model frozen and trains small low-rank adapter matrices, drastically reducing the number of trainable parameters. QLoRA goes further by keeping the base model quantized — commonly at 4-bit — while training LoRA adapters. For many organization-specific behavior changes, LoRA or QLoRA is the practical first option; use RAG instead when the main need is frequently changing factual knowledge.
Fine-tune only when the target behavior is trainable
Fine-tuning is useful when you want consistent behavioral changes: a specialized response format, domain-specific terminology, tool-selection patterns, classification behavior, style, or repeated task structure. It is less attractive when the main problem is “the model needs today’s policy document.” That is usually a retrieval problem.
Before training, establish a baseline with prompting and RAG. If those approaches solve the task, they are easier to update and debug. Fine-tuning should address a clear gap that examples can teach.
Full fine-tuning
Full fine-tuning updates the pretrained model weights using task-specific data. This can provide maximum adaptation freedom, but it is expensive. Training stores not only weights but gradients, optimizer states and activations, so memory requirements can be several times inference memory.
Full tuning also creates a new checkpoint lifecycle: versioning, evaluation, redistribution rights, security review and rollback. For large models, distributed training infrastructure is often required.
Use it when adapter capacity is insufficient, the organization has strong training infrastructure, and the expected performance gain justifies the operational burden.
LoRA: update small low-rank matrices instead of the full model
Low-Rank Adaptation freezes the original model weights and adds small trainable matrices to selected layers. The Hugging Face PEFT documentation describes LoRA as a common parameter-efficient fine-tuning method because it drastically reduces the number of trainable parameters.
The result is operationally useful: the base model remains unchanged while different LoRA adapters can represent different tasks or customers. Adapter artifacts are much smaller than a complete copy of the base checkpoint.
Important configuration choices include rank, target modules, scaling and dropout. Higher rank increases adaptation capacity but also increases trainable parameters.
QLoRA: fine-tuning on a quantized base model
QLoRA keeps the pretrained model frozen in a 4-bit quantized representation and backpropagates into LoRA adapters. The original paper introduced techniques such as NF4, double quantization and paged optimizers to make large-model adaptation practical with much less GPU memory.
This makes QLoRA attractive for teams that need to adapt larger models but cannot afford full-precision training infrastructure. It is still training: data preparation, optimizer configuration, evaluation and safety testing remain necessary.
Data quality matters more than raw example count
Fine-tuning data should represent the desired behavior, not merely contain lots of domain text. Instruction tuning requires examples that show inputs and high-quality target outputs. Duplicated, contradictory or low-quality examples can teach undesirable behavior.
Separate training, validation and held-out evaluation sets. Include negative and edge cases. If the model will use tools, include realistic tool failures and ambiguous situations rather than only perfect trajectories.
For regulated or sensitive domains, review whether personal or confidential data is actually necessary before using it in training.
Evaluate the adapter and the base model separately
Measure improvement on the target task and regression on important general capabilities. A fine-tune can become better at the specialized task while losing robustness, calibration or instruction-following behavior elsewhere.
Record base model version, adapter version, dataset version, training configuration and evaluation suite. Without those identifiers, reproducing a result or rolling back a deployment becomes difficult.
A practical decision framework
| Need | Preferred first tool | Why |
|---|---|---|
| Current internal documents | RAG | Knowledge changes without retraining |
| Consistent output style/format | LoRA | Behavioral adaptation with small artifacts |
| Adapter training with limited VRAM | QLoRA | Quantized base reduces memory |
| Deep model transformation | Full fine-tuning | Maximum parameter freedom |
Frequently asked questions
What is the difference between LoRA and QLoRA?
LoRA trains low-rank adapters while the base model stays frozen. QLoRA additionally keeps the base model quantized, typically at 4-bit, to reduce memory.
Should I fine-tune a model to learn company documents?
Usually not as the first choice. RAG is better for frequently changing factual knowledge and source traceability.
Can LoRA adapters be swapped at runtime?
Many serving stacks support adapter workflows, but operational details vary. Treat base model and adapter compatibility as versioned deployment metadata.
Does fine-tuning change the model license?
The legal effect depends on the original model terms and the adapter/checkpoint distribution plan. Review the exact license before training or redistributing derivatives.
Apply this knowledge
Move from the concept to a concrete deployment shortlist.
Primary sources and technical references
OpenWeightModels prefers publisher documentation, standards bodies, official repositories and original research papers. The source material remains authoritative where it changes.