TempatGunting

LLM Fine-Tuning with LoRA and QLoRA: Practical Implementation Guide

Full fine-tuning of a 70B parameter language model requires over 560 GB of GPU memory for the model weights, gradients, and optimizer states. Even with gradient checkpointing and mixed precision, you need a cluster of eight A100 80GB GPUs. LoRA and QLoRA make this impractical setup unnecessary by training only a small number of additional parameters while keeping the base model frozen.

The insight behind LoRA is that the weight updates during fine-tuning have low intrinsic rank. Instead of updating the full weight matrix W ∈ ℝ^(d×k), LoRA decomposes the update into two small matrices: ΔW = BA, where B ∈ ℝ^(d×r) and A ∈ ℝ^(r×k), with rank r much smaller than d and k. For a typical attention layer where d = k = 4096 and r = 16, this reduces the trainable parameters from 16.7 million to 131 thousand — a 127x reduction.

LoRA Implementation

The Hugging Face PEFT library provides the standard implementation. Configuration requires three decisions: which layers to target, what rank to use, and what alpha scaling to apply.

from peft import LoraConfig, get_peft_model, TaskType
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B",
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

lora_config = LoraConfig(
    task_type=TaskType.CAUSAL_LM,
    r=16,                          # rank
    lora_alpha=32,                 # scaling factor
    lora_dropout=0.05,
    target_modules=[
        "q_proj", "v_proj",        # attention projections
        "k_proj", "o_proj",        # optional: more attention layers
        "gate_proj", "up_proj",    # optional: MLP layers
        "down_proj"
    ],
    bias="none"
)

model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# trainable params: 83,886,080 || all params: 8,114,770,944 || trainable%: 1.034

The lora_alpha parameter controls the scaling of the LoRA update. The actual scaling applied is alpha / r, so with alpha=32 and r=16, the LoRA contribution is scaled by 2x. Higher alpha values amplify the fine-tuning effect, which helps when the training data distribution differs significantly from the base model's pretraining data. A common starting point is alpha = 2 × r.

QLoRA: 4-Bit Quantized Fine-Tuning

QLoRA extends LoRA by quantizing the base model to 4-bit precision using NormalFloat4 (NF4) quantization. The key insight is that pretrained model weights follow a normal distribution, so a quantization scheme optimized for normal distributions preserves more information than uniform quantization at the same bit width.

from transformers import BitsAndBytesConfig

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True  # quantize the quantization constants
)

model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-70B",
    quantization_config=bnb_config,
    device_map="auto"
)

# Apply LoRA on top of the quantized model
model = get_peft_model(model, lora_config)

Double quantization (bnb_4bit_use_double_quant=True) quantizes the quantization constants themselves, saving an additional 0.37 bits per parameter. For a 70B model, this saves roughly 3 GB of memory — significant when operating near the memory limit of a single GPU.

The memory savings are dramatic. A 70B parameter model requires 140 GB in float16 but only 35 GB in 4-bit quantization. Adding LoRA adapters with rank 16 adds less than 200 MB. The total memory footprint fits on a single A100 80GB GPU with room for activations and optimizer states — a setup that would be impossible with full precision fine-tuning. This makes vector database integration with custom-tuned models practical even on modest hardware.

Training Configuration

Fine-tuning configuration differs from pretraining in several important ways. Learning rates should be 10-100x smaller than pretraining to avoid catastrophic forgetting. Gradient accumulation compensates for smaller batch sizes imposed by memory constraints. Warmup prevents destabilizing the pretrained representations during early training.

from transformers import TrainingArguments
from trl import SFTTrainer

training_args = TrainingArguments(
    output_dir="./output",
    num_train_epochs=3,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,     # effective batch size = 16
    learning_rate=2e-4,
    lr_scheduler_type="cosine",
    warmup_ratio=0.03,
    weight_decay=0.001,
    bf16=True,
    logging_steps=10,
    save_strategy="epoch",
    optim="paged_adamw_8bit",          # memory-efficient optimizer
    max_grad_norm=0.3,
    group_by_length=True,              # minimize padding waste
)

trainer = SFTTrainer(
    model=model,
    train_dataset=dataset,
    args=training_args,
    max_seq_length=2048,
    packing=True,                      # pack multiple examples per sequence
)

The paged_adamw_8bit optimizer uses 8-bit quantized optimizer states and CPU offloading for memory that exceeds GPU capacity. This reduces optimizer memory from 16 bytes per parameter (AdamW) to 2 bytes, saving over 80% of optimizer memory. The paging mechanism moves unused optimizer states to CPU RAM automatically, which prevents out-of-memory errors during training when memory usage spikes.

Data Preparation

Fine-tuning data quality determines the outcome more than any hyperparameter. Each training example should demonstrate the exact behavior you want the model to learn. Noisy, inconsistent, or poorly formatted examples teach the model to produce noisy, inconsistent, or poorly formatted outputs.

# Data format for instruction fine-tuning
dataset = [
    {
        "instruction": "Summarize this clinical trial report in 3 bullet points.",
        "input": "A randomized controlled trial of 450 patients...",
        "output": "• Primary endpoint met: 23% reduction in...\n• Adverse events occurred in 12% of...\n• Subgroup analysis showed stronger effect in..."
    },
    # ...
]

# Convert to chat format for Llama-style models
def format_example(example):
    return f"""<|begin_of_text|><|start_header_id|>system<|end_header_id|>
You are a medical research assistant that summarizes clinical trial reports.
<|eot_id|><|start_header_id|>user<|end_header_id|>
{example['instruction']}

{example['input']}
<|eot_id|><|start_header_id|>assistant<|end_header_id|>
{example['output']}<|eot_id|>"""

Dataset size guidelines: 100-500 examples for style and format adaptation, 1,000-5,000 for domain terminology and reasoning patterns, 5,000-50,000 for complex multi-step tasks. Deduplicate aggressively — the model memorizes repeated examples, which inflates training metrics without improving generalization. Use attention pattern analysis to verify the model attends to the right parts of the input after fine-tuning.

Rank Selection Strategy

Choosing the right rank involves balancing quality, training cost, and inference overhead. Lower ranks train faster and produce smaller adapter files but may underfit complex adaptation tasks. Higher ranks capture more nuanced patterns but risk overfitting on small datasets.

RankTrainable Params (8B model)Best ForTraining Time
4~21M (0.26%)Format/style transferFastest
8~42M (0.52%)Simple domain adaptationFast
16~84M (1.03%)Domain knowledge + reasoningModerate
32~168M (2.07%)Complex multi-task adaptationSlow
64~335M (4.13%)Near full fine-tuning qualitySlowest

Start with rank 16 and adjust based on validation loss. If validation loss plateaus early and the training loss is much lower, the model is overfitting — reduce the rank or add more training data. If validation loss has not converged after several epochs, increase the rank to give the model more capacity to learn the task.

Merging and Deployment

After training, merge the LoRA adapters into the base model for deployment. The merged model has identical architecture and inference speed to the original — no PEFT library needed at inference time.

# Merge LoRA weights into base model
merged_model = model.merge_and_unload()

# Save the merged model
merged_model.save_pretrained("./merged-model")
tokenizer.save_pretrained("./merged-model")

# For quantized deployment, convert to GGUF
# llama.cpp handles the quantization for inference
# python convert_hf_to_gguf.py ./merged-model --outtype q4_k_m

For serving multiple LoRA adapters from the same base model, keep adapters separate and load them dynamically. Frameworks like vLLM and LoRAX support concurrent serving of multiple LoRA adapters from a single base model in GPU memory, switching between adapters per request with negligible overhead. This enables multi-tenant deployments where each customer gets a personalized model without duplicating the base model weights.

Common Pitfalls

Catastrophic forgetting happens when the learning rate is too high or the fine-tuning data is too narrow. The model loses general capabilities while specializing on the fine-tuning task. Prevent this by using learning rates below 5e-4, including a small percentage of general-purpose data in the training mix, and evaluating on both task-specific and general benchmarks throughout training.

Overfitting on small datasets produces a model that memorizes training examples rather than learning generalizable patterns. Signs include: training loss near zero while validation loss increases, and the model producing exact copies of training examples when prompted. Mitigate with dropout (0.05-0.1), reduced rank, fewer training epochs, and data augmentation through paraphrasing. Monitoring the training with tools from your time-series analysis toolkit helps identify divergence points early.

Quantization artifacts in QLoRA occasionally produce degraded outputs for specific input patterns. If QLoRA quality is noticeably worse than LoRA on the same task, try increasing the compute dtype to float32 for the LoRA forward pass while keeping the base model in 4-bit. This costs additional memory but eliminates numerical precision issues in the adapter computation.