TempatGunting

LLM Fine-Tuning: LoRA, QLoRA, and Parameter-Efficient Adaptation Methods

The Full Fine-Tuning Problem

Full fine-tuning of a large language model updates every parameter in the model. For a 7B parameter model in BF16, this requires 14GB just for the model weights, another 14GB for gradients, and 56GB for Adam optimizer states (two momentum buffers per parameter). That's 84GB of GPU memory before accounting for activations and batch data — exceeding the capacity of any single consumer GPU and requiring multi-GPU setups even for modest model sizes.

The memory cost scales linearly with parameter count. A 70B model needs roughly 840GB for full fine-tuning — a cluster of 8-12 A100 80GB GPUs. At cloud GPU rates, a single fine-tuning run costs hundreds to thousands of dollars. For organizations that need to fine-tune multiple models for different tasks, maintain multiple versions, or iterate rapidly on data composition, full fine-tuning is economically impractical.

Parameter-efficient fine-tuning (PEFT) methods address this by training only a small fraction of the parameters while keeping the rest frozen. LoRA is the dominant approach, combining strong performance with minimal memory overhead and a clean separation between the base model and task-specific adaptations. This separation turns out to be as important as the memory savings — it enables deployment patterns that aren't possible with full fine-tuning.

Input x [B, S, d] d = 4096 W (frozen) d × d, no grad A d × r B r × d Output W·x + B·A·x [B, S, d] + Only A and B are trained (r = 16: 0.3% of total params) A initialized with Gaussian, B initialized with zeros → Δ starts at zero

LoRA: The Mathematics of Low-Rank Adaptation

LoRA's key insight comes from the observation that the weight changes during fine-tuning have low intrinsic rank. When you fine-tune a pretrained model and compute ΔW = W_finetuned - W_pretrained, the matrix ΔW can be well-approximated by a low-rank decomposition ΔW ≈ B·A, where A is a d×r matrix and B is an r×d matrix, with r much smaller than d.

Instead of learning ΔW directly (which has d² parameters), LoRA learns A and B separately (which have 2·d·r parameters). For a hidden dimension of 4096 and rank 16, this reduces the trainable parameters per weight matrix from 16.7M to 131K — a 128x reduction. Across all adapted layers in a 7B model, total trainable parameters drop from 7B to approximately 20M.

from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-2-7b-hf",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

lora_config = LoraConfig(
    r=16,
    lora_alpha=32,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                     "gate_proj", "up_proj", "down_proj"],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
)

model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# trainable: 20,971,520 (0.29% of 7,244,738,560)

The initialization matters. Matrix A is initialized with a random Gaussian, while B is initialized with zeros. This means the LoRA modification starts as exactly zero (B·A = 0), and the model begins training from the pretrained weights unchanged. This zero-initialization is critical for training stability — it ensures that the model's behavior at the start of fine-tuning matches the pretrained model exactly.

The lora_alpha parameter controls the scaling of the LoRA output. The actual modification applied is (alpha / r) * B·A·x. Setting alpha = 2*r is a common default that balances the LoRA contribution against the frozen weights. Higher alpha values amplify the adaptation, which can speed convergence but risks overshooting.

QLoRA: Fine-Tuning at 4-Bit Precision

QLoRA takes LoRA's parameter efficiency and adds memory efficiency by quantizing the frozen base model to 4-bit precision. The key innovation is the NormalFloat4 (NF4) data type, which is information-theoretically optimal for normally distributed weights — and pretrained neural network weights are approximately normally distributed.

NF4 quantization divides the weight range into 16 bins (4 bits) placed at the quantiles of a standard normal distribution. Compared to uniform INT4 quantization, NF4 assigns more precision to the dense center of the distribution and less to the sparse tails, reducing quantization error by 20-30%. The base model stays in NF4 throughout training; only the LoRA adapter matrices are kept in BF16.

Double quantization further reduces memory by quantizing the quantization constants. Standard blockwise quantization uses one FP32 scaling constant per block of 64 weights, adding 0.5 bits per parameter overhead. Double quantization quantizes these constants to FP8, reducing the overhead to 0.127 bits per parameter. For a 7B model, this saves approximately 400MB of memory — seemingly small, but significant when working at the boundary of GPU capacity.

from transformers import BitsAndBytesConfig

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_use_double_quant=True,
    bnb_4bit_compute_dtype=torch.bfloat16,
)

model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-2-70b-hf",
    quantization_config=bnb_config,
    device_map="auto",
)

# 70B model now fits in ~35GB GPU memory
# Add LoRA on top for fine-tuning
model = get_peft_model(model, lora_config)

Paged optimizers complete QLoRA's memory optimization. When GPU memory is exhausted during training, optimizer states are automatically moved to CPU memory using CUDA unified memory. This adds latency when those states are needed, but prevents OOM crashes during the occasional long-sequence batch that temporarily spikes memory usage. The approach connects to general GPU memory management strategies for large model training.

Choosing the Right Configuration

The three key decisions in LoRA configuration are rank, target modules, and alpha scaling. Each affects quality, memory, and training speed differently.

Rank Selection

Rank determines the expressiveness of the adaptation. Think of it as the number of independent "concepts" the adaptation can encode. For a formatting change (JSON output, specific template), rank 4-8 is sufficient. For domain adaptation (medical, legal, financial), rank 16-32 captures the necessary distributional shift. For knowledge-intensive tasks requiring significant new information, rank 64 may help, though the returns diminish rapidly above 32 for most tasks.

We recommend starting with rank 16 and measuring validation loss convergence. If the model converges quickly (within 1 epoch) and validation loss plateaus far above training loss, the rank is likely too low and the model is capacity-constrained. If validation loss closely tracks training loss and continues decreasing throughout training, the rank is adequate.

Target Module Selection

Which layers receive LoRA adapters matters more than the rank. The original paper adapted only the query and value projections in attention. Recent practice adapts all linear layers in the transformer: all four attention projections (Q, K, V, O) plus the MLP layers (gate, up, down projections in LLaMA-style architectures).

ConfigurationTarget ModulesTrainable Params (7B)Memory (7B, BF16)Typical Quality
Minimalq_proj, v_proj~8M15 GB85-90% of full FT
StandardAll attention~16M16 GB92-96% of full FT
Full PEFTAll linear layers~21M17 GB95-100% of full FT
QLoRA FullAll linear (4-bit base)~21M5.5 GB93-98% of full FT
Full fine-tuningAll parameters7B84 GB100% (baseline)

For chat and instruction tuning, adapting all linear layers consistently outperforms attention-only configurations. The MLP layers carry factual knowledge, and instruction tuning often requires adjusting how the model accesses and formats that knowledge. The additional memory cost is modest — roughly 5M more parameters — and the quality improvement justifies it.

Training Recipes That Work

LoRA fine-tuning shares some hyperparameter sensitivities with full fine-tuning but introduces new ones. Learning rate is the most critical: LoRA adapters typically train best at 1e-4 to 3e-4, which is 10-100x higher than full fine-tuning learning rates. The intuition is that LoRA weights start at zero and need larger steps to reach their target values quickly.

Batch size interacts with learning rate in the usual way, but gradient accumulation is particularly important for QLoRA training where per-device batch size is limited by memory. We typically use a per-device batch size of 4 with 4 gradient accumulation steps for an effective batch of 16. Larger effective batches (32-64) improve training stability at the cost of slower convergence.

from transformers import TrainingArguments
from trl import SFTTrainer

training_args = TrainingArguments(
    output_dir="./lora-output",
    num_train_epochs=3,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,
    learning_rate=2e-4,
    lr_scheduler_type="cosine",
    warmup_ratio=0.03,
    bf16=True,
    logging_steps=10,
    save_strategy="epoch",
    evaluation_strategy="epoch",
    optim="paged_adamw_8bit",
    gradient_checkpointing=True,
    max_grad_norm=0.3,
)

trainer = SFTTrainer(
    model=model,
    args=training_args,
    train_dataset=train_data,
    eval_dataset=eval_data,
    max_seq_length=2048,
    dataset_text_field="text",
)

Gradient checkpointing trades compute for memory by recomputing intermediate activations during the backward pass instead of storing them. For LoRA training, this reduces memory by 30-50% at a 20-30% training speed penalty — a worthwhile trade when memory is the binding constraint. The training recipe ties into broader distributed training strategies for scaling beyond single-GPU setups.

Adapter Composition and Multi-Task Deployment

LoRA's cleanest advantage over full fine-tuning isn't memory — it's composability. Because LoRA adapters are additive modifications to frozen base weights, multiple adapters can share the same base model. A serving system loads the base model once and hot-swaps adapters per request, serving different tasks (summarization, translation, code generation) from a single model instance.

Adapter merging takes this further: multiple LoRA adapters can be arithmetically combined. Linear interpolation between two task-specific adapters produces a model that performs both tasks, often with minimal quality loss. More sophisticated merging techniques — TIES-Merging, DARE — selectively combine adapter weights based on their magnitude and direction, resolving conflicts where two adapters modify the same weights in opposite directions.

In production, adapter hot-swapping enables serving patterns impossible with full fine-tuning. A single GPU hosts the base model (14GB for a 7B model) and dozens of LoRA adapters (40MB each). When a request arrives tagged for a specific task, the appropriate adapter is loaded (if not already cached), applied to the base weights, and inference proceeds. The adapter swap takes milliseconds. This multi-tenant serving architecture connects to broader model serving patterns.

Beyond LoRA: The PEFT Landscape

LoRA isn't the only parameter-efficient method, though it's the most widely adopted. Several alternatives offer different trade-offs.

Prefix tuning prepends learned "virtual tokens" to the input at each transformer layer. These tokens don't correspond to real text — they're learned vectors that steer the model's behavior. Prefix tuning is effective for generation tasks but adds inference latency proportional to the prefix length (typically 20-100 tokens).

Adapters insert small bottleneck layers after the attention and MLP layers. Each adapter has a down-projection, a nonlinearity, and an up-projection, forming a narrow bottleneck. Adapters modify the residual stream directly and can express functions that LoRA cannot (due to the nonlinearity), but they add latency to every forward pass — unlike LoRA, which can be merged into the base weights for zero-overhead inference.

IA3 (Infused Adapter by Inhibiting and Amplifying Inner Activations) learns per-element scaling vectors for keys, values, and feed-forward intermediate activations. It trains even fewer parameters than LoRA (roughly 10x fewer) but with correspondingly lower expressiveness. IA3 is best suited for tasks where the model already has the necessary knowledge and just needs to adjust its output distribution.

When LoRA Isn't Enough

LoRA struggles in three scenarios. First, when the fine-tuning task requires substantial new factual knowledge that isn't in the base model. LoRA can adjust how the model uses existing knowledge but has limited capacity to inject new knowledge — the low-rank constraint restricts how much the weight matrix can change. For knowledge-intensive fine-tuning, consider retrieval augmentation alongside LoRA rather than pushing the rank higher.

Second, when the target domain's data distribution is radically different from the pretraining data. Fine-tuning a text model for protein sequences or musical scores requires changes across the entire weight space that a low-rank approximation can't capture. Full fine-tuning or continued pretraining on domain data followed by LoRA fine-tuning produces better results for extreme domain shifts.

Third, when serving latency is critical and the adapter cannot be merged into base weights. Some deployment scenarios require swapping adapters per-request too frequently for merging, and the separate adapter forward pass adds 5-10% overhead. For latency-sensitive applications, consider distilling the LoRA-adapted model into a standalone model that runs without the adapter mechanism. The distillation process relates to techniques covered in knowledge distillation for production-ready models.