🔧 Fine-tuning Strategies
Adapting pre-trained models to specific tasks with full fine-tuning, LoRA, QLoRA, and other PEFT methods
On this page
Before you start
Use linear algebra for matrix products and rank, calculus and optimization for gradients and Adam, and transformers for attention projections. Numerical computing explains precision and memory units; evaluation explains held-out comparisons.
By the end, you should be able to choose an adaptation method from a measured failure, compute a LoRA update and its parameter count, budget training memory with explicit assumptions, and evaluate specialization without overlooking regressions.
The problem: a support model keeps breaking the output contract
Suppose an assistant must return a category and a short explanation for each support ticket. A prompted baseline recognizes many tickets but inconsistently follows the output schema and misses rare categories. Meanwhile, refund policies change weekly.
These are different failure sources. A schema validator or constrained decoder can enforce structure; targeted supervised examples can teach classification behavior; retrieval can supply the current refund policy. Fine-tuning changes model parameters, so it can adapt behavior, skills, and stable domain knowledge. It does not provide a reliable update mechanism or source citation for each learned fact. Diagnose the failing part before paying for training.
Worked example: a rank-one update
Consider a frozen linear layer with two inputs and three outputs:
W = [[1, 0], [0, 1], [1, 1]]
x = [3, 1]
W x = [3, 1, 4]
Instead of training all six entries of W, train two smaller matrices:
A = [[1, -1]] shape 1 × 2
B = [[0.5], [0], [-0.5]] shape 3 × 1
A x = [2]
B(A x) = [1, 0, -1]
With scaling equal to one, the adapted output is [3, 1, 4] + [1, 0, -1] = [4, 1, 3]. The effective update is:
B A = [[0.5, -0.5], [0, 0], [-0.5, 0.5]]
All its nonzero rows are multiples of one direction, so its rank is one. This tiny example trains five parameters instead of six; the saving becomes substantial for large layers. At d = k = 4096 and r = 8, a full update has dk = 16,777,216 parameters, while LoRA has r(d + k) = 65,536, exactly 1/256 as many.
Notation and the learning objective
| Symbol | Meaning |
|---|---|
| W, shape d × k | Frozen base projection, output dimension d and input dimension k |
| A, shape r × k; B, shape d × r | Trainable factors with rank budget r |
| α | Adapter scale; this chapter uses scaling α/r |
| θ | All parameters selected for training |
| mₜ | Loss mask: one for a supervised token, zero otherwise |
| pθ(yₜ given prefix) | Model probability of the target token |
For supervised fine-tuning (SFT), a common token-averaged objective is:
L = −Σₜ mₜ log pθ(yₜ given prefix) / Σₜ mₜ.
The denominator must be nonzero. In chat SFT, supervising assistant tokens while masking user and system tokens is common. Continued pretraining instead usually supervises ordinary next-token predictions throughout the document. Other masks and per-example weighting are valid, but change the objective; specify them. Two supervised tokens with probabilities 0.8 and 0.5 have mean loss −(ln 0.8 + ln 0.5)/2 ≈ 0.458 nats.
How LoRA changes training
The forward pass is h = Wx + (α/r)BAx. The base W remains frozen, while backpropagation updates A and B. A low-rank update is an architectural restriction that often works well; it is not a theorem that every useful adaptation is low rank. Rank limits the subspace of the update, not its magnitude: scaling B can make ‖BA‖ arbitrarily large.
A common initialization makes A random and B zero. Then BA = 0, so the wrapped layer initially matches the base. The first gradient for A is zero because it passes through B, while B can receive a nonzero gradient through A. Initializing both factors to zero would leave both unable to start learning.
Target layers are an experimental choice. Attention query/value projections are a compact starting configuration; other attention and MLP projections add capacity and memory cost. Compare target sets and ranks under the same data and evaluation budget. Changing rank also changes scaling unless α is adjusted under the chosen convention.
At inference, a dense base can absorb the update as Wmerged = W + (α/r)BA, removing the separate adapter operations. Unmerged adapters allow switching tasks without duplicating the base. Merging into a quantized representation may require dequantization and requantization; re-evaluate the resulting model.
Memory: count bytes, then profile
For 7 billion parameters, use decimal GB and assume BF16 weights and gradients, FP32 Adam first and second moments, and optional FP32 master weights:
| Persistent state | Bytes per parameter | 7B total |
|---|---|---|
| BF16 weights | 2 | 14 GB |
| BF16 gradients | 2 | 14 GB |
| Two FP32 Adam moments | 8 | 56 GB |
| Optional FP32 master weights | 4 | 28 GB |
| Total with master weights | 16 | 112 GB |
| Total without master weights | 12 | 84 GB |
These are illustrative persistent-state totals, before activations, temporary buffers, communication, and allocator overhead. Implementations vary: gradients or optimizer states may use different precision, and sharding/offloading changes per-device residency. A GPU fit claim needs a sequence length, microbatch size, checkpointing strategy, optimizer, and distributed configuration.
LoRA removes base-weight gradients and optimizer states, but still stores the frozen base and backpropagates through layers to reach the adapters. Activations remain important. QLoRA stores the frozen base in a low-bit representation; 7B raw 4-bit values occupy 3.5 GB before quantization metadata and all other state. Neither number is total training memory.
QLoRA's original method combines NF4 quantization, double quantization of quantization constants, and paged optimizers to manage memory pressure. Computation uses dequantized values in a supported compute precision; the frozen 4-bit values are not directly trained. Its reported hardware experiments demonstrate particular configurations, not a universal fit or speed guarantee.
A training experiment you can trust
- Establish the baseline. Evaluate prompting, constrained output, and retrieval where appropriate. Measure task accuracy, valid output rate, latency, and cost.
- Split before iteration. Group related documents, customers, or templates to prevent near-duplicate leakage across training and evaluation. Keep a separate final test set and a regression suite for capabilities that must survive.
- Validate the actual training tokens. Inspect chat-template special tokens, stop tokens, truncation, label masks, and packed-example attention boundaries. Check that important target tokens are not discarded.
- Run a small controlled sweep. Compare learning rate, rank, target layers, and training duration. Freeze data and evaluation while changing one factor or a small planned grid. Adapter learning rates need their own tuning; there is no required ratio to full fine-tuning.
- Track task metrics alongside loss. Lower cross-entropy can coexist with worse structured output, privacy leakage, or general-task performance. Compare checkpoints and retain a rollback artifact.
- Validate serving parity. Use the same tokenizer/template, adapter and base versions, quantization, and decoding policy as deployment. Record these with the dataset version and training configuration.
Gradient accumulation helps fit a desired batch into memory. With 9,000 training examples, one device, microbatch 4, and accumulation 8, the nominal effective batch is 32. If the final partial batch is flushed each epoch, there are ceil(9000/32) = 282 optimizer steps per epoch, or 564 for two epochs. Packing and distributed sampling can change this accounting.
Boundaries and optional methods
Full fine-tuning allows unrestricted updates to the selected base weights; it may help when a constrained adapter underfits, but can also overfit or forget. Replay data and penalties against a reference can reduce regression; their weights must be chosen using both target-task and retained-capability evaluations. Low rank alone does not guarantee retention.
Other PEFT methods move the trainable parameters elsewhere: prompt tuning learns input embeddings, prefix tuning learns layer-level attention prefixes, bottleneck adapters insert small modules, and IA³ learns activation scaling vectors. Their quality, latency, and parameter counts depend on the architecture and task. Choose them when their deployment or capacity properties solve a measured constraint.
SFT learns from demonstrations and can generalize beyond exact training examples. RL learns from rewards on sampled behavior and can also fail to explore or exploit a weak grader. They are complementary sources of supervision, not a hierarchy in which SFT cannot reason and RL always improves.
Check your understanding
For a 2048 × 1024 projection with rank 4, how many LoRA parameters are trained? Does that ratio also describe the reduction in total GPU memory? If every update row is a multiple of one row, can its norm still be large?
Solution
The adapter has 4 × (2048 + 1024) = 12,288 parameters versus 2,097,152 for a full update, about 0.586%. Total memory does not shrink by the same ratio because the frozen base, activations, and workspaces remain. A rank-one update can have arbitrarily large norm by scaling its nonzero factors; low rank does not bound update size.
Continue learning
Use RAG to supply current, citable evidence, RLHF and DPO to learn from preferences, and Harness Engineering to connect reproducible training with evaluation and serving.
References
- Hu et al., LoRA: low-rank updates, initialization, and empirical adaptation experiments.
- Dettmers et al., QLoRA: NF4, double quantization, and paged optimizers.
- Kingma and Ba, Adam: first- and second-moment optimizer state.
- Ouyang et al., InstructGPT: supervised demonstrations within a larger post-training pipeline.