Lesson 3 of 6 · 50 min

LoRA & QLoRA — PEFT and the memory math

The low-rank update that froze the base model and changed production fine-tuning — the BA decomposition and what actually gets trained, rank/alpha/target-modules, why merged LoRA has zero inference latency, QLoRA’s NF4 + double-quant + paged-optimizer memory math (65B on one 48GB GPU), and the cases where LoRA underperforms full fine-tuning.

Why nobody full-fine-tunes by default anymore

Full fine-tuning a 175B model in fp16 needs ~1.2 TB of VRAM just for weights, gradients, and optimizer state — out of reach for almost everyone. LoRA (Hu et al., 2021) made that number ~350 GB and the saved checkpoint 35 MB, at parity with full fine-tuning on the tasks it was tested on. QLoRA then put a 65B fine-tune on a single 48 GB GPU. This is why “LoRA, probably QLoRA” is the default production answer, and why the memory math is a near-guaranteed interview drill. This lesson is the mechanism, the numbers, and the limits.
The mechanism, stated precisely because interviewers probe it. LoRA freezes the pretrained weights and injects a trainable rank decomposition into selected linear layers. For a weight matrix W (d×k), instead of learning a full update ΔW you learn ΔW = B·A, where A is r×k and B is d×r, and r (the rank) is tiny — often 8 to 64. The forward pass becomes h = W·x + (alpha/r)·B·A·x. A is initialised from a small Gaussian and B is initialised to zero, so at step 0 the adapter contributes nothing and the model starts exactly at the base distribution. The hypothesis behind it: the update a fine-tune needs lives on a low-dimensional subspace, so a low-rank B·A captures it.
code
1LoRA: freeze W, learn a low-rank update  ΔW = B·A23         x ──►  W (frozen, d×k) ───────────────►(+)──► h4                │                                  ▲5                └─► A (r×k, Gaussian) ─► B (d×r, =0)┘   scaled by alpha/r67  A: r×k   B: d×r   r = rank (e.g. 8-64) << min(d,k)8  trainable params = r·(d+k)   vs   full = d·k     -> ~0.01-0.5% of the model9  B initialised to ZERO  ->  step 0 == base model (no shock)10  At deploy:  W_merged = W + (alpha/r)·B·A   ->  ZERO extra inference latency
The single most striking empirical fact in the LoRA paper: on GPT-3 175B, a rank of r = 1 or 2 was sufficient even though the attention weights are 12,288-dimensional. You are approximating a 12,288-wide update with a rank-1 or rank-2 factorisation and matching full fine-tuning. That is the whole PEFT thesis in one number — the adaptation manifold is shockingly low-rank. Interview angle. “Why does adding two small matrices approximate full fine-tuning?” → the fine-tuning delta is intrinsically low-rank, so B·A with tiny r captures it; you’re not approximating W, you’re approximating the much-simpler ΔW. Candidates who frame it as a delta-on-the-update-manifold (not “freezing weights”) stand out.
QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers (AMLD)

What actually gets trained — rank, alpha, target modules

Three knobs decide a LoRA run. Rank (r) — typically 16–128 — sets the capacity of the update; higher r learns more but costs more memory and risks overfitting on small data. Alpha — usually 1–2× rank — is the scaling on the adapter (the forward multiplier is alpha/r), so the common recipe alpha = 2r keeps the effective update scale stable as you change r. Target modules — which linear layers get an adapter — is the highest-leverage and most-misunderstood knob. The cheapest config adapts only the attention q_proj and v_proj. But the QLoRA ablation on LLaMA-7B Alpaca showed that to match full fine-tuning you must apply LoRA to all linear transformer-block layers, not just q/v.
code
1LoRA KNOBS  (what to set and why)23  Knob            Typical          Effect / rule4  -------------   --------------   ----------------------------------------5  rank r          16-128           capacity of the update; too high overfits6  alpha           1-2x rank        forward scale = alpha/r; alpha=2r is common7  target modules  q_proj,v_proj    cheapest; matches full FT only if you cover8                  ...ALL linear     ALL linear layers (QLoRA LLaMA-7B ablation)9  lora_dropout    0-0.1            regularise on small data10  rank-stabilized alpha/sqrt(r)    keeps effective LR stable across ranks (rsLoRA)1112  Start: r=16, alpha=32, q_proj+v_proj -> measure -> expand coverage + rank13  ONLY if held-out quality misses the bar. Coverage > rank for matching full FT.
The target-modules finding is the senior gotcha and a favourite follow-up. “What changes if you target all linear layers vs just q/v?” → broader coverage is what lets LoRA match full fine-tuning; rank alone can’t buy that. So the practical recipe is: start at r=16, alpha=32, q_proj+v_proj only (cheap, fast), measure on a held-out set, and expand coverage to all-linear and raise rank only if quality misses the bar. Two refinements worth naming: rank-stabilised LoRA uses alpha/√r instead of alpha/r to keep the effective learning rate stable across ranks, and LoRA-GA aligns the adapter init with the full-fine-tuning gradient and reports 2–4× faster convergence.
A clean, memorisable footprint number from the interview research: for Llama-2-7B at r=8, LoRA trains 4,194,304 parameters — about 16.78 MB, roughly a 1000× reduction versus the full model. That tiny adapter is the entire artifact you ship and version. Interview angle. Being able to produce “r=8 on a 7B → ~4.2M trainable params, ~17 MB adapter, ~1000× fewer than full” on demand is exactly the kind of concrete number that separates a strong LoRA answer from a hand-wave.

The inference story — and the train-vs-inference memory trap

LoRA has zero additional inference latency, and the reason is the merge: at deployment you compute W_merged = W + (alpha/r)·B·A once, and the result is a dense matrix bit-identical to what full fine-tuning would have produced. A merged LoRA model is not “a base model plus an adapter at runtime” — it’s a normal dense model. This is the architectural reason LoRA dominates adapter methods in production: classic adapters insert an extra layer and pay a forward hop on every token; a merged LoRA pays nothing. (You can also keep adapters unmerged and hot-swap them per request — the basis for multi-tenant adapter serving — at the cost of a small runtime add.)
The trap that catches people: train-time memory savings are large; inference-time memory savings are essentially zero. At training, freezing the base weights means no optimizer state for them — and Adam’s optimizer state (two moments per trainable parameter) is the dominant memory term, so cutting trainable params ~1000× slashes optimizer + gradient memory. But at inference, a merged LoRA is the same size as the dense model; you saved nothing on serving memory unless you’re hot-swapping adapters over a shared base. Sebastian Raschka’s practical write-ups call this out explicitly. Interview angle. “What memory does LoRA actually save, train vs inference?” → train: huge (no optimizer/gradient state for frozen weights); inference: ~none after merge. Conflating the two is a classic weak answer.

QLoRA — the three tricks that put 65B on one 48GB GPU

QLoRA (Dettmers et al., 2023) backpropagates gradients through a frozen, 4-bit-quantized base model into the LoRA adapters — compressing the dominant memory line item (base-model weights) without disturbing gradient flow into the adapters. Three innovations make it work, and a strong answer names all three rather than saying “it’s 4-bit LoRA.” (1) NF4 (4-bit NormalFloat) — a data type that is information-theoretically optimal for normally-distributed weights, with bins holding equal probability mass, which fits trained weights’ Gaussian shape better than uniform int4. (2) Double quantization — quantizing the quantization constants themselves, cutting the per-parameter overhead from 0.5 bits to ~0.127 bits (≈3 GB saved on a 65B model). (3) Paged optimizers — using NVIDIA unified memory to page optimizer state to CPU during the gradient-checkpoint memory spike on long sequences, killing the OOM-on-long-context failure mode.
code
1QLoRA = 4-bit FROZEN base  +  BF16 LoRA adapters23  Forward: dequantize W (NF4 -> BF16) on the fly, compute in BF16, add B·A4    Y_BF16 = X · doubleDequant(W_NF4) + X · B_BF16 · A_BF1656  THE THREE TRICKS7  1. NF4            4-bit type, info-theoretically optimal for ~Gaussian weights8  2. Double-quant   quantize the quant constants: 0.5 bits/param -> ~0.127 bits9                    (~3 GB saved on a 65B model)10  3. Paged optim    page optimizer state CPU<->GPU on the long-seq memory spike11                    -> no more OOM-on-long-context1213  Adapters stay BF16; only the frozen base is NF4. "Matches 16-bit LoRA and14  full-finetuning task performance." (Dettmers et al., 2023)
The memory math is the headline interview drill, so know the table. QLoRA’s total training footprint at batch size 1, sequence length 512 is 45.0 GB for LLaMA-65B and 24.7 GB for LLaMA-33B — both fit on a single 48 GB GPU, where full 16-bit fine-tuning would need several hundred GB. A 7B QLoRA model’s base weights are ~5 GB, so a single 24 GB consumer card can fine-tune it with headroom. The interview research adds a concrete tradeoff measurement: on Llama-2-7B, QLoRA hit ~17.86 GB peak vs ~36.66 GB for a two-GPU full fine-tune — a ~33% memory cut at a ~39% runtime increase. That runtime cost (the dequantize-on-the-fly tax) is QLoRA’s tradeoff, and naming it is the senior signal.
code
1QLoRA MEMORY  (training, bs=1, seq=512)            FITS ON23  Model        QLoRA total      Base weights      single GPU?4  ----------   --------------   --------------    -----------------5  LLaMA 7B     ~base 5 GB       ~5 GB             yes (24 GB)6  LLaMA 33B    24.7 GB          ~21 GB            yes (24-32 GB)7  LLaMA 65B    45.0 GB          ~41 GB            yes (48 GB)89  Tradeoff (Llama-2-7B, interview-research numbers):10    QLoRA ~17.86 GB peak  vs  full (2-GPU) ~36.66 GB11    => ~33% LESS memory at ~39% MORE runtime (dequant-on-the-fly tax)
The measured-quality result is what makes QLoRA safe to default to: it “replicates 16-bit LoRA and full-finetuning task performance,” and the Guanaco family it produced reached 99.3% of ChatGPT on the Vicuna benchmark from a single GPU in 24 hours — with a 7B Guanaco beating a 26 GB Alpaca by 20+ points (the quality-over-quantity story from L2). Interview angle. “When would you skip QLoRA and use plain LoRA?” → when serving/training hardware is fp16-plentiful and you don’t need the 4-bit compression, because QLoRA pays a dequantize tax at every forward step; QLoRA is the move when memory is the binding constraint, which at 33B+ on a single GPU it almost always is.

When LoRA underperforms full fine-tuning

LoRA is the default, not a free lunch — and a senior answer names where it falls short. (1) Under-coverage: the most common cause of a LoRA that misses full-FT quality is adapting too few modules. The QLoRA ablation is explicit — q/v-only does not match full fine-tuning; you need LoRA on all linear layers. (2) Rank too low for the task: a genuinely high-rank adaptation (a large domain shift, a new capability rather than a behaviour tweak) can exceed what a small r captures, so quality plateaus below full-FT until you raise r and coverage. (3) Large distribution shift / continued pretraining: when you’re trying to move the model a long way (new language, new modality-style data, heavy domain adaptation), full or near-full fine-tuning still wins, because you’re changing the base prior, not nudging it.
The decision rule that ties it together: LoRA matches full fine-tuning for behaviour-shaped adaptation (style, format, instruction-following, moderate domain) when you cover all linear layers at adequate rank; full fine-tuning earns its cost when the change is capability-shaped or a large distribution shift. Coverage and rank are the two dials you turn before concluding LoRA can’t reach the bar. Interview angle. “Your LoRA plateaus below full fine-tuning — what do you try before giving up on PEFT?” → expand target modules to all-linear, raise rank (and use rsLoRA/alpha=2r to keep the effective LR sane), check the data isn’t the real ceiling (L2), and only then consider full FT. Jumping straight to full FT signals you don’t know coverage is usually the culprit.
paperLoRA: Low-Rank Adaptation of Large Language ModelsHu et al. (arXiv)paperQLoRA: Efficient Finetuning of Quantized LLMsDettmers et al. (arXiv)articlePractical Tips for Finetuning LLMs Using LoRASebastian RaschkadocsHugging Face PEFT — LoRA reference (rank, alpha, target_modules)Hugging Face

Checkpoint

An interviewer asks why a rank-8 B·A update can approximate a full fine-tune of a 12,288-dimensional attention weight. Strongest answer?

ABecause rank-8 matrices can represent any 12,288-dim matrixBYou’re approximating the fine-tuning delta ΔW, not W — and the delta is intrinsically low-rank, so a small-r B·A captures it (LoRA found r=1–2 sufficed on GPT-3)CBecause the base weights are frozen, the update can be any size
Sign up free to answer and see why

Checkpoint

Your LoRA fine-tune (q_proj + v_proj, r=16) consistently lands ~3 points below a full fine-tune on the held-out set. The data is clean. What do you try first?

AExpand target modules to all linear layers (and raise rank if needed) — coverage, not rank alone, is what matches full FTBSwitch to full fine-tuning immediatelyCLower the learning rate to stabilise training
Sign up free to answer and see why

Checkpoint

A teammate says “we used LoRA, so serving will use much less GPU memory than the base model.” Correct them.

AThey’re right — LoRA models are smaller at inferenceBThe savings are at training (no optimizer state for frozen weights); a merged LoRA serves at full size — the only inference win is hot-swapping many adapters over one shared baseCOnly if you also quantize the adapter to 4-bit
Sign up free to answer and see why

Checkpoint

You need to fine-tune a 33B model and you have a single 48GB GPU. What’s the defensible approach, and what’s the cost you should flag?

AFull fine-tuning in fp16 — it’s the highest qualityBQLoRA (4-bit NF4 base + BF16 adapters, paged optimizer) — ~24.7GB fits 48GB; flag the ~30–40% runtime tax from dequantizing on the fly and cover all linear layers for qualityCPlain LoRA in fp16 — skip quantization
Sign up free to answer and see why

Checkpoint

When is full (or near-full) fine-tuning genuinely worth its cost over LoRA?

AAlways — full fine-tuning is strictly betterBNever — LoRA is mathematically equivalentCFor large distribution shifts or new capabilities (new language/modality-style, heavy domain adaptation) where you’re moving the base prior, not nudging behaviour
Sign up free to answer and see why

Interview prep

LoRA/QLoRA is the most mechanism-and-numbers-heavy topic in the track, and the memory math is a near-guaranteed drill. Interviewers want the BA decomposition derived, the three QLoRA tricks named, a memory number produced on demand, and an honest account of where LoRA underperforms. The recurring failure point is people who say “lower memory” with no numbers and conflate train-time with inference-time savings.
  1. 01“Explain LoRA.” → freeze W, learn ΔW = B·A at rank r; forward h = Wx + (alpha/r)·B·A·x; B init 0 so step 0 = base. The delta is low-rank.
  2. 02“Why does low rank work?” → you approximate the fine-tuning delta, not W; it’s intrinsically low-rank — GPT-3 needed only r=1–2.
  3. 03“Rank, alpha, target modules?” → r 16–128 (capacity), alpha 1–2× r (scale alpha/r), and you must cover ALL linear layers to match full FT, not just q/v.
  4. 04“Train vs inference memory savings?” → train: huge (no optimizer state on frozen weights); inference: ~none after merge — merged LoRA is full-size.
  5. 05“Why zero inference latency?” → merge W_merged = W + (alpha/r)·B·A once; it’s a normal dense matrix, unlike adapters that add a forward hop.
  6. 06“QLoRA — 65B on one 48GB GPU how?” → NF4 4-bit base + double-quant (0.5→0.127 bits) + paged optimizers; ~45GB total at bs1/seq512.
  7. 07“QLoRA quality / cost?” → matches 16-bit LoRA & full FT; Guanaco hit 99.3% of ChatGPT in 24h on one GPU; tradeoff is ~30–40% slower (dequant tax).
  8. 08“When does LoRA underperform?” → under-coverage (q/v-only), rank too low, or large distribution shift / new capability — expand coverage + rank before full FT.
Going deeper, expect: “r=8 on a 7B — how many trainable params?” (~4.2M, ~17 MB adapter, ~1000× fewer than full); “skip QLoRA when?” (fp16 memory is plentiful and you don’t want the per-step dequant tax); “what speeds it up?” (Unsloth ~2× faster / ~70% less VRAM; Liger kernels ~20% throughput / ~60% memory — name them as the production tooling); and “how do you serve many fine-tunes cheaply?” (keep adapters unmerged and hot-swap over one base — multi-tenant LoRA serving, the lone inference-memory win). Always pair a mechanism with a number.

Could you derive the BA update, set rank/alpha/target modules, do QLoRA’s memory math for a 33B/65B model, and name where LoRA underperforms?

New to itGetting thereConfident

Takeaways

  • LoRA freezes W and learns ΔW = B·A at tiny rank; the fine-tuning delta is intrinsically low-rank (GPT-3 needed r=1–2).
  • Coverage > rank for matching full FT: cover ALL linear layers, not just q/v (QLoRA’s LLaMA-7B ablation).
  • Merged LoRA = full dense model: zero inference latency, but inference memory savings are ~none — savings are at training.
  • QLoRA = NF4 + double-quant (0.5→0.127 bits) + paged optimizers → 65B on one 48GB GPU (~45GB at bs1/seq512).
  • QLoRA matches full-FT quality (Guanaco 99.3% of ChatGPT, 24h, 1 GPU) at a ~30–40% runtime tax from on-the-fly dequant.
  • Quantization sets the memory floor; adapter coverage sets the quality ceiling — independent dials.

Next: preference optimization — DPO vs RLHF, the alignment intuition, when preference tuning helps, and how it can regress capabilities.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.