QLoRA (Dettmers et al., 2023) combines 4-bit (NF4) quantization of the frozen base model with LoRA adapters trained in higher precision. The motivation is that even with LoRA the frozen model copy must reside in GPU memory during training, which dominates VRAM at 70B+ scales. By quantizing the frozen weights to 4-bit, you cut base-model memory 4x, letting you fine-tune a 65B model on a single 48GB GPU instead of an 80GB A100. Three additional tricks stabilize training: double quantization (quantize the quantization constants), paged optimizers (paged-adam to handle optimizer state spikes), and NF4 (normal-float 4) which is information-theoretically optimal for normally-distributed weights. Trade-off: a tiny quality hit (usually <1% on benchmarks) for 4x memory; some workloads show bigger regressions, especially on smaller base models. Senior caveat: an inference-time "QLoRA" model still needs to be dequantized, so don't confuse QLoRA's memory savings with decoding speed.