LoRA (Low-Rank Adaptation, Hu et al. 2021) freezes the pretrained weight matrix W and learns a low-rank update Delta W = BA, where B is d x r, A is r x k, r << min(d,k). During the forward pass, h = W x + BA x. Because BA is low-rank, the trainable parameter count drops from dk to r(d+k), often 100-1000x compression. Empirically this matches full fine-tuning quality on many tasks while using a fraction of GPU memory and storage (you only store the adapter, not a full model copy). Trade-offs: (1) rank r exposes a quality/capacity knob: too low underfits, too high approaches full FT; (2) merging adapters back to W at deployment is lossless; (3) LoRA assumes the task-relevant update lives in a low-rank subspace, which holds empirically but is not theoretically guaranteed; (4) compared to adapters / prompt tuning, LoRA is the sweet spot between quality and efficiency.