PTQ (post-training quantization) trades memory and latency for some accuracy loss. Quantization levels: FP16/FP32 (training precision), INT8 weights + FP16 compute (SmoothQuant, LLM.int8()), INT4 weights + group-wise scaling (GPTQ, AWQ), FP8 weights (Hopper native FP8). Picking one: (1) AWQ / GPTQ (INT4): aggressive memory savings (~4x), small quality loss, popular for self-hosted 70B-class models. (2) INT8 (SmoothQuant, LLM.int8()): safer quality, 2x memory savings. (3) FP8: native on H100, near-zero accuracy loss, faster than INT8 on Hopper. (4) INT4 + QLoRA-style tricks: finetuning on quantized base. Memory math decides: a 70B model fp16 = 140 GB, INT8 = 70 GB, INT4 = 35 GB, FP8 = 70 GB. Senior nuance: KV cache quantization (FP8 KV) is orthogonal: 2-4x memory savings on cache without quantizing weights. Quality tests are essential at each level because outliers vary per model.