Lessons
1Fine-tune vs prompt vs RAG — the decision46 min read
The lever that costs the least and survives a model swap, the one that fixes knowledge, and the one that fixes behaviour — the cost/quality/latency triangle, the orthogonal-failure-modes mental model, and how to justify the choice in an interview without reaching for the most expensive tool.
- →Adaptation Strategy
2Datasets for SFT — quality over quantity48 min read
The dataset is the fine-tune — instruction-data construction and chat templates, why a small high-quality set beats a large noisy one, the QDC frontier (quality/diversity/complexity), synthetic data done right with judge filtering, and the data failure modes (contamination, format drift, low-diversity collapse) that silently wreck a run.
- →Training Data Quality
3LoRA & QLoRA — PEFT and the memory math50 min read
The low-rank update that froze the base model and changed production fine-tuning — the BA decomposition and what actually gets trained, rank/alpha/target-modules, why merged LoRA has zero inference latency, QLoRA’s NF4 + double-quant + paged-optimizer memory math (65B on one 48GB GPU), and the cases where LoRA underperforms full fine-tuning.
- →Efficient Fine-Tuning
4Preference optimization — DPO vs RLHF48 min read
After SFT teaches format, preference tuning teaches taste — the RLHF three-stage pipeline (SFT → reward model → PPO) and its failure modes, DPO’s closed-form collapse of that pipeline into one loss, when each wins (dense preferences vs sparse shaped rewards), and how alignment regresses capabilities and diversity.
- →Preference Optimization
5Quantization & efficient serving at scale50 min read
The model is trained — now serve it cheaply. Quantization quality is a function of method not bit-width (GPTQ/AWQ/fp8/GGUF with measured perplexity deltas), PagedAttention kills the 60–80% KV-cache waste, continuous batching is the 23–36x system win, speculative decoding buys 2–3x single-stream latency, and the vLLM/TGI/SGLang/TensorRT-LLM choice is workload- not vendor-driven — with the production numbers to budget against.
- →Efficient Model Serving
6Capstone: fine-tune + eval a small model46 min read
Tie the track together: pick the lever (it’s a behaviour gap), LoRA/QLoRA-tune a small model on a clean stratified set, then PROVE the gain with a three-layer eval — capability benchmarks, a decontaminated golden set, and a calibrated LLM-as-judge — gated by a capability floor and a task ceiling. The discipline that separates "it got better" from "it got different."
- →Adaptation Strategy
- →Training Data Quality
- →Efficient Fine-Tuning
Skills in this course
- 01Adaptation StrategyChoose prompting, retrieval, fine-tuning, or a combined approach for the actual gap.
- 02Training Data QualityBuild diverse, clean, correctly formatted post-training datasets and eval splits.
- 03Efficient Fine-TuningConfigure LoRA or QLoRA with sound rank, target-module, and memory choices.
- 04Preference OptimizationChoose DPO or RLHF based on preference data and reward requirements.
- 05Efficient Model ServingSelect quantization and serving techniques for measured quality, latency, and cost.