Lessons

1Fine-tune vs prompt vs RAG — the decision46 min read

The lever that costs the least and survives a model swap, the one that fixes knowledge, and the one that fixes behaviour — the cost/quality/latency triangle, the orthogonal-failure-modes mental model, and how to justify the choice in an interview without reaching for the most expensive tool.

  • →Adaptation Strategy
Read lesson
2Datasets for SFT — quality over quantity48 min read

The dataset is the fine-tune — instruction-data construction and chat templates, why a small high-quality set beats a large noisy one, the QDC frontier (quality/diversity/complexity), synthetic data done right with judge filtering, and the data failure modes (contamination, format drift, low-diversity collapse) that silently wreck a run.

  • →Training Data Quality
Read lesson
3LoRA & QLoRA — PEFT and the memory math50 min read

The low-rank update that froze the base model and changed production fine-tuning — the BA decomposition and what actually gets trained, rank/alpha/target-modules, why merged LoRA has zero inference latency, QLoRA’s NF4 + double-quant + paged-optimizer memory math (65B on one 48GB GPU), and the cases where LoRA underperforms full fine-tuning.

  • →Efficient Fine-Tuning
Read lesson
4Preference optimization — DPO vs RLHF48 min read

After SFT teaches format, preference tuning teaches taste — the RLHF three-stage pipeline (SFT → reward model → PPO) and its failure modes, DPO’s closed-form collapse of that pipeline into one loss, when each wins (dense preferences vs sparse shaped rewards), and how alignment regresses capabilities and diversity.

  • →Preference Optimization
Read lesson
5Quantization & efficient serving at scale50 min read

The model is trained — now serve it cheaply. Quantization quality is a function of method not bit-width (GPTQ/AWQ/fp8/GGUF with measured perplexity deltas), PagedAttention kills the 60–80% KV-cache waste, continuous batching is the 23–36x system win, speculative decoding buys 2–3x single-stream latency, and the vLLM/TGI/SGLang/TensorRT-LLM choice is workload- not vendor-driven — with the production numbers to budget against.

  • →Efficient Model Serving
Read lesson
6Capstone: fine-tune + eval a small model46 min read

Tie the track together: pick the lever (it’s a behaviour gap), LoRA/QLoRA-tune a small model on a clean stratified set, then PROVE the gain with a three-layer eval — capability benchmarks, a decontaminated golden set, and a calibrated LLM-as-judge — gated by a capability floor and a task ceiling. The discipline that separates "it got better" from "it got different."

  • →Adaptation Strategy
  • →Training Data Quality
  • →Efficient Fine-Tuning
Read lesson

Skills in this course

  1. 01Adaptation StrategyChoose prompting, retrieval, fine-tuning, or a combined approach for the actual gap.
  2. 02Training Data QualityBuild diverse, clean, correctly formatted post-training datasets and eval splits.
  3. 03Efficient Fine-TuningConfigure LoRA or QLoRA with sound rank, target-module, and memory choices.
  4. 04Preference OptimizationChoose DPO or RLHF based on preference data and reward requirements.
  5. 05Efficient Model ServingSelect quantization and serving techniques for measured quality, latency, and cost.