Lesson 6 of 6 · 46 min

Capstone: fine-tune + eval a small model

Tie the track together: pick the lever (it’s a behaviour gap), LoRA/QLoRA-tune a small model on a clean stratified set, then PROVE the gain with a three-layer eval — capability benchmarks, a decontaminated golden set, and a calibrated LLM-as-judge — gated by a capability floor and a task ceiling. The discipline that separates "it got better" from "it got different."

The eval is the deliverable, not the fine-tune

Anyone can run a LoRA job; the senior skill is proving it helped. The interview research is blunt: candidates who describe a three-layer eval pipeline with concrete numbers get the senior offers, and those who say "we ran some examples and it looked better" get screened out. This capstone walks the full loop — decide, tune, and above all measure — on a concrete scenario: fine-tune Mistral-7B on 50k support tickets. The thesis of the whole track lands here: the data and the eval are the load-bearing parts; the algorithm is the smaller decision on top.
The scenario, end to end: a product team wants a support assistant that answers in their house voice, follows a strict resolution/escalation format, and reliably emits their tool schema. Step 0 — pick the lever (L1). Voice + format + tool-call reliability is a behaviour gap, not a knowledge gap, so this is the rare case where fine-tuning is the right primary tool (the live ticket data still goes through RAG; the rules go in the prompt). Step 1 — build the data (L2). Curate a clean, stratified slice of the 50k tickets — not all 50k — judge-filtered and decontaminated. Step 2 — tune (L3). QLoRA on the 7B, all-linear coverage, r=16/alpha=32 to start. Step 3 — prove it (this lesson). The three-layer eval, gated on a capability floor and a task ceiling, against the same base-model pipeline.
Lessons from a Year of Building with LLMsHamel Husain, Eugene Yan, et al.

Step 0–2: decide, data, tune (the track in three moves)

Anchor each move to its lesson so the capstone reads as one decision chain. The lever follows the L1 frame: classify the failure mode first — house voice, strict format, and tool-schema reliability are stable behavioural contracts the base prior won’t reach in context even with a careful prompt, so they go to a fine-tune; the weekly-changing ticket knowledge and current policy stay in RAG; per-turn rules stay in the prompt template. Stating that split out loud — "fine-tune the floor, RAG the ceiling, prompt the rules" — is the senior framing the interview rubric rewards, and it is the reason "just fine-tune everything" is wrong: facts go stale in weights and can’t be cited.
The data follows L2: from 50k tickets you do not train on all 50k. You stratify by ticket type, product area, and resolution outcome (coverage = the diversity axis), judge-filter for quality, compute inter-annotator agreement (Cohen’s kappa) on a labelled sample and drop non-converging examples, and — critically — decontaminate against the held-out eval set by exact and fuzzy match before training. A few thousand clean, stratified examples beat the full noisy 50k (the Guanaco/InstructGPT evidence). The tune follows L3: QLoRA (4-bit NF4 base + BF16 adapters, paged optimizer) so a 7B trains on one consumer GPU; cover all linear layers (not just q/v) to match full-FT quality; start r=16, alpha=32, low LR (2e-5–5e-5), 1–3 epochs; mix in ~5–15% general-instruction replay to fight catastrophic forgetting.
code
1THE CAPSTONE CHAIN  (each step cites its lesson)23  STEP 0  LEVER (L1)   behaviour gap (voice/format/tools) -> fine-tune4                       knowledge (live tickets) -> RAG;  rules -> prompt5  STEP 1  DATA  (L2)   stratify 50k -> clean subset; kappa; DECONTAMINATE vs eval6  STEP 2  TUNE  (L3)   QLoRA, ALL-linear coverage, r=16/a=32, LR 2-5e-5, 1-3 ep,7                       +5-15% general replay (anti-forgetting)8  STEP 3  PROVE (here) 3-layer eval vs SAME base pipeline; gate on floor+ceiling9  STEP 4  SERVE (L5)   merge adapter -> AWQ/FP8 + paged KV + continuous batching1011  The fine-tune is one line of STEP 2. Steps 1 and 3 are where seniors win.

Step 3: the three-layer eval — prove it got better, not different

The core deliverable. A domain fine-tune shifts the distribution, so "different" is guaranteed and "better" must be demonstrated on three layers that disambiguate them. Layer 1 — capability benchmarks: run a harness (lm-evaluation-harness, 60+ tasks: MMLU, GSM8K, HumanEval, TruthfulQA) before and after to detect regressions — this is your forgetting detector. Layer 2 — task-specific golden set: curate 200–500 held-out tickets, hand-label ground truth (resolution, sentiment, escalation), report accuracy/F1, and decontaminate against the training set so the score is real. Layer 3 — LLM-as-judge + human spot-check: have a stronger model grade 1,000+ outputs against a rubric, calibrated against 100 human grades. All three run against the same pipeline on the base model so the comparison is apples-to-apples.
code
1THREE-LAYER EVAL  (run base AND fine-tune through the SAME pipeline)23  Layer 1  CAPABILITY    lm-eval-harness: MMLU/GSM8K/HumanEval/TruthfulQA4           (regression)  -> the catastrophic-forgetting detector5  Layer 2  TASK GOLDEN   200-500 held-out tickets, hand-labeled ground truth,6           (in-domain)   accuracy/F1; DECONTAMINATE vs training set (or score is fake)7  Layer 3  JUDGE+HUMAN   stronger model grades 1000+ vs rubric, calibrated on8           (open-ended)  100 human grades (judge ~85% vs humans ~81% agreement)910  GATE:  capability FLOOR (e.g. MMLU must not drop > 2 pts)11   AND   task CEILING (golden-set metric >= target)12  Ship only if BOTH hold. Store the golden set as a FROZEN regression CI gate.
The calibration numbers anchor Layer 3 and a real interview probe. LLM-as-judge reaches ~85% agreement with humans, versus ~81% inter-annotator agreement between humans — so a well-calibrated judge is roughly as reliable as a second human reviewer, which is what makes it defensible at 1,000+ samples where human grading would cost ~52 person-days/month. But it must be calibrated: sample 100–200 outputs graded by humans, compute agreement, and tune the judge prompt if alignment drops below ~80%. Interview angle. "How do you know your fine-tune got better, not just different?" → the three layers, the decontaminated golden set, and the judge-vs-human agreement numbers; "we eyeballed a few outputs" is the answer that caps the round at mid-level. The honest framing is a marriage: LLM-as-judge for breadth, humans for calibration — pushing for either alone misses the empirical bounds of each.

Calibrating the judge — why Layer 3 is defensible

Layer 3 carries the most scrutiny, because "we let GPT-4 grade it" sounds like hand-waving until you show the calibration. The economics force the issue: grading 100k outputs by hand is ~52 person-days/month at $1–5+ per example, so human-only eval doesn’t scale past a spot-check — but an uncalibrated judge is just a second opinion you can’t trust. The fix is to measure the judge against humans: sample 100–200 outputs, have humans grade them, and compute agreement. The research anchor: a tuned LLM-as-judge reaches ~85% agreement with humans, while humans agree with each other only ~81% — so a calibrated judge is statistically about as reliable as a second human reviewer, which is what makes it defensible at the 1,000+ scale Layer 3 needs. If judge-human alignment drops below ~80%, you iterate the judge prompt (sharpen the rubric, add few-shot exemplars) before trusting its grades.
code
1JUDGE CALIBRATION  (breadth from the judge, calibration from humans)23  Human-only:  ~100k outputs x ~30s = ~52 person-days/mo at $1-5+/example -> spot-check only4  Judge-only:  cents + minutes, but UNTRUSTED until calibrated56  CALIBRATE:  sample 100-200 outputs -> human grade -> compute agreement7              judge ~85% vs humans ~81%  =>  judge ~= a second reviewer8              if alignment < ~80%: iterate the JUDGE PROMPT (rubric + few-shot)910  Use judge for BREADTH (grade 1000+), humans for CALIBRATION (the 100-200).11  Pushing for either alone misses the empirical bounds of each modality.

Reading the result: the multi-objective gate

Now interpret the numbers like a senior. The realistic outcome of a domain fine-tune: the task metric and human-eval scores rise, and a general benchmark dips — because specialisation trades some general capability for in-domain skill. That is expected, not a failure, and the right frame is multi-objective: set a capability floor (e.g. MMLU must not drop more than 2 points) and a task ceiling (golden-set metric ≥ target). If both hold, ship. If the task ceiling needs a higher learning rate that breaks the floor, that tension is the real decision — and naming it as a tradeoff (raise replay fraction, lower LR, bound the LoRA update) rather than dismissing the regression is the senior signal. The floor is set by running the base model through the identical eval and allocating a tolerance band.
The failure to be ready for: the fine-tune lifts support accuracy but MMLU drops 6 points. The weak answers are "ship it, MMLU is irrelevant" (ignores that a 6-point general drop surfaces as odd failures off the narrow task distribution) and "revert to prompting, fine-tuning failed" (over-correction — the task gain is real). The senior answer treats it as catastrophic forgetting (L1): it’s a train-data composition problem, so add ~5–15% general-distribution replay, lower the LR / bound the update (LoRA rank, fewer epochs), and re-gate. You only saw the regression because Layer 1 held out a general benchmark — which is the entire reason the three-layer eval exists. Interview angle. "Higher human-eval but lower MMLU — did it help?" → multi-objective: floor + ceiling, ship only if both pass; the tradeoff is the answer, not a side note.
A fine-tune that wins the task and clears the capability floor is a ship; a fine-tune that wins the task but breaks the floor is a tradeoff to negotiate, not a victory to announce. "It got better" without a decontaminated baseline and a floor is "it got different" wearing a number.

Step 4: serve it, and close the regression loop

A proven fine-tune still has to be served cheaply (L5), and the capstone choices follow directly. Merge the LoRA adapter into the base (W + (alpha/r)·B·A) so there’s zero inference latency — unless you’re serving many tenants’ adapters, in which case keep them unmerged and hot-swap over one shared base. Quantize for the target hardware: AWQ/GPTQ int4 if memory-bound, FP8 on Hopper/Blackwell for the best cost-per-token. Serve on vLLM with paged KV + continuous batching + prefix caching (the support prompt is shared, so prefix caching pays off). The eval numbers should be re-checked post-quantization — a 4-bit serve can shift quality slightly, so the golden set runs again on the actually-served artifact, not just the BF16 checkpoint.
Then make the eval permanent. The senior move the research names explicitly: store the golden set as a frozen regression suite alongside the model, re-run it on every new candidate (a new base model, a data refresh, a quantization change), and put a CI gate on minimum scores — task ceiling must hold, capability floor must not breach. This is what turns a one-time "did it help?" into a durable guarantee that the next change doesn’t silently regress. It also closes the track’s recurring theme: every silent failure mode — contamination, forgetting, diversity collapse, a quantization quality dip — is caught only by a clean, decontaminated, re-runnable eval. The eval is the product you actually maintain.
code
1SHIP + GUARD  (serve the proven artifact, then lock the eval)23  MERGE      W_merged = W + (alpha/r)·B·A  -> 0 extra latency4             (or keep adapters unmerged + hot-swap for multi-tenant)5  QUANTIZE   AWQ/GPTQ int4 (memory-bound)  |  FP8 (Hopper/Blackwell, best $/tok)6  SERVE      vLLM: paged KV + continuous batching + prefix cache (shared prompt)7  RE-EVAL    rerun the GOLDEN SET on the QUANTIZED artifact (4-bit can shift quality)8  GUARD      freeze golden set as a regression CI gate: ceiling holds, floor unbroken9             -> every future change (new base, data refresh, requant) must pass
articleUsing LLM-as-a-Judge for evaluation: a complete guideHamel Husainrepolm-evaluation-harness — 60+ benchmarks (MMLU, GSM8K, HumanEval)EleutherAIpaperQLoRA: Efficient Finetuning of Quantized LLMs (Guanaco recipe)Dettmers et al. (arXiv)articleApplied LLMs — lessons from a year of building (eval & iteration)Husain, Yan, Liu, et al.

Checkpoint

For the support-assistant scenario (house voice + strict format + tool schema, over weekly-changing tickets), what’s the right primary architecture?

AFine-tune the voice/format/tool behaviour, keep RAG for the live ticket knowledge, and put per-turn rules in the prompt — classify the failure mode per layerBFine-tune on everything including the ticket contents so the model knows it allCRAG only — retrieve everything and skip fine-tuning
Sign up free to answer and see why

Checkpoint

You have 50k support tickets. What’s the strongest data move before tuning the 7B?

ATrain on all 50k to maximise signalBStratify by ticket type/product/outcome, judge-filter for quality, drop low-kappa examples, and decontaminate against the held-out eval — then train a clean subset and let error analysis say where to add dataCGenerate 200k synthetic tickets to add volume
Sign up free to answer and see why

Checkpoint

How do you prove the fine-tune got better, not just different?

ARun 20 prompts through both models and read the outputsBCompare task-eval scores only; if higher, shipCThree layers vs the same base pipeline: capability benchmarks (regression), a decontaminated golden set (in-domain), and a calibrated LLM-as-judge — gated on a capability floor and a task ceiling
Sign up free to answer and see why

Checkpoint

The fine-tune lifts support-task accuracy by 9 points but MMLU drops 6. What do you do?

AShip it — MMLU doesn’t matter for supportBTreat it as catastrophic forgetting: add ~5–15% general-instruction replay, lower the LR / bound the LoRA update, re-train and re-gate on the capability floor (MMLU −2 max) and task ceilingCAbandon fine-tuning and go back to prompting
Sign up free to answer and see why

Checkpoint

The BF16 fine-tune passed your gates. You quantize to int4 for serving. What’s the disciplined final step?

AShip the quantized model — int4 quality loss is negligible so re-eval is unnecessaryBRe-run the golden set (and capability check) on the actual quantized artifact, and freeze the golden set as a regression CI gate for every future changeCTrust the BF16 eval numbers since quantization only affects speed
Sign up free to answer and see why

Interview prep

The capstone is the integration question: companies increasingly bundle fine-tuning and serving into one role because the value of a fine-tune depends on whether you can prove it helped and serve it cheaply. The eval is the highest-signal part — a three-layer pipeline with calibration numbers and a multi-objective gate beats any clever training trick. Lead with "better vs different," name decontamination unprompted, and treat the capability/task tension as the decision, not a footnote.
  1. 01“How do you know the fine-tune helped?” → 3 layers vs the same base pipeline: capability benchmarks (regression), decontaminated golden set, calibrated LLM-as-judge.
  2. 02“Better vs different?” → a domain fine-tune always shifts the distribution; "better" needs a task ceiling AND a capability floor, both measured against the base.
  3. 03“Is your LLM-judge trustworthy?” → calibrate on 100–200 human grades; judge ~85% vs human ~81% agreement, so a tuned judge ≈ a second reviewer — breadth from judge, calibration from humans.
  4. 04“Higher human-eval, lower MMLU — ship?” → multi-objective: floor (MMLU −2 max) + ceiling; ship only if both hold; the tradeoff is the answer.
  5. 05“Fine-tune regressed a general capability — fix?” → catastrophic forgetting: 5–15% general replay, lower LR, bound the update; you saw it because Layer 1 held out a general benchmark.
  6. 06“How much data from 50k tickets?” → a clean stratified subset, judge-filtered and decontaminated — quality over quantity; let error analysis drive more data.
  7. 07“Why decontaminate?” → eval examples leaking into training inflate scores; exact+fuzzy dedup against a sacred held-out set or the number is fake.
  8. 08“Serve the proven model?” → merge the adapter (0 latency), AWQ/FP8 quantize, vLLM + paged KV + prefix cache, then re-eval the quantized artifact and freeze a regression CI gate.
Going deeper: “what metric if ground truth is open-ended?” (LLM-as-judge with a calibrated rubric, or pairwise preference comparisons — not BLEU on free text); “how do you set the capability floor?” (run the base through the identical pipeline, allocate a tolerance band, gate on it); “eval without any ground truth?” (LLM-as-judge on faithfulness + a 100-sample human spot-check); and “why bundle fine-tuning and serving in one role?” (the operational impact of a fine-tune depends on serving it cheaply — the eval is the integration point of the whole loop). Always anchor to "prove it, against the base, on a decontaminated set."

Could you take a behaviour-gap scenario end to end — pick the lever, build a clean decontaminated dataset, QLoRA-tune a small model, and PROVE the gain with a three-layer, floor+ceiling-gated eval?

New to itGetting thereConfident

Track complete

  • The chain: classify the failure (L1) → clean, decontaminated, stratified data (L2) → QLoRA all-linear (L3) → three-layer eval → serve cheaply (L5).
  • Fine-tune only the behaviour gap; RAG the volatile knowledge; prompt the per-turn rules — "just fine-tune everything" stales facts and can’t cite.
  • Prove better-vs-different on three layers vs the SAME base pipeline: capability benchmarks, a decontaminated golden set, a calibrated LLM-as-judge (~85% vs ~81%).
  • Gate the ship on a capability floor (MMLU −2 max) AND a task ceiling; a general regression is catastrophic forgetting — fix with 5–15% replay + a bounded update.
  • Serve the proven artifact: merge the adapter (0 latency), AWQ/FP8 quantize, vLLM + paged KV + prefix cache — then re-eval the quantized model.
  • Freeze the golden set as a regression CI gate: the eval, not the fine-tune, is the durable deliverable that catches every silent failure.

Track complete — you can now decide the lever, fine-tune with LoRA/QLoRA, align with DPO/RLHF, serve cheaply at scale, and PROVE the gain against the base. That full loop is exactly what a senior LLM-engineer interview tests.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.