The dataset is the fine-tune — instruction-data construction and chat templates, why a small high-quality set beats a large noisy one, the QDC frontier (quality/diversity/complexity), synthetic data done right with judge filtering, and the data failure modes (contamination, format drift, low-diversity collapse) that silently wreck a run.
The fine-tune is the data
Across the entire literature, one finding dominates: data quality outweighs parameter optimisation. mlabonne’s course states it flatly — “always prioritize data quality over parameter optimization.” You can pick the perfect rank, the perfect learning rate, the perfect base model, and still ship a worse fine-tune than a team that obsessed over 5,000 clean examples. This lesson is about the part of fine-tuning that actually decides the outcome — and the data failure modes that look fine in a spreadsheet while they quietly poison the run.
Mechanically, supervised fine-tuning (SFT) maximises the log-likelihood of (prompt, response) pairs from your curated set. As mlabonne frames it, SFT “turns base models into helpful assistants … capable of answering questions and following instructions,” because the model learns to structure answers and reactivate a subset of knowledge it already learned in pre-training. That last clause is the senior insight: SFT is mostly teaching behaviour and format over latent knowledge, not stuffing in new facts — which is exactly why the decision lesson sends knowledge problems to RAG, not SFT.
The format layer matters as much as the content, and juniors skip it. Every (prompt, response) row is rendered through a chat template — ChatML, Alpaca, Llama-style — that inserts the special tokens (roles, turn markers, BOS/EOS) the model was post-trained to expect. Get the template wrong — a missing EOS, the wrong role tag, training on the prompt tokens as if they were the response — and the model learns to never stop, or to echo prompts, or to drift from the base model’s expected structure. Interview angle. “What’s a non-obvious way an SFT run silently fails?” → a mismatched or hand-rolled chat template; the loss looks healthy, the model is quietly broken. Use the tokenizer’s built-in template, don’t reinvent it.
The canonical alignment corpus is far smaller than people guess, and the numbers are worth memorising because interviewers use them to puncture the “more data is better” reflex. InstructGPT’s SFT stage used about 13k training prompts; the reward-model stage about 33k; the PPO stage about 31k — and the whole thing was produced by just 40 contractor labellers working to a clear rubric. Two implications: tens of thousands of high-quality demonstrations sit at the upper edge of what a single team can label, and that volume was enough to flip a 175B model’s behaviour so hard that the 1.3B variant beat it on preference.
code
1THE INSTRUCTGPT CORPUS (Ouyang et al., 2022) -- smaller than you think23 Stage ~Training prompts Produced by4 ----------------- ----------------- ---------------------------5 SFT ~13,000 40 contractor labellers6 Reward model (RM) ~33,000 labeller rankings of outputs7 PPO ~31,000 prompts only (RL rollouts)89 Lesson: a few tens of thousands of CLEAN, well-rubric'd demonstrations10 flipped a 175B model's behaviour. Quality + a clear rubric > raw volume.
The Guanaco counterpoint (from the QLoRA paper, L3) makes the quality-over-quantity case even sharper: Guanaco fine-tuned on a relatively tiny slice of carefully curated OASST1 data and reached 99.3% of ChatGPT on the Vicuna benchmark, while a 7B Guanaco outscored a 26GB Alpaca model by more than 20 percentage points. Alpaca had vastly more (synthetic, noisier) data. The smaller, cleaner set won. Interview angle. “How much data do I need?” is a trap — the strong answer is “enough clean examples that cover the task distribution; quality and coverage decide it, not a row count,” and you cite Guanaco/InstructGPT as the evidence that small-and-clean beats large-and-noisy.
The QDC frontier — quality, diversity, complexity
The most useful framework for reasoning about SFT data is the QDC generalisation frontier (Havrilla et al., 2024), because it refuses the lazy “just maximise quality” answer. The finding: quality is essential for in-distribution generalisation (clean examples on the task you’ll see), diversity is essential for out-of-distribution generalisation (coverage of the cases you didn’t anticipate), and complexity benefits both (harder, multi-step examples teach more per row). Crucially they trade off: high-quality synthetic corpora tend to be less diverse, and highly diverse synthetic corpora tend to be lower quality. So you tune the mixture, you don’t maximise one axis.
code
1THE QDC FRONTIER (Havrilla et al., 2024) -- you tune the mix, not one axis23 Axis Drives If you over-optimise it4 ---------- ---------------------------- ----------------------------5 Quality in-distribution accuracy low diversity -> brittle OOD6 Diversity out-of-distribution generaliz. lower avg quality per example7 Complexity BOTH (harder = more per row) too hard -> noisy labels, drift89 Synthetic tension: high-Q sets skew low-D; high-D sets skew low-Q.10 Operational rule: filter for Q (judge), score for D (embedding spread),11 control C by prompt design. Measure all three; don't maximise one.
A second QDC finding sets up the whole alignment lesson (L4): RLHF-aligned models generalise better than SFT-only models but exhibit worse sample diversity. Alignment improves in-distribution accuracy at the cost of output variety — a recurring tax worth flagging at design time, and the reason rejection-sampling and seed-diversity tricks exist. The senior framing for this lesson: your SFT mixture is a deliberate point on the QDC frontier, chosen for your task’s in-distribution-vs-OOD needs — not a pile of whatever data you could scrape.
Constructing an instruction dataset — the senior recipe
The canonical interview task: “build a 10,000-example SFT dataset for a legal-contract-review assistant — what does ‘good’ look like and how do you verify it?” The strong answer is a four-step pipeline, not “find 10k contracts and train.” (1) Source diversification — prefer human-written instructions paired with model responses over model-only synthetic data, because human prompts carry richer intent signal. (2) Coverage audit — stratify by document type, clause category, and edge cases (multi-party agreements, amendments, indemnity), so the set spans the real distribution. (3) Quality filtering — have a domain expert review a random sample, compute inter-annotator agreement, drop examples below threshold. (4) Decontamination — exclude any text that appears in your eval set, or your numbers are a lie.
code
1BUILD AN SFT SET (the four-step recipe the rubric rewards)23 1. SOURCE human-written prompts > model-only synthetic (richer intent)4 2. COVERAGE stratify: doc types, clause categories, edge cases5 -> this is the DIVERSITY axis; don't let one type dominate6 3. QUALITY FILTER expert-review a sample, compute inter-annotator agreement,7 drop below threshold; LLM-as-judge to scale the filter8 4. DECONTAMINATE remove any text overlapping the eval set (leakage = fake wins)910 Disagreement handling: Cohen's kappa / Krippendorff's alpha; re-label conflicts11 with a senior third reviewer; DROP examples that never converge.
The annotator-agreement step is where seniors separate from juniors. “What if your annotators disagree?” is a standard probe; the answer is a number and a process: compute Cohen’s kappa (or Krippendorff’s alpha) to quantify agreement, re-label the conflicts with a senior third reviewer, and drop examples that never converge — a label two experts can’t agree on is noise you’d be training the model to imitate. This is the same discipline RLHF preference data needs (a documented interview probe is literally “your preference data has low annotator agreement — how do you ensure quality?”), and the answer transfers directly.
Interview angle. The follow-up that catches people: “how much data is enough?” after you’ve described the pipeline. Don’t bite — the strong move is to restate quality-over-quantity with the Guanaco/InstructGPT evidence and say you’d start with a few thousand clean, stratified examples, train, eval, and then let the error analysis tell you which slices need more data. Adding rows blindly is how you slide down the QDC frontier into low diversity or contamination. Data work is iterative and eval-driven, not a one-shot collection.
Synthetic data — done right, with judge filtering
When human data is too expensive or too scarce, you generate it — but synthetic data is where the QDC frontier bites hardest, so it needs guardrails. mlabonne’s recipe: generate “instruction-response pairs based on seed data using frontier models,” with the deliberate constraint of “designing diverse seed tasks and effective system prompts,” and use “reward models and judge LLMs” for “fine-grained, customizable quality control.” The two techniques to know by name: Evol-Instruct (the WizardLM family) iteratively rewrites a seed instruction into harder variants along axes like deepening, concretisation, reasoning, and breadth; GLAN generalises this to taxonomy-driven synthesis and its 7B variant beat WizardLM-13B on most difficulty cohorts.
code
1SYNTHETIC DATA PIPELINE (generate, then PROVE quality + diversity)23 seed tasks (diverse, hand-written)4 | frontier model generates instruction-response pairs5 v6 Evol-Instruct: rewrite each seed harder (deepen / reason / broaden) [C up]7 |8 v QUALITY GATE: LLM-as-judge (or reward model) filters low-Q rows [Q]9 |10 v DIVERSITY GATE: sentence-transformer embedding spread; drop near-dupes [D]11 |12 v DECONTAMINATE against eval set, then mix with human data13 Failure if skipped: generator over-prefers ONE style -> diversity collapse.
The QDC warning is the operational rule for synthetic data: a generator that over-prefers one style collapses the diversity axis — every example reads the same, and the model overfits to that voice and fails OOD. So every synthetic corpus needs two gates, not one: judge-LLM (or reward-model) filtering for quality (Q), and sentence-transformer embedding spread for diversity (D), with complexity (C) controlled by the seed/prompt design (Evol-Instruct). The WizardLM result proves synthetic data can beat human data on specific benchmarks — but only when complexity is deliberately escalated and quality/diversity are measured, not assumed.
The data failure modes that silently wreck a fine-tune
These are the bugs that don’t throw — the loss curve looks textbook, the model trains, and the damage shows up weeks later as “it got worse” or “the eval numbers were a mirage.” (1) Contamination / leakage: eval examples (or near-duplicates) in the training set inflate your scores; you decontaminate by exact and fuzzy match against the held-out set before training, every time. (2) Format / template drift: a chat template that doesn’t match the base model’s, a missing EOS so the model never stops, or training on prompt tokens — all silent, all corrupting. (3) Diversity collapse: a synthetic generator (or an over-filtered human set) that narrows to one style, killing OOD generalisation. (4) Label noise: examples two experts can’t agree on, taught to the model as ground truth.
code
1DATA FAILURE MODES (silent -- the loss curve looks fine)23 Failure Symptom Defense4 ------------------ ------------------------------ ----------------------5 Contamination eval scores too good to be true exact+fuzzy dedup vs eval6 Template drift model never stops / echoes promp use tokenizer's template;7 mask prompt tokens; keep EOS8 Diversity collapse great in-dist, brittle OOD embedding-spread gate (D)9 Label noise confident wrong answers kappa; drop non-converging10 Catastrophic forget general skills regress (L1) replay 5-15% general data1112 None of these throw an error. Only a clean, decontaminated eval catches them.
The reason these are uniquely dangerous is the same as in the RAG track’s silent-recall-decay story: no request errors, healthy metrics, real damage. A contaminated eval is the worst because it actively misleads — you ship a fine-tune you believe is +8% that is really −2%, and you find out from users. The senior discipline: treat the held-out eval set as sacred (built before training, never touched, decontaminated against the training set), and re-run a general eval alongside the task eval to catch the forgetting from L1. Interview angle. “How do you know your fine-tune actually helped?” → a decontaminated, held-out task eval plus a general-capability eval, compared against the same base-model pipeline — never “it looks better in a few examples.”
You’re told you have 2M scraped instruction pairs and a competitor shipped a strong model on ~8k curated ones. Where do you start?
ATrain on all 2M — more data dominates curationBCurate a clean, stratified subset, judge-filter and decontaminate it, train, eval, then add data where error analysis says it’s thinCUse the 2M as-is but raise the learning rate to absorb it faster
Your synthetic SFT set scores great on your in-distribution eval but the model is brittle on slightly-unusual real queries. Most likely cause?
AThe model is too small — scale it upBDiversity collapse — the generator over-preferred one style, so the set lacks OOD coverage; add a diversity gate (embedding spread) and diversify seedsCQuality is too low — filter harder with the judge
Two expert annotators persistently disagree on ~12% of your legal SFT labels. What’s the senior move?
AKeep all examples — disagreement adds useful varietyBAverage the two labels into oneCQuantify with Cohen’s kappa, re-label conflicts with a senior third reviewer, and drop examples that never converge
Your fine-tune reports a big jump on the task eval, but you suspect the win is inflated. What’s the most likely culprit and the check?
AOverfitting — train for fewer epochsBContamination — eval examples (or near-duplicates) leaked into training; decontaminate with exact + fuzzy match against the held-out set and re-measureCThe judge model is too lenient
SFT-data rounds test whether you treat the dataset as the product — with a construction pipeline, measured quality/diversity/complexity, and a healthy paranoia about silent failures — or as a pile of rows. The single highest-signal move is refusing the “more data is better” bait and replacing it with the QDC frontier and a number (InstructGPT’s 13k, Guanaco beating Alpaca). Lead with the failure mode and the metric.
01“How much data for an SFT?” → quality + coverage decide it, not a row count; InstructGPT flipped 175B with ~13k clean prompts; Guanaco beat larger Alpaca.
02“Quality vs quantity?” → quality dominates parameter tuning; a small clean set beats a large noisy one — measure Q/D/C, don’t maximise volume.
03“What is the QDC frontier?” → quality → in-distribution, diversity → OOD, complexity → both; synthetic Q and D trade off, so you tune the mix.
04“How do you build a 10k SFT set?” → diversify sources, stratify coverage, quality-filter (kappa + judge), decontaminate against the eval set.
05“Annotators disagree — what now?” → Cohen’s kappa, re-label with a senior reviewer, drop non-converging examples (noise, not diversity).
06“Synthetic data — how, safely?” → seed-diverse generation + Evol-Instruct for complexity, then a judge gate (Q) and embedding-spread gate (D); decontaminate.
07“A silent SFT failure?” → chat-template / EOS / prompt-masking bugs, or eval contamination — both look fine on the loss curve.
08“How do you know it helped?” → decontaminated held-out task eval + a general-capability eval vs the same base pipeline — never a few cherry-picked outputs.
Going deeper, the follow-ups that separate offers: “your preference/SFT data has low agreement — fix it” (kappa, escalate, drop — same answer as annotator disagreement); “defend using synthetic data to a skeptic” (cite WizardLM beating human baselines, but only with complexity escalation and Q/D gates); “how do you keep diversity while filtering for quality?” (name the tension — high-Q skews low-D — and gate the two axes separately rather than filtering ever-harder on quality alone); and “how does RLHF change the data story?” (it generalises better but reduces output diversity — the L4 tax). Always name the QDC axis a question is poking at.
Could you design an SFT-data pipeline (sources, stratification, quality/diversity gates, decontamination) and name three silent data failure modes?
New to itGetting thereConfident
Takeaways
The dataset is the fine-tune: quality dominates parameter optimisation — small-and-clean beats large-and-noisy.
Numbers to cite: InstructGPT ~13k SFT prompts / 40 labellers; Guanaco beat larger Alpaca by 20+ points.