Lesson 1 of 6 · 46 min

Fine-tune vs prompt vs RAG — the decision

The lever that costs the least and survives a model swap, the one that fixes knowledge, and the one that fixes behaviour — the cost/quality/latency triangle, the orthogonal-failure-modes mental model, and how to justify the choice in an interview without reaching for the most expensive tool.

Lesson 1 · Foundations

Fine-tune vs prompt vs RAG

The most expensive default in the room

When a feature underperforms, the reflex answer — “let’s fine-tune” — is usually the wrong first move, and the senior signal is knowing why. Prompting touches no weights and survives any model upgrade; RAG fixes a knowledge gap with citations and freshness; fine-tuning is the only durable fix for a behaviour gap. The single most repeated LLM-engineer interview question is some flavour of “fine-tune vs RAG vs prompt?” — and the people who lead with a feature list lose to the people who lead with a decision frame. This lesson builds that frame.
Start with the mechanism, because the three levers act on different parts of the system. Prompting reshapes the input distribution the model sees at inference — no weight changes, so the only cost is context-window tokens and the only engineering surface is the prompt string. RAG keeps the base model frozen and prepends retrieved evidence, so the model conditions its next-token distribution on fresh, attributable text at every query. Fine-tuning updates the weights so the desired behaviour becomes a prior the model already carries, which is what removes the per-query cost of a long prompt or a retrieved chunk. Three levers, three different objects: the input, the context, the weights.
The mental model that wins interviews: behaviour and knowledge are orthogonal failure modes. If the model doesn’t know something — your Q3 board deck, yesterday’s ticket, a regulation that changed last week — that is a knowledge gap, and no amount of fine-tuning reliably injects volatile facts (you’d retrain on every edit, lose citations, and can’t enforce per-user permissions). If the model knows enough but won’t act right — wrong tone, wrong format, ignores your tool schema, can’t do the domain’s reasoning even with a perfect prompt — that is a behaviour gap, and that is what fine-tuning durably fixes. A real system often needs all three at once: prompt for format, RAG for knowledge, fine-tune for the domain voice and skill.
Finetuning Open-Source LLMsSebastian Raschka

The decision tree — cheapest lever that closes the gap

The discipline is to climb the ladder, not jump to the top. Benchmark the best prompt-and-RAG configuration first, then invest in fine-tuning only for the residual error that is a stable behaviour the model cannot acquire in context. Prompting is the default baseline because it has the smallest footprint and is the only intervention that survives a model swap — when the provider ships a better base model next quarter, your prompt still works while your fine-tune has to be redone. The order below is the order of increasing cost and decreasing reversibility.
code
1THE DECISION LADDER  (climb from cheap/reversible to expensive/sticky)23  1. PROMPT      no weights, survives model swaps, zero infra4                 -> try first; fixes "missing instruction", format demo, tone hint5  2. RAG         frozen model + retrieved context6                 -> the failure is MISSING KNOWLEDGE: freshness, citations, ACL7  3. FINE-TUNE   update the weights8                 -> the failure is BEHAVIOUR the model can't reach in context:9                    consistent style/format, domain jargon, tool-call reliability,10                    sub-200ms latency (no long prompt), a capability gap1112  Rule: benchmark best prompt+RAG FIRST. Fine-tune only the residual that is a13  stable behaviour, not a fact. Most production stacks run all three together.
Why does a vanilla model fall short on behaviour even when it clearly “knows” the material? Because the base model’s prior is the average of the public web, not your house style or your domain’s conventions. The single most quoted evidence here is InstructGPT (Ouyang et al., 2022): outputs from the 1.3B aligned InstructGPT were preferred to outputs from the 175B unaligned GPT-3 — a 130x smaller model winning on human preference purely because its behaviour was tuned. The lesson for the decision: behavioural alignment is not something a bigger base model gives you for free, which is exactly why “just use a bigger model” is not a substitute for fine-tuning when the gap is behaviour.

The cost/quality/latency triangle — say the numbers

Interviewers reward candidates who put numbers on the triangle. Prompting adds context tokens to every call — a long few-shot prompt is paid on every request, forever, and output tokens cost ~3–5x input. RAG adds retrieval latency plus the chunk tokens to the context, and an indexing pipeline to build and keep fresh. Fine-tuning front-loads a training cost and then has zero extra per-query latency — a merged LoRA is bit-identical to a dense fine-tuned model (we prove this in L3), so a fine-tune that lets you drop a 2,000-token instruction prompt can pay for itself in inference savings at volume.
code
1THE TRIANGLE  (what each lever costs you)23  Lever        Touches   Fixes        Fixes        Per-query      Cold-start4               weights   knowledge?   behaviour?   latency cost   eng cost5  ----------   -------   ----------   ----------   ------------   ----------6  Prompting    no        weakly       moderately   + ctx tokens   low7  RAG          no        YES (+cite)  no           + retrieve     medium (index)8  Fine-tune    YES       indirectly   YES durable  ~0 (merged)    high (data+train)910  Anchor: GPT-4.1 fine-tune training ~$25 / 1M tokens -> a 20M-token run ~= $50011  of training alone, before eval. RAG re-indexes for free; prompting is just tokens.
That ~$25 / 1M training-token figure (hosted GPT-4.1 fine-tuning) is worth memorising because it grounds the build-vs-buy half of the conversation: a 20M-token SFT run is ~$500 of training cost alone, before the data labelling and eval that dominate the real bill. Interview angle. When asked “would you fine-tune for this?”, the strong answer names a cost number and a cadence: “the knowledge rotates weekly, so a weekly retrain at ~$500+ is wasteful versus a re-index that’s effectively free — RAG. But if they also need a fixed schema and refusal pattern, I’d fine-tune that behaviour once and keep RAG for the facts.” That is the cost-freshness-control triangle stated as a decision.

The canonical interview scenario, worked

The question almost every LLM-engineer loop asks, verbatim: “A product team wants a Q&A system over their internal 200-page knowledge base. They also want it to cite sources and update weekly. Prompt, RAG, or fine-tune — walk me through your decision.” The strong answer clarifies the goals, then reasons from the failure modes. The knowledge half is volatile and must be citable — that is textbook RAG: the corpus lives outside the model, re-indexes cheaply, and chunk-level attribution gives you citations a fine-tune cannot natively produce. Fine-tuning the recall half is rejected because retraining weekly is wasteful and still can’t cite.
But the senior answer doesn’t stop at “RAG.” It surfaces where fine-tuning does earn its place in the same system: if the base model refuses to discuss the domain, or you need a strict house format/tone, or the assistant must reliably emit your tool schema, those are behaviour gaps RAG can’t touch — so the production pattern is the hybrid: fine-tune the floor of competence and voice, wrap it in RAG for the latest documents, and externalise hard rules through the prompt template. The fine-tune sets the floor (never invents an order code); RAG sets the ceiling (today’s inventory and policy). Naming the layered choice — and why each layer is there — is what separates a 5 from a 3 on the rubric.
The fine-tune sets the floor of competence; RAG sets the ceiling of freshness; the prompt carries the rules of this turn. “We tried RAG and it didn’t work” is usually “we used RAG for a behaviour problem” — and the reverse for fine-tuning.

When fine-tuning genuinely wins — and the catastrophic-forgetting trap

Fine-tuning is the right tool when the success criterion is a stable behavioural contract: a consistent tone, a domain-specific output format, tool-calling reliability, a sub-200ms latency target that forbids long prompts, or a capability the base model can’t express even with optimal context. The Guanaco result (QLoRA, L3) is the proof point — a 7B model fine-tuned on a small, carefully curated set reached 99.3% of ChatGPT on the Vicuna benchmark from a single GPU in 24 hours. That is behaviour bought cheaply, exactly where prompting and RAG would have stalled.
The failure mode every senior must be ready for: catastrophic forgetting. Aggressive domain fine-tuning pulls weights away from general competence — the model gets better at your tickets and quietly worse at everything else. The interview probe is “your fine-tune regressed on a previous capability — diagnose and fix.” The strong answer treats it as a train-data composition problem, not an architecture one: mix in a small fraction (~5–15%) of replay examples from the original/general distribution to anchor general skills, lower the learning rate, and bound the update (LoRA rank, fewer epochs). And you only see forgetting if you held out a general eval (MMLU, MT-Bench) alongside your task eval — measure both, every run.
code
1WHEN EACH LEVER WINS  (map the failure to the tool)23  Symptom                                   -> Lever4  ---------------------------------------------  --------------------------5  "doesn't know our latest docs / prices"   -> RAG (freshness + citations)6  "needs to cite its sources"               -> RAG (chunk-level attribution)7  "wrong tone / format, ignores schema"     -> fine-tune (behaviour)8  "can't do our domain's reasoning"         -> fine-tune (capability)9  "too slow / prompt too long at volume"    -> fine-tune (drop the prompt)10  "just needs a clearer instruction"        -> prompt (cheapest, reversible)11  "regressed on general tasks after FT"     -> data composition: replay 5-15%1213  Behaviour gap -> weights.  Knowledge gap -> retrieval.  Instruction gap -> prompt.
Interview angle. A favourite follow-up: “how would you measure forgetting quantitatively?” The answer that lands: hold out the original instruction suite plus a general benchmark, run both before and after, and gate the deploy on “MMLU must not drop more than ~2 points” (a capability floor) while the task metric clears its bar (a task ceiling). If both pass, ship; if the task ceiling needs a higher learning rate that breaks the floor, that tension is the real decision — and naming it as a multi-objective tradeoff is the senior signal.

How real teams sequence the levers

The pattern across documented production systems is the same staircase, applied pragmatically. Teams start with prompting to validate the use case (zero infra, instant iteration). They add RAG the moment the failure is “doesn’t know our stuff,” because re-indexing beats retraining for anything that changes. They reach for fine-tuning last and narrowly — for the persistent voice, the format, the latency win of dropping a long instruction, or a capability gap — and they keep RAG underneath it for the facts. The instinct to protect: a fine-tune is a project (data curation, training, eval, a re-do on the next base model), so it should clear a bar that prompting and RAG demonstrably could not.
There is also a quieter operational reason to prefer the cheaper levers: maintenance. A prompt is one string in version control. A RAG index is an ingestion pipeline you already need for freshness. A fine-tune adds a training pipeline, a dataset you must keep clean, an eval gate, and a standing dependency on whatever base model you tuned — when the provider deprecates it or ships a better one, you re-pay the whole cost. The senior framing: fine-tuning is not just a quality decision, it’s a lifecycle commitment, and that commitment is the thing juniors under-price.
repomlabonne/llm-course — the fine-tuning vs RAG vs prompting framingMaxime LabonnearticleRAG vs fine-tuning vs prompt engineeringIBMpaperTraining language models to follow instructions (InstructGPT)Ouyang et al. (arXiv)

Checkpoint

A team wants a support assistant that answers from a knowledge base that changes weekly and must cite the document it used. What’s the strongest primary lever?

AFine-tune the model weekly on the latest knowledge baseBRAG — retrieve at query time for freshness and chunk-level citations; fine-tune only if a behaviour gap remainsCA larger base model so it memorises more of the knowledge base
Sign up free to answer and see why

Checkpoint

A model answers factually fine but won’t reliably emit your strict JSON tool schema and ignores your house tone, even with a careful prompt and examples. Best durable fix?

AFine-tune on examples of the correct schema and tone — this is a behaviour gapBAdd the company knowledge base via RAGCSwitch providers and hope the new model formats better
Sign up free to answer and see why

Checkpoint

Your fine-tune lifted task accuracy on support tickets but MMLU dropped 6 points. What does a senior do?

AShip it — MMLU is irrelevant to the support taskBConclude fine-tuning was the wrong call and revert to promptingCTreat it as catastrophic forgetting: add ~5–15% general-distribution replay, lower LR / bound the update, and gate on a capability floor (e.g. MMLU −2 max)
Sign up free to answer and see why

Checkpoint

A latency-critical, high-volume classifier currently uses a 2,000-token few-shot prompt per call. The labels are stable. What’s the strongest optimisation?

ACache the prompt and call it a dayBMove to RAG so the examples are retrieved instead of in the promptCFine-tune a small model on the labelled examples and drop the few-shot prompt entirely
Sign up free to answer and see why

Checkpoint

An interviewer asks: “Why not just fine-tune everything — wouldn’t a model that knows our domain and behaves right be best?” Strongest response?

AAgree — one fine-tuned model is simpler than maintaining three systemsBSeparate the failure modes: fine-tune behaviour, RAG the volatile knowledge, prompt the per-turn rules — and only fine-tune the residual that prompting+RAG can’t reachCFine-tune for knowledge and prompt for behaviour
Sign up free to answer and see why

Interview prep

This is the most reliably asked topic in the whole track, and it tests one thing: do you reach for the cheapest lever that closes the gap, or do you default to the most expensive tool? Interviewers want a decision frame kept in reserve (cost, freshness, control) that you can drop into the moment they probe — and an honest tradeoff with a number attached. Lead with the failure-mode classification (knowledge vs behaviour vs instruction), then pick, then name what it costs and how it’s maintained.
  1. 01“Fine-tune vs RAG vs prompt?” → classify the failure: volatile/attributable knowledge → RAG; stable behaviour/format/skill → fine-tune; missing instruction → prompt. They compose.
  2. 02“Why not fine-tune the knowledge base?” → stale on every edit, can’t cite, can’t enforce per-user ACL, and you re-pay on each base-model upgrade.
  3. 03“What does fine-tuning actually buy?” → durable behaviour (tone, format, tool-call reliability, a skill) and a latency/cost win from dropping a long prompt.
  4. 04“Cost of a fine-tune?” → training ~$25/1M tokens hosted (a 20M run ≈ $500) before the data + eval that dominate; RAG re-indexes ~free; prompting is just tokens.
  5. 05“Evidence behaviour ≠ scale?” → InstructGPT: a 1.3B aligned model was preferred over 175B GPT-3 — alignment, not size, drove preference.
  6. 06“Catastrophic forgetting — diagnose + fix?” → it’s data composition: replay 5–15% general data, lower LR, bound the update; detect with a held-out general eval (MMLU/MT-Bench).
  7. 07“Higher human-eval but lower MMLU — did it help?” → multi-objective: set a capability floor (MMLU −2 max) and a task ceiling; ship only if both hold.
  8. 08“200-page KB, must cite, updates weekly?” → RAG primary (freshness + citations); fine-tune only a residual behaviour gap (refusals/format); prompt the rules.
To go deeper, expect the follow-ups that separate “read a blog” from “shipped one.” “What if the base model refuses to discuss the domain?” (that’s behaviour → SFT on domain Q&A; RAG alone won’t unblock a refusal). “How do you eval the system with no ground truth?” (LLM-as-judge on faithfulness + a 100-sample human spot-check — the eval lesson’s playbook). “When is the hybrid worth the complexity?” (when you have both a behaviour gap and a freshness need — fine-tune the floor, RAG the ceiling). In every case, classify the failure mode out loud before you name a tool.

Could you take “fine-tune vs RAG vs prompt for X?” and answer with a failure-mode classification, a pick, a cost number, and the maintenance tradeoff?

New to itGetting thereConfident

Takeaways

  • Three levers, three objects: prompt (input), RAG (context/knowledge), fine-tune (weights/behaviour).
  • Climb the ladder: prompt → RAG → fine-tune; benchmark prompt+RAG before investing in a fine-tune.
  • Behaviour ≠ knowledge: fine-tune fixes style/format/skill; RAG fixes volatile, citable facts. They compose.
  • Put numbers on it: hosted fine-tune ~$25/1M tokens; InstructGPT 1.3B beat 175B on preference (alignment, not size).
  • Catastrophic forgetting is a data-composition problem: replay 5–15%, bound the update, gate on a capability floor.
  • A fine-tune is a lifecycle commitment (data, eval, base-model dependency), not just a quality knob.

Next: the dataset is the product — how to construct SFT data, why quality beats quantity, and the data failure modes that wreck a fine-tune.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.