Lesson 3 of 6 · 49 min

RAG vs fine-tuning vs API

The single most-tested AI-PM decision: prompt → RAG → fine-tune → custom, and the build-vs-buy call underneath it. What each approach actually changes, what it costs, when each wins, the named-company cases (Notion, Harvey, Bloomberg), and the decision tree interviewers want to hear.

The question you cannot fumble

Across every AI-PM interview source, one question appears more than any other: RAG vs fine-tuning vs prompting — when do you use each? It shows up in OpenAI-style technical rounds, Microsoft’s guides, and 2026 question banks, and the research is blunt: if you cannot answer it crisply in about 90 seconds, you fail the round regardless of other strengths. But it is not trivia — it is the core architecture decision that sets your product’s freshness, cost, trust, and time-to-market. This lesson gives you the decision tree and the named-company evidence to defend it.
Start with what each lever actually changes, because the weak answer conflates them. Prompting changes the input — same model, you steer it with instructions and examples (lesson 2). RAG (retrieval-augmented generation) changes the context — same model, but you fetch relevant text at query time and put it in the prompt, so the answer is grounded in data the model never trained on. Fine-tuning changes the weights — you retrain on focused examples so a behaviour is baked in. The one-line mechanism that lands in interviews: RAG keeps the model the same and changes what it sees; fine-tuning changes the model itself.
code
1WHAT EACH LEVER CHANGES (the distinction interviewers test)23  Lever        Changes        Best for                         Cost / effort4  ----------   ------------   ------------------------------   ---------------------5  Prompting    the INPUT      open-ended gen, format, tone     lowest (manual, instant)6  RAG          the CONTEXT    fresh/proprietary facts,         medium (data + retrieval7                              citations, access control          engineering)8  Fine-tuning  the WEIGHTS    consistent tone/format, a        high (data prep, compute,9                              narrow skill, lower per-call        retrain cadence)10                              cost at scale11  Custom train the MODEL      a domain moat (huge proprietary  highest (months; rarely12                              corpus, specialized vocab)          justified now)1314  Default order: PROMPT -> RAG -> FINE-TUNE -> CUSTOM. Each step gated by evals.
The ordering is the recommendation, and committing to it is what separates strong from weak answers. The default path taught by both OpenAI’s optimization guide and Microsoft’s curriculum is prompt first, then RAG, then fine-tune, custom-train only with clear evidence. Interviewers are explicitly listening for a default you commit to — hedging with “it depends” on every clause is a known weak-answer signal. The senior framing: pick the default, then name the specific evidence that would make you deviate (freshness, tone at scale, a cost cliff, a proprietary corpus). Interview angle. The strong opener is the hierarchy, not the definitions: “prompt first; RAG when I need proprietary or current info; fine-tune when RAG-plus-prompting has plateaued on tone, format, or cost; custom only with a real data moat.”

RAG: the default for knowledge that moves

RAG is the right tool whenever the user expects accuracy, citations, or freshness over data that is proprietary or changing. Its three structural advantages over baking facts into weights are the ones to recite: freshness (re-index when a document changes — no retrain), attribution (cite the retrieved chunk, which builds trust and is auditable), and access control (filter retrieval by who is allowed to see what — impossible if the facts are smeared across the weights). IBM frames RAG plainly as “connect the LLM to your proprietary data for accuracy.” For most enterprise features — support over a help center, search over a corpus, Q&A over internal docs — RAG is the default, not an advanced option.
The decisive test for RAG-vs-fine-tuning, from IBM, is worth memorising: if the answer changes because the world changes, fine-tuning is the wrong lever and RAG is right. A PM who fine-tunes on last quarter’s product catalog has bought a stale model that cannot cite, with a multi-week retrain lag baked in. The failure mode is silent: the model keeps answering confidently from frozen facts while reality moved. The corollary for cost (lesson 5): RAG’s hidden expense is the embedding bill — re-embedding a large corpus on every change is one of the most common surprise costs, so freshness has a price you schedule, not a free lunch.
Case study — Notion. Notion bought the model API but built the data substrate, because the moat was the corpus (millions of users’ docs), not model weights. As its data grew from 20 billion to over 200 billion block rows across roughly three years — hundreds of terabytes compressed — it built a pipeline (change-data-capture into a lake) that let it ship AI Search and embedding-based RAG, cut ingestion from over a day to minutes, and bank seven-figure infra savings. The PM lesson: the “build” decision here was about data infrastructure for retrieval, not training a model. The model was a buy; the proprietary data plumbing was the build.
There is also an ownership angle that makes RAG the modern default beyond the three structural wins. Because the knowledge lives outside the model, you keep control of it: you can correct a wrong answer by fixing one document (not retraining), you can prove provenance to a regulator or a customer via the citation, and you can delete a user’s data from the index on request — a real compliance requirement that is effectively impossible once facts are smeared across frozen weights. Interview angle. “A customer asks us to delete their data and stop the model answering from it — can we?” With RAG, yes (drop it from the index); with fine-tuned-in facts, not without a retrain. That asymmetry alone settles many enterprise designs in RAG’s favour.

Fine-tuning: buys style and unit-economics, not facts

Fine-tuning earns its (real) cost in three situations, and you should be able to name them cold. (1) Tone/format you cannot lock with prompts — a precise brand voice or output structure across millions of calls. (2) A narrow, repeated skill — one task done extremely well, where you have labeled examples. (3) Unit economics at scale — distilling the behaviour into a smaller, cheaper, faster model so you stop paying frontier prices on the hot path, and your prompts get shorter (logic moves from the prompt, billed every call, into the weights, paid once). What fine-tuning is not for: adding knowledge that changes. That is RAG, every time.
Case study — Bloomberg (the rare from-scratch case). BloombergGPT was a 50-billion-parameter model trained on a 700-billion-token corpus — about 363 billion tokens of English financial documents plus public data. Bloomberg’s justification was a genuine domain moat: deeply specialized financial vocabulary and a proprietary corpus no frontier API of the day covered. Crucially, that was March 2023, before long context windows and cheap retrieval made RAG viable for very large custom corpora. The PM takeaway is not “build a model” — it is that custom training is the historical exception justified only by a severe vocabulary gap plus a huge proprietary corpus, and is rarely the right bet now.
Case study — Harvey (when even RAG falls short). Harvey, with OpenAI, custom-trained a model on the equivalent of 10 billion tokens spanning all U.S. case law. OpenAI reported an 83% increase in factual responses vs GPT-4, and attorneys preferred the custom model’s answers 97% of the time because they were more complete and nuanced. The stated reason matters for your mental model: it “overcomes limitations found in RAG and fine-tuning via public APIs for complex, open-ended case-law research.” In other words, you go past API-level tuning only when the failure mode of RAG (retrieve-and-cite without deep synthesis) is itself the bottleneck and the trust requirement is existential (legal hallucination). That is a high bar, met by few products.

They compose: most real systems stack the levers

A word on RAG plus fine-tuning together, because the strongest answer refuses the false binary. Production systems routinely stack them: a model fine-tuned for your domain’s tone and format that also retrieves the current facts at query time. The split stays clean — fine-tuning carries the durable behaviour (how to sound, what shape to return, a specialized reasoning skill), retrieval carries the volatile knowledge (today’s docs, this user’s data). When an interviewer frames it as “RAG or fine-tuning,” the senior tell is to note that most real systems do both, then explain which job each one is doing. Treating them as mutually exclusive is a weak-answer signal the rubric flags.

Build vs buy: the moat test underneath it all

Zoom out and RAG-vs-fine-tune-vs-API is one face of a bigger decision: build vs buy. Marty Cagan’s durable frame — build when it is your core competency, buy when it is not — combines with Hatchworks’ 2026 trigger list to give a usable rule. Default to buy (wrap an API), and only build the differentiated layer when one of four triggers fires: requirements change weekly (your roadmap is the moat), vendors are immature or keep sunsetting what you depend on, the real problem is workflow and incentives (not software), or the integrations are far more complex than anyone admitted. Commoditizing capabilities — speech-to-text, generic translation, basic embeddings — are default-buy, because the leading APIs move faster than your team can.
The senior tier-2 move — the one that separates a strong build-vs-buy answer from a coin-flip — is optionality: start with a buy decision behind a clean abstraction layer (a thin model wrapper, an eval harness, a prompt-config) so you can swap in a custom model later without rewriting the product. This keeps switching costs near zero while you learn whether the capability is actually your wedge. Interview angle. The strong build-vs-buy answer names the commoditizing-vs-wedge axis, calls out proprietary data as the structural build trigger, and describes the buy-behind-abstraction middle path. The weak answer picks a side, treats cost as the only factor, and forgets the data moat entirely.
code
1BUILD-VS-BUY: DEFAULT BUY, BUILD ONLY ON A TRIGGER23  BUY (default) when...                  BUILD when... (Hatchworks 2026 triggers)4  -----------------------------------    ------------------------------------------5  capability is commoditizing fast       requirements change weekly (roadmap = moat)6  (STT, translation, basic embeddings)   vendors immature / keep sunsetting features7  it's outside your core competency      the real problem is workflow + incentives8  a clean API exists at your scale       integrations are the hidden complexity9                                         your proprietary DATA is a structural edge1011  Senior move: BUY behind a clean abstraction layer so a future BUILD costs ~zero.12  Three tiers: (1) wrap API  ->  (2) fine-tune on top  ->  (3) train from scratch.13  Default (1); jump to (2) only when prompts plateau on tone/format/cost; (3) only14  with a measurable trust or unit-economics bar that nothing else meets.
One more honest caveat the research stresses: the “build” call scales with stakes and reversibility. Hatchworks is talking about enterprise vendors with quarterly roadmaps; a single PM shipping an internal automation should be far more conservative. And the hidden risk across every choice is the same — without evals and a cost-per-task dashboard you ship a feature you cannot triage. The PM who pairs a build-vs-buy decision with an eval harness has a feedback loop; the one who pairs it with vibes is gambling. We make that harness concrete in the capstone.

Interview prep

This is the highest-frequency technical round for an AI PM, and it is scored on a two-tier rubric: tier 1 is the decision hierarchy (do you commit to a default and know when to deviate), tier 2 is the mechanics (do you know what each lever changes, its hidden costs, and how to keep optionality). Interviewers want a decision tree, not definitions, and they penalise “it depends” with no default. Lead with the hierarchy, ground it in a named case, and name the cost no one mentions.
  1. 01“RAG vs fine-tuning vs prompting — when each?” → prompt first; RAG for fresh/proprietary/cited facts; fine-tune when tone/format/cost has plateaued; custom only with a data moat.
  2. 02“One-line difference between RAG and fine-tuning?” → RAG changes what the model sees (context); fine-tuning changes the model (weights). RAG adds knowledge; fine-tuning changes behaviour.
  3. 03“The answer changes when the world changes — which lever?” → RAG; fine-tuning bakes in stale facts, can’t cite, and adds a retrain lag.
  4. 04“When is fine-tuning actually right?” → locked tone/format at scale, a narrow repeated skill, or distilling to a cheaper/faster small model for high volume — not for facts.
  5. 05“Build or buy this AI capability?” → default buy (commoditizing, non-core), build only on a trigger (weekly-changing reqs, immature vendors, workflow-is-the-value, integration depth, proprietary data).
  6. 06“How do you keep build-vs-buy reversible?” → buy behind a clean abstraction (model wrapper + eval harness) so swapping in a custom model later is near-free.
  7. 07“Could you do this with prompting alone?” → often yes — reach for the expensive lever last; default prompt → RAG → fine-tune, each gated by an eval.
  8. 08“What’s the hidden cost of RAG / fine-tuning?” → RAG: the embedding/re-index bill on every change; fine-tuning: data prep plus a retrain cadence and monitoring.
Going deeper, the follow-ups separate “read the framework” from “shipped the decision”: “who owns indexing and freshness for your RAG?” (operational depth — is the retrieval system actually shippable, and how often do you re-embed?); “when has buy become build for you?” (they want a real inflection, usually proprietary data or a cost cliff); “how do you handle vendor lock-in?” (the abstraction layer answer); and “most production systems combine all three — why?” (a fine-tuned model that also retrieves is common; they compose). Name a company case (Notion buy-the-model-build-the-data, Harvey custom-because-the-corpus, Bloomberg from-scratch-because-the-domain) to show you reason from evidence, not slogans.
articleRAG vs fine-tuning vs prompt engineering (clean side-by-side)IBMarticleThe Build vs Buy Framework in the Age of AI (2026 trigger list)HatchworksarticleCustomizing models for legal professionals (Harvey: +83% factual, 97% preferred)OpenAIarticleNotion: scaling data infrastructure for AI features and RAGZenML LLMOps DBarticleIntroducing BloombergGPT (50B params, 700B-token finance corpus)Bloomberg

Checkpoint

Your product answers questions over a help center that changes weekly, and legal requires every answer to cite its source. An engineer proposes fine-tuning monthly on the articles. Best stance?

AUse RAG — retrieve the current articles at query time for freshness, citations, and per-user access; fine-tuning would bake in stale, uncitable factsBFine-tune monthly so the articles live in the weights and the model is fasterCPrompt the base model to “always be accurate and cite sources”
Sign up free to answer and see why

Checkpoint

A high-volume classification feature works well on a frontier model but the per-call bill is unsustainable at your scale, and the task is narrow with plenty of labeled data. Most defensible move?

AAdd RAG to make it cheaperBKeep paying frontier prices — model choice shouldn’t be a product concernCFine-tune (or distill to) a smaller model on the labeled data to hit the same quality at a fraction of the cost and latency
Sign up free to answer and see why

Checkpoint

Leadership wants to “train our own model” to differentiate a new feature whose capability (summarizing meetings) is well served by frontier APIs. What is the senior recommendation?

ATrain a custom model now to own the capability end-to-endBBuy the API behind a clean abstraction layer; revisit building only if proprietary data, a cost cliff, or a vendor failure becomes a real triggerCFine-tune a base model on generic meeting transcripts to differentiate
Sign up free to answer and see why

Checkpoint

In a technical round you’re asked “RAG or fine-tuning?” and you’re tempted to say “it depends on the use case.” Why is that risky, and what’s the stronger frame?

AIt’s fine — “it depends” shows nuance and covers all casesBCommit to the default hierarchy (prompt → RAG → fine-tune → custom) and name the specific evidence that would make you deviateCAlways answer “fine-tuning,” since it changes the model most
Sign up free to answer and see why

Checkpoint

A vertical legal-research product finds that RAG retrieves the right cases but its answers lack the deep synthesis attorneys need, and hallucinated reasoning is an existential risk. What does the Harvey case suggest?

AAdd more documents to the RAG context until synthesis improvesBSwitch back to prompting the base model more firmlyCThis is the rare case to go past API-level tuning — a custom-trained model on the domain corpus — because RAG’s retrieve-and-cite failure mode is itself the bottleneck
Sign up free to answer and see why

Could you defend, in 90 seconds with a named case, the prompt → RAG → fine-tune → custom hierarchy and the build-vs-buy triggers underneath it?

New to itGetting thereConfident

Takeaways

  • Default order: prompt → RAG → fine-tune → custom; commit to it and name the evidence that makes you deviate.
  • RAG changes the context (adds knowledge: freshness, citations, access control); fine-tuning changes the weights (behaviour: tone, format, a skill, cheaper-at-scale).
  • If the answer changes when the world changes, it’s RAG — fine-tuning bakes in stale, uncitable facts.
  • Build vs buy: default buy (commoditizing, non-core); build only on a trigger, and keep proprietary data as the structural build signal.
  • Preserve optionality — buy behind a clean abstraction so a future custom build costs near-zero.
  • Named cases anchor the answer: Notion (buy model, build data), Harvey (custom because the corpus + trust), Bloomberg (from-scratch because the domain).

Next: agents and MCP — what they actually are, and the narrow conditions under which they beat a simple workflow.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.