Lesson 1 of 6 · 46 min

LLMs for PMs

What an LLM actually is — a probabilistic next-token predictor with finite working memory you rent by the token — with no math. Tokens, context, why output varies, what it can and cannot reliably do, and the interview questions that probe whether you treat it as a magic box.

Why “add AI” is not a spec

The single most expensive mistake an AI PM makes is treating “the model” as one undifferentiated thing you sprinkle onto a feature. An LLM is a probabilistic next-token predictor with finite working memory you rent by the token and latency that scales with how much it writes. Almost every scoping, pricing, and reliability decision — and almost every technical-round interview question you will get — falls directly out of those three facts. This lesson is the ground floor: no math, but the exact mechanisms a senior PM is expected to reason from.
Here is the whole mechanism in one breath, because interviewers open with it. The model reads your text as tokens (sub-word pieces), and for each step it scores every token in its vocabulary (~100k–200k entries), turns those scores into probabilities, samples one, appends it, and repeats — until it emits a stop token or hits a length cap. Everything we call “reasoning,” “following instructions,” or “knowing things” is emergent behaviour on top of that one loop. The first product consequence: the output is a sample from a distribution, not a return value. The same prompt can — and will — produce different text.
Why does it “know” anything at all? The weights are a lossy compression of its training data at a fixed point in time. That single sentence carries three product implications you will lean on all track. It has a knowledge cutoff (it has no row for your Q3 board deck, yesterday is genuinely unknown to it). It cannot cite by default (it is reconstructing plausible text, not looking anything up). And it will confidently invent when the right answer was never in its weights — what we call a hallucination. None of these are bugs you patch; they are properties you design around. Interview angle. “Why does the model hallucinate on our internal data?” The strong opener is exactly this: the data was never in the weights, so the model fills the gap with fluent plausibility — which is why the fix is retrieval (lesson 3), not a bigger model.
Andrej Karpathy — Intro to Large Language ModelsAndrej Karpathy

Tokens: the unit of cost, latency, and limits

Models do not see words or characters — they see tokens. The rule of thumb worth memorising: for English, ~4 characters ≈ 1 token, and roughly 0.75 words per token, so ~750 words is about 1,000 tokens. Code, JSON, tables, and non-English text tokenize less efficiently (more tokens per character), which quietly inflates cost. Tokens are the unit of everything commercial: pricing is per token, latency scales with tokens, and the context limit is counted in tokens. A PM who cannot estimate tokens cannot estimate cost — so this is the first number you learn to do in your head.
Two pricing facts a senior PM keeps in their pocket. First, output tokens cost meaningfully more than input tokens — typically 3–8× more, depending on the model — because generating each output token is sequential and compute-heavy, while reading the prompt is parallelised. We will see exact numbers in lesson 5; for now, the instinct is “the model writing a lot is what gets expensive.” Second, a long, stable prompt prefix can be cached for up to ~90% off its input cost. Put together: the cheapest feature has a big cached prompt and a short output — the exact opposite of the naive intuition that “shorter prompts are cheaper.”
code
1TOKENS, IN PLAIN NUMBERS (English text, rough rules of thumb)23  1 token        ~= 4 characters ~= 0.75 words4  1,000 tokens   ~= 750 words ~= 1.5 pages of prose5  a 20-page PDF  ~= 12,000-15,000 tokens of input6  a 1-paragraph answer ~= 100-200 output tokens78  Cost asymmetry to remember:9    output tokens cost ~3-8x input tokens (the model WRITING is what's pricey)10    a stable prompt prefix can be cached for up to ~90% off its input cost1112  So: cheapest feature = big CACHED prompt + SHORT output. Counterintuitive but real.

The context window is working memory — and more is not better

The context window (say 200k tokens) is the model’s working memory for one call. The system prompt, the conversation history, any retrieved documents, the tool definitions, the tool results, and the output all share that one budget. It is not long-term memory: nothing persists between calls unless your product puts it back in. A chatbot does not “remember” your last session — your code re-sends the history every turn, and you pay for it every turn. Treating the window like a database is the root of a whole class of bugs and surprise bills.
Two empirical facts every AI PM should be able to name, because they overturn the intuition that “more context is always better.” (1) Lost in the Middle (Liu et al., 2023): recall is U-shaped — facts placed at the start or end of a long context are recalled reliably, facts buried in the middle are frequently missed. (2) Context Rot (Chroma, 2025): answer quality degrades as the input grows even in frontier models, and it is a reasoning degradation, not just a retrieval miss. The practical upshot you will repeat in scoping debates: stuffing more into the prompt can make answers worse, not just slower and more expensive.

Why the same input gives different output

Because the model samples from a probability distribution, the same prompt can return different text on different runs. There is one knob a PM should understand by name: temperature. Low temperature makes the model greedier — it leans hard toward the single most likely next token, producing consistent, repetitive, “safe” output. High temperature flattens the distribution — more varied, more “creative,” more prone to wander. The product mapping is clean: extraction, classification, and anything you parse → low temperature; brainstorming and multiple drafts → high temperature; a balanced assistant → somewhere in the middle.
The trap that catches senior people too: temperature 0 is not a determinism guarantee. It makes decoding greedy, but real systems still vary run-to-run because of batched inference, floating-point quirks on GPUs, mixture-of-experts routing, and — the big one for product — the provider silently updating the model under you. Your downstream product must never assume byte-identical output across calls. Interview angle. “You set temperature 0 and your JSON parser still crashes intermittently in prod — why?” The answer is exactly this: greedy ≠ deterministic, so the fix is to constrain the output to a schema and validate it, never assert output == expected_string.
  1. 01Extraction / classification / structured fields → low temperature (you want the single most likely answer, repeatably).
  2. 02Balanced assistant / Q&A → mid temperature (some variety, still grounded).
  3. 03Brainstorming / creative / multiple drafts → high temperature (explore the distribution on purpose).
  4. 04Never promise users or downstream code exact reproducibility — even at temperature 0, design for variation (schemas, evals, graceful UI).

What it can vs cannot reliably do (the capability tiers)

A senior AI PM carries a shared, accurate map of capability — otherwise you ship features the model cannot do, or you waste quarters solving problems it already covers. MIT Sloan’s 2025 inventory is a clean anchor: generative AI is reliably deployed for summarization, classification, drafting, transformation, and question-answering over context you provide; it is not reliable for high-stakes autonomous decisions, long-horizon reasoning chains, multimodal action without verification, or causal forecasting. Model progress is monotonic on benchmarks but lumpy on real workflows, so “the demo worked” is almost never evidence the feature works.
code
1CAPABILITY TIERS (a PM mental model for matching task to risk)23  Tier 1  ~90%+ reliable today    summarize, classify, extract, draft, Q&A over4                                  PROVIDED context. Ship with light scaffolding.5  Tier 2  ~60-85%, context-       synthesize across many docs, plan, multi-step6          and domain-dependent    reasoning. Needs retrieval + evals + a human path.7  Tier 3  genuinely risky         reliable multi-step AGENCY, causal forecasting,8                                  long-horizon execution. Needs hard HITL scaffolding.910  The PM move: match the capability TIER to the COST OF ERROR in the workflow.11  A tier-mismatch (tier-3 experience, tier-1 scaffolding) is the product bug.
The discipline this buys you is to scope around the cost of error, not the cleverness of the demo. Mapping a feature to a tier and an error-cost (reversible / expensive / irreversible / regulated) is, per the research, the single most durable move a senior AI PM makes — it converts a vague “let’s add AI” conversation into a tractable engineering scope. Interview angle. “How would you launch a feature whose predictions are probabilistic and sometimes wrong?” Strong answers name the failure modes (hallucination, calibration error, stale knowledge), then design the customer experience around them (confidence cues, citations, an easy correction path, escalation) — they do not apologise for the model, they engineer the product around it.

Case studies: production teams taming these exact facts

Notion cut its AI chat latency from ~2s to ~350ms (about 4×) for 100M+ users by serving a smaller, fine-tuned model on dedicated infrastructure rather than a frontier model on the hot path — their framing, worth quoting in any room, is that “latency is perceived as quality.” GitHub Copilot deliberately caps prompts (around 6,000 characters) to keep fast models inside its latency envelope, and earned measured wins from context rather than model size. These are the same three facts — tokens cost, the window is scarce, latency tracks output — turned into product decisions.
And a cautionary one on the silent-failure tax. Klarna publicly claimed an AI agent was doing the work of 853 employees, then walked the AI-only customer-service strategy back after quality and customer-experience metrics fell, re-hiring humans. The trap is textbook and worth internalising as a PM: the demo succeeded, the long tail (escalated complaints, regulated disputes, emotionally charged cases) was underweighted, and the “replace headcount” framing masked a missing product measurement. The lesson: frame AI features as augmentation with a measurable quality bar, not headcount replacement, and budget for the long tail before you scale.
Latency is perceived as quality. — the recurring lesson across Notion, Copilot, and Perplexity: the team that shrinks what hits the model before the model runs wins on speed, cost, and trust at once.

Interview prep

Foundational AI-PM screens test whether you treat the model as a system with real mechanisms and limits — or as a magic box. Interviewers probe four things here: can you walk the generation loop in plain language, do you think in tokens (cost and latency), do you know why output varies and how to design for it, and can you state what the model cannot reliably do and scope around it. Lead with the mechanism, then the product implication.
  1. 01“Walk me through what happens when we call an LLM.” → tokenize → score every token → sample one → repeat to a stop token; the output is a sample, not a return value.
  2. 02“Why does it hallucinate on our data?” → the data was never in the weights (a lossy snapshot at a cutoff); it fills the gap with fluent plausibility — fix with retrieval, not a bigger model.
  3. 03“Why should an AI PM care about tokens?” → they drive cost (billed per token, output ~3–8× input), latency (scales with output), and the context limit — they are the product’s unit economics.
  4. 04“Is temperature 0 deterministic?” → no — greedy ≠ deterministic; batching, hardware, MoE routing, and silent provider updates vary it, so design with schemas and evals.
  5. 05“The document is bigger than the window — what do you do?” → retrieve the few relevant passages (RAG) or map-reduce summarise; not “use the biggest-context model.”
  6. 06“Is more context always better?” → no — Lost-in-the-Middle and Context Rot mean quality can drop as you fill the window; curate and order it.
  7. 07“What can’t this model reliably do?” → high-stakes autonomous decisions, long-horizon reasoning, action without verification, causal forecasting — match tier to cost of error.
  8. 08“How do you launch a probabilistic, sometimes-wrong feature?” → name the failure modes, then design CX around them (confidence cues, citations, correction path, escalation) and ship a quality metric.
To go deeper, expect the follow-ups that separate “read a thread” from “shipped one”: “how would you measure whether it’s good?” (you need an eval set and a quality metric, not vibes — the through-line of lesson 6); “the same prompt gives different answers — is that a bug?” (no, it is the medium; constrain and validate); “why not just wait for a bigger model to fix this?” (some failures are structural — tokenization, no row for private data, the window — and architecture beats scale for them); and “what’s the cheapest version of this feature?” (cached prompt, short output, smallest model that passes the eval — lesson 5). In every case, name the mechanism and the metric before proposing a fix.
repoGenerative AI for Beginners (Lesson 01: intro to generative AI & LLMs)MicrosoftpaperLost in the Middle: How Language Models Use Long ContextsLiu et al. (arXiv)articleContext Rot: how increasing input tokens degrades LLM performanceChroma ResearcharticleMachine learning and generative AI: what are they good for?MIT Sloan

Checkpoint

An engineer says your support assistant “forgot” a fact the user gave it three turns ago. A teammate proposes switching to a 1M-token-context model. What is the senior read?

AAgree — a bigger context window is the direct fix for forgettingBThe window is per-call memory — confirm the history is actually being re-sent and curated each turn before touching model sizeCRaise the temperature so the model attends to more of the conversation
Sign up free to answer and see why

Checkpoint

A stakeholder wants a feature that answers questions about your customers’ live account balances from the base model alone. Best framing to bring back?

AThe base model can answer once we prompt it more firmly to be accurateBRaise the temperature to 0 so the answers are exactCThe model has a knowledge cutoff and cannot see live private data — this needs retrieval/tool access to the balance, not the base model
Sign up free to answer and see why

Checkpoint

Your team wants to ship an AI feature that auto-approves insurance claims end-to-end with no human review. Using capability tiers, what is the right product stance?

AThis is a tier-3, irreversible, regulated workflow — scope it as assistive (draft/recommend) with a human approval gate, not full autonomyBShip it autonomously — claims are a well-defined task and the demo handled the examplesCUse a bigger model so it can be trusted to auto-approve
Sign up free to answer and see why

Checkpoint

A PM proposes pricing the feature by “number of requests” and is surprised the bill is dominated by long answers. What did the model miss?

ARequests are the right unit; the provider must be overchargingBOutput tokens are the expensive unit (~3–8× input) — cost tracks tokens written, so price and design around output length and cachingCSwitch to a model with a larger context window to reduce cost
Sign up free to answer and see why

Checkpoint

QA reports that the model’s JSON output “randomly” breaks the parser about once every few hundred calls, even at temperature 0. Best fix to spec?

AIt is a parser bug — JSON parsing is deterministic so the code must be wrongBLower the temperature below 0CConstrain the output to a schema (structured outputs) and validate, instead of assuming stable text
Sign up free to answer and see why

Could you explain to a non-technical exec, in plain language, what an LLM is and the three facts that drive its cost, limits, and reliability?

New to itGetting thereConfident

Takeaways

  • An LLM is a probabilistic next-token predictor — output is a sample, not a return value; design for variation.
  • Tokens are the unit of cost, latency, and limits; output costs ~3–8× input, and a stable prefix can be cached ~90% off.
  • The context window is scarce, ordered working memory that does not persist — more context can mean worse, slower, pricier answers.
  • Hallucination, knowledge cutoff, and no-citations come from the weights being a lossy snapshot — fix with architecture, not scale.
  • Match the capability tier to the cost of error; scope around failure modes, never around the demo.
  • Latency is perceived as quality — shrink what hits the model before the model runs (Notion, Copilot).

Next: how prompts and context actually steer the model — and how to critique and improve a bad one.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.