Lesson 1 of 6 · 42 min

How LLMs actually behave

An LLM is not a function — same input, different output; finite memory you pay for by the token; latency that scales with what it writes. The generation loop, tokenization, the context window, sampling, reasoning models, and the interview questions that probe all of it.

Why your code can’t treat it like a function

A function returns the same output for the same input, instantly, for free. An LLM does none of that: it’s a probabilistic text generator with finite working memory you pay for by the token and latency that scales with how much it writes. Almost every reliability, cost, and structure decision you’ll make — and almost every foundational interview question you’ll get — is a direct consequence of those three facts. This lesson is the ground floor the rest of the track is built on.
Mechanically, the model is autoregressive: given the tokens so far, it produces a score (a logit) for every token in its vocabulary (~100k–200k entries), turns those into probabilities with a softmax, samples one, appends it, and repeats — until it emits a stop token or hits a length cap. Everything we call “reasoning” or “instruction-following” is emergent behaviour on top of that one loop. The first engineering consequence: your output is a sample from a distribution, not a return value.
Three properties fall out of that loop, and they organise this whole track: it is probabilistic (same input → different output), it has finite working memory you rent by the token (the context window), and its latency scales with how much it writes (one forward pass per output token). Interviewers love to open with “walk me through what actually happens when you call an LLM API” — by the end of this lesson, that walkthrough is yours.
Andrej Karpathy — Intro to Large Language ModelsAndrej Karpathy

The generation loop: prefill, then decode

A single call runs in two phases. Prefill ingests your entire prompt in one parallel forward pass, building the KV cache (the attention keys/values for every input token). Decode then generates output tokens one at a time, each a forward pass that attends to the growing KV cache. Prefill is parallel and compute-bound; decode is sequential and memory-bandwidth-bound. That asymmetry explains nearly every latency behaviour you’ll observe.
code
1ONE LLM CALL = PREFILL  +  DECODE23  prefill   ── process all N input tokens in ONE parallel pass ──►  build KV cache4              (cost ∝ input length; this is most of your time-to-first-token)56  decode    ── token ── token ── token ── ... ── <stop>7              (each = one forward pass reusing the KV cache; ~constant per token)8              total decode time ∝ NUMBER OF OUTPUT TOKENS910  TTFT  ≈ prefill time           (shrink the PROMPT to improve it)11  total ≈ TTFT + out_tokens × inter-token-latency   (shrink the OUTPUT to improve it)</stop>
This is why “why is the first token slow but the rest stream out fast?” is a classic warm-up question: the wait before the first token is the prefill pass over your whole prompt; after that, each token is a cheap incremental step. It’s also why a 4,000-token prompt that returns 50 tokens feels snappy, while a 500-token prompt that returns 2,000 tokens feels slow — output length dominates total latency, input length dominates time-to-first-token.

Tokens: the unit of cost, latency, and limits

Models don’t see words or characters — they see tokens, sub-word pieces from a byte-pair-encoding (BPE) vocabulary. Rule of thumb for English: ~4 characters ≈ 1 token, ~0.75 words per token. Code, JSON, and non-English text tokenize less efficiently (more tokens per character), which quietly inflates cost. Tokens are the unit of everything: pricing, latency, and the context-window limit are all counted in tokens, so counting them is the first thing a production client does.
python
1import tiktoken23enc = tiktoken.encoding_for_model("gpt-4o")4def n_tokens(text: str) -> int:5    return len(enc.encode(text))67prompt = open("contract.txt").read()8print(n_tokens(prompt))          # e.g. 9,142 -> ~$0.02 input on a mid-tier model910# tokenization is why models "can't spell": they never see characters.11enc.encode("strawberry")          # -> a few sub-word tokens, not 10 letters12# ...so "how many r's in strawberry?" is genuinely hard -- it's a tokenization artifact,13# not a reasoning failure. Same root cause as shaky digit-by-digit arithmetic.
Two pricing facts to memorise. First, output tokens cost ~3–5× more than input tokens on most providers — because decode is sequential and compute-heavy, while prefill is parallelised. Second, a long stable prefix can be prompt-cached (covered in L5) for up to ~90% off its input cost. Together these mean the cheapest feature is one with a big cached prompt and a short output — exactly backwards from the intuition that “shorter prompts are cheaper.”

Interview drill: estimate cost & latency on a whiteboard

Interview angle. “Roughly what will this feature cost per month, and what’s its latency?” is a standard senior screen — they’re testing whether you think in tokens. The back-of-envelope: monthly_cost ≈ requests × (in_tokens × price_in + out_tokens × price_out), and latency ≈ prefill(in_tokens) + out_tokens × inter_token_latency. Always state your assumptions out loud (model, token counts, price per million) — that’s what they’re grading.
python
1# Back-of-envelope a support assistant: 200k requests/day, mid-tier model2reqs_per_day = 200_0003in_tokens, out_tokens = 1_200, 2504price_in, price_out = 2.50 / 1e6, 10.00 / 1e6     # note: output ~4x input56daily = reqs_per_day * (in_tokens * price_in + out_tokens * price_out)7print(round(daily), round(daily * 30))            # ~$1,100/day  ~$33k/month89# latency intuition (rough): TTFT tracks the 1,200-token prefill;10# total tracks the 250 output tokens. Cut output -> cut total time AND cost.

The context window is working memory — and it rots

The context window (e.g. 200k tokens) is the model’s working memory for one call — and the system prompt, conversation history, retrieved documents, tool definitions, tool results, and the output all share that one budget. It is not long-term memory; nothing persists across calls unless you put it back in. Treating the window like a database is the root of a whole class of bugs.
Two empirical facts every senior must know. (1) Lost in the Middle (Liu et al., 2023): recall is U-shaped — facts at the start or end of a long context are recalled reliably, facts buried in the middle are frequently missed. (2) Context Rot (Chroma, 2025): answer quality degrades as input grows even in frontier models, and it’s a reasoning degradation, not just a retrieval miss. The practical upshot: more context can make answers worse, not just slower.
There’s also a cost/latency reason not to overfill it: attention cost grows super-linearly with sequence length, and the KV cache grows linearly in memory, so a 100k-token context is far more expensive to serve than its input-token price alone suggests. GitHub Copilot deliberately capped prompts around 6,000 characters specifically to keep fast models inside its latency envelope — a concrete example of treating the window as scarce.

Interview drill: “handle a document bigger than the window”

Interview angle. “The user uploads a 500-page PDF — how do you answer questions over it?” The wrong answer is “use a bigger-context model.” The senior answer walks the menu and picks by constraint: chunk + retrieve (RAG) the few relevant passages (default; cheapest, best recall on targeted questions); map-reduce summarise when the question needs the whole document (“summarise this contract”); hierarchical summaries for very large corpora; and reserve long-context stuffing for when the whole doc genuinely fits and the task is holistic. Name the tradeoff: stuffing is simplest but hits Lost-in-the-Middle, cost, and latency.
python
1# map-reduce summarization: the answer to "summarize a doc bigger than the window"2def summarize_large(doc, chunk):3    parts = [summarize(c) for c in chunk(doc)]   # MAP: summarize each chunk (parallel)4    while n_tokens("\n".join(parts)) > WINDOW_BUDGET:5        parts = [summarize("\n".join(g)) for g in groups_of(parts, 5)]  # REDUCE: fold up6    return summarize("\n".join(parts))           # final pass over the folded summaries

Temperature, top-p, top-k: the sampling knobs

Temperature rescales the logits before the softmax — high temperature flattens the distribution (more varied, more “creative”), low sharpens it toward the single most likely token (more greedy, more repetitive). top-p (nucleus) samples only from the smallest set of tokens whose cumulative probability exceeds p; top-k from the k most likely. In practice you tune temperature and leave top-p near its default; they interact, so changing both at once makes behaviour hard to reason about.
The senior trap: temperature 0 is not a determinism guarantee. It makes decoding greedy (always take the argmax), but real systems still vary run-to-run — from batched inference, floating-point non-associativity on GPUs, mixture-of-experts routing, and silent model updates by the provider. A seed parameter helps reproducibility but is best-effort, not a contract. Your downstream code must never assume byte-identical output across calls.
  1. 01Extraction / classification / structured output → temperature 0–0.2 (you want the most likely answer).
  2. 02Balanced assistant / Q&A → ~0.7 (some variety, still grounded).
  3. 03Brainstorming / creative / multiple drafts → 0.9–1.2 (explore the distribution).
  4. 04Never depend on exact reproducibility — even at temperature 0, design for variation (schemas, evals).

Reasoning models & the knobs that came with them

Modern reasoning models (extended-thinking modes) generate a hidden chain of “thinking” tokens before the visible answer. They lift accuracy on hard, multi-step problems — but you pay for those hidden tokens, they add latency, and streaming them is a common source of surprise bills. Newer APIs expose a reasoning-effort knob (low/medium/high) to trade quality for cost and speed. The senior instinct: reasoning models for genuinely hard planning/math/code, standard models for the latency-sensitive, high-volume, “just extract this” majority.
Interview angle. “When would you use a reasoning model versus a standard one?” Strong answer: reach for reasoning when the task has multiple dependent steps or needs verification (complex code, multi-hop analysis), and stay standard when latency, cost, or volume dominate and the task is shallow — and note that you can route between them per request (L5) rather than choosing globally.

The probabilistic tax (the through-line of this track)

Because output is probabilistic and free-form, four problems follow — and each gets its own lesson: parsers break on a stray token (→ structured outputs, L3); calls time out, rate-limit, and 500 (→ reliable client, L4); tokens cost money and time (→ cost & latency, L5); and “looks right” isn’t “is right” (→ testing & evals, the capstone).
There’s a fifth, sneakier one: silent regressions. Anthropic’s own April 2026 Claude Code postmortem traced two months of “it got dumber” reports to three interacting changes shipped weeks apart — a default reasoning-effort cut (high → medium), a caching bug that dropped older thinking from idle sessions, and a verbosity-reduction system-prompt tweak. Individually small; together they read as broad, inconsistent intelligence loss. The lesson that protects you: pin model and prompt versions, and run an eval on every change.

Case studies: how production teams tame these facts

Notion cut chat latency from ~2s to ~350ms (4×) for 100M+ users by serving a smaller, fine-tuned model on dedicated infra rather than a frontier model on the hot path — their framing, worth quoting in an interview, is that “latency is perceived as quality.” GitHub Copilot caps prompts (~6,000 chars) to keep fast models in budget, and earned measured wins from context, not model size: fill-in-the-middle gave +10% relative acceptance, neighbouring-tab context +5%. Perplexity sustains ~200M queries/day at p50 358ms / p95 <800ms by doing cheap retrieval first and running an expensive reranker only on the top candidates.
Latency is perceived as quality. — the recurring lesson across Notion, Copilot, and Perplexity: the team that shrinks what hits the model before the model runs wins on both speed and cost.

Interview angle: the questions you’ll actually get

Foundational LLM screens cluster around a handful of questions. Be able to answer each in 60–90 seconds, leading with the mechanism, then the production implication.
  1. 01“Walk me through what happens when you call an LLM.” → tokenize → prefill (parallel, builds KV cache) → decode (sequential, one token/pass) → stop. Output is a sample, not a return value.
  2. 02“Why is the first token slow but streaming is fast?” → prefill processes the whole prompt at once (TTFT); decode is cheap per token.
  3. 03“Estimate cost and latency for this feature.” → tokens × price (output ~4× input); TTFT ∝ prompt, total ∝ output. State assumptions.
  4. 04“Is temperature 0 deterministic?” → no — greedy ≠ deterministic; batching/hardware/MoE/updates vary it. Design for variation.
  5. 05“The doc is bigger than the window — what do you do?” → chunk+RAG / map-reduce / hierarchical; not “bigger model.”
  6. 06“Why are LLMs bad at counting letters / arithmetic?” → tokenization (sub-word units); give it a tool.
  7. 07“Reasoning model vs standard?” → hard multi-step vs latency/cost-sensitive; route per request.
  8. 08“How do you handle nondeterminism in tests and prod?” → schemas + validation + eval gates, not exact-match assertions.
videoDeep Dive into LLMs like ChatGPT (tokenization, sampling, the whole stack)Andrej KarpathydocsTiktokenizer — see exactly how text becomes tokensinteractive playgroundpaperLost in the Middle: How Language Models Use Long ContextsLiu et al. (arXiv)articleContext Rot: how increasing input tokens degrades LLM performanceChroma ResearchdocsHugging Face LLM Course (ch. 1–2: how LLMs work)Hugging Face

Checkpoint

You set temperature=0 and parse the model’s reply with json.loads(). It works in dev, then crashes intermittently in production. Best explanation + fix?

AA bug in your parsing code — json.loads is deterministicBtemperature 0 isn’t a determinism guarantee — enforce a schema (structured outputs) and validate, instead of assuming stable textCRaise the temperature so the model is more confident
Sign up free to answer and see why

Checkpoint

You stuff 60 documents into one prompt; the model nails facts from the first and last few but ignores a crucial one in the middle. Best fix?

AMove to a model with a larger context windowBRetrieve only the few relevant chunks and place them at the start/end of the context (or use RAG)CIncrease temperature so the model explores more of the context
Sign up free to answer and see why

Checkpoint

An interviewer asks: “why is the first token slow but subsequent tokens stream out quickly?” Best answer?

AThe model loads its weights for the first token, then caches themBPrefill processes the whole prompt in one parallel pass (TTFT); then each output token is a cheap incremental decode step reusing the KV cacheCThe first token uses a slower sampling method
Sign up free to answer and see why

Checkpoint

A feature sends a 1,000-token prompt and gets a 1,000-token answer; another sends 2,000-token prompts but caps answers at 100 tokens. Which is cheaper/faster, and why?

AThe first — shorter prompts are always cheaperBThe second — output tokens cost ~4× input and dominate total latency, so a short output beats a short promptCThey cost the same — only total tokens matter
Sign up free to answer and see why

Checkpoint

Users report your support bot “got noticeably worse this week,” but you shipped nothing. What’s the most likely cause and the right safeguard?

ARandom bad luck — LLMs vary, nothing to doBThe provider silently updated the model; pin the model version and run an eval gate on every change so you detect (and prevent) itCYour context window shrank
Sign up free to answer and see why

Could you whiteboard the generation loop, estimate a feature’s cost/latency, and field the foundational interview questions above?

New to itGetting thereConfident

Takeaways

  • A call is prefill (parallel, sets TTFT) + decode (sequential, one pass per output token, sets total time).
  • Everything is tokens: output costs ~3–5× input; tokenization is why char/number tasks need tools.
  • Context is scarce and ordered — Lost-in-the-Middle + Context Rot mean more context can mean worse answers.
  • temperature 0 is greedy, not deterministic; design for variation with schemas + evals.
  • Reasoning models trade cost/latency for multi-step accuracy — route per task, don’t default globally.
  • Pin versions and eval every change, or a silent provider update becomes your incident.

Next: prompting & context engineering — designing the input so you get the output you want, reproducibly.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.