Lesson 2 of 4 · 25 min

Make cache reuse measurable and safe

Evaluate prefix caching without assuming universal speedups.

Mechanism and reasoning

Prefix caching reuses attention state for an exact shared prefix. It can save prompt computation when many requests begin with the same system instructions or document. It does not remove the work required to generate different continuations. A high cache hit rate can therefore coexist with slow output generation.
Ordering affects reuse. Stable content placed before variable content can share a longer prefix than content that changes near the start. This is a semantic design choice too. Do not rearrange instructions merely to improve a metric if doing so changes model behavior. Model version, tokenizer behavior, adapters and other execution details can determine whether cached state is compatible.
A shared cache also creates an isolation question. The application must enforce authorization before serving content, regardless of cache placement. Reuse mechanisms should not expose another tenant's material through debugging output, unsafe key construction or timing assumptions. Where the runtime supports tenant cache separation or salts, assess the actual mechanism and configuration. Do not infer a security boundary from a performance feature's name.
Measure useful saved work rather than hit rate alone. Token-weighted hit rate and request hit rate answer different questions. One large shared prompt can save more computation than hundreds of tiny prefixes. Memory used by idle cached blocks can also compete with active sequences. Under pressure, eviction may remove the expected benefit.
The experiment should include cold starts, warm repeated prefixes, unique prompts and mixed tenants. Record prefill duration and first-token time separately from decode. That lets you explain whether a change saved computation, reduced queueing or merely changed the workload.

Separate compatibility, ownership and measurement

A cache key is part of the correctness boundary. Two prompts with the same visible text need not have the same token sequence or execution context. A tokenizer change, different adapter, or different model revision can make previously computed state unsuitable. Exact compatibility rules belong to the engine version and configuration. The application should record those dimensions rather than infer compatibility from an endpoint name such as “latest.”
code
1Illustrative compatibility record2model_revision: m173tokenizer_revision: t44adapter_revision: none5cache_format: runtime-specific6tenant_partition: tenant-427prefix_token_ids_digest: digest-of-exact-token-sequence
This record is conceptual. It does not claim that these fields are the exact internal key used by vLLM. It shows the questions an engineer must resolve when reviewing reuse. A tenant partition can intentionally prevent cross-tenant reuse even when prefixes match. That may reduce performance but can simplify isolation requirements. Where an engine supplies a cache-salt mechanism, verify its documented scope and ensure the caller cannot choose another tenant's partition. A salt does not authorize access to source documents.
Consider a retrieval application. Authorization must determine which document text may enter the prompt before the inference call. The cache cannot repair an application that accidentally retrieved another tenant's document. Conversely, a correct retrieval check does not make debug endpoints safe if they expose prompts or state identifiers. Treat application authorization, cache partitioning and diagnostic access as distinct controls.
Traffic sliceRequestsInput tokens/requestReused tokens/requestTotal reused
Shared instructions805001008,000
Shared long document208,0006,000120,000
Total100MixedMixed128,000
All one hundred requests have some reuse, so request hit rate is 100 percent. Total input is 80 × 500 + 20 × 8,000 = 200,000 tokens. Token reuse is 64 percent. The twenty document requests account for most saved tokens. If a report gives only the request hit rate, it hides which class receives the benefit and how much work remains.
Even token reuse is not a direct speedup ratio. Cached blocks still need to be located and used, and the rest of the prompt must be processed. Queueing, generation and transport remain. Suppose a request originally spends 400 ms in prefill, 100 ms elsewhere before the first token, and 1,500 ms generating the remainder. Saving 300 ms of prefill reduces total duration from 2,000 to 1,700 ms, a 15 percent reduction. A large prefill improvement can therefore produce a modest end-to-end improvement without anything being broken.
Cache pressure can alter the result over time. A tiny repeated test may fit every prefix in the cache and never exercise eviction. A production mix can introduce many unique long documents and displace useful blocks. Test a realistic working set, not just a repeated single prompt. Report warm-up behavior, occupancy, eviction or recomputation evidence, and the input distribution. If the runtime does not expose one of these measurements, state the gap instead of inventing a precise hit explanation.
A controlled test can replay the same requests with and without reuse, while keeping model revision, output limits and arrival timing fixed. Do not compare a cached test of short outputs with an uncached test of long outputs. Use client timestamps and server stage metrics to identify whether the saved time came from prompt computation or reduced queueing. The latter is a real benefit, but it should not be described as a faster decode kernel.
Finally, consider a release. A new model revision may start with a cold cache while the old version has a warm working set. Comparing raw canary latency can confound model performance with cache state. Report both representative cold and warm conditions, and maintain version separation during rollback. The deployment decision should account for the temporary cold-start penalty rather than hiding it inside an average.

Worked example

In a teaching trace, 100 requests each have 2,000 input tokens. Sixty share a 1,500-token prefix already in cache; forty are unique. Potential reused tokens are 60 × 1,500 = 90,000 of 200,000 input tokens, or 45%. A request-based metric may report 60% with some reuse. Both numbers are correct under different definitions. Neither says generation time fell by 45%, since output generation remains. If the same trace has long outputs, end-to-end savings can be much smaller.

Exercise

Forty requests have 3,000 input tokens. Ten reuse 2,000 tokens each; thirty have no reuse. Calculate token reuse and request reuse. Name the timing measurement needed to check the benefit.

Model solution and rubric

Reused input is 20,000 of 120,000 tokens, about 16.7%. Ten of forty requests have a hit, or 25%. Compare prefill duration and first-token time on the same trace with cache enabled and disabled. Track cache memory and evictions because a nominal hit rate does not show the effect on active capacity. Keep output lengths fixed in the comparison.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

“Similar text is reusable prefix state.” Reuse requires the engine's exact compatibility rules, including token sequence and execution identity.
“A tenant salt replaces authorization.” It can partition reuse, but it does not decide which documents or outputs a caller may access.

Interview probe

Evidence class: recommended. Original practice.
Our cache hit rate is 80%, but cost barely changed. Explain.
Strong answer: I would first ask whether that is request-weighted or token-weighted. Small shared prefixes can create many hits while saving little prefill. Long decode work, cache memory pressure or other bottlenecks can dominate total cost.
Follow-up: How would a variable request ID near the start of every prompt change prefix reuse?
Weak answer indicators: Equating hit rate with end-to-end speedup; using approximate text similarity as proof of exact prefix compatibility; ignoring tenant isolation.

Sources

Technical references: vLLM prefix caching; vLLM metrics design. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsvLLM prefix cachingdocs.vllm.aidocsvLLM metrics designdocs.vllm.ai

Checkpoint

A tokenization revision changes while visible prompt text stays the same. What must be checked?

AReuse the state whenever the UTF-8 prompt bytes matchBWhether the exact cache compatibility rules still holdCReuse the state whenever old and new token counts matchDReuse the state whenever the model weight revision matches
Sign up free to answer and see why

Checkpoint

All 100 requests reuse a 100-token prefix, but each input has 10,000 tokens. Token reuse is?

A100%B10%C1%D0.1%
Sign up free to answer and see why

Checkpoint

Prefill falls from 400 to 100 ms; all other work remains 1,600 ms. End-to-end duration falls by?

A75%B64%C30%D15%
Sign up free to answer and see why

Checkpoint

A tenant cache partition is configured. Which authorization requirement remains?

ACheck whether the caller may use the source documentsBAllow any retrieved document because cache is partitionedCTrust a caller-supplied tenant partition without validationDRemove application checks once hit rate is high
Sign up free to answer and see why

Checkpoint

A warm single-prefix benchmark is fast, but diverse production traffic is slow. What next test addresses the gap?

AIncrease output lengths onlyBReplay a representative prefix working set with eviction pressureCReport request hit rate aloneDCompare different model revisions simultaneously
Sign up free to answer and see why

Explain how you would evaluate prefix caching without assuming universal speedups without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Explain what a hit counts. Verify saved prefill work and preserve the application's authorization rules.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.