Grounding the demo in the customer’s own documents
The “see, it works on YOUR data” moment is what wins the technical evaluation — and it’s the moment most likely to hallucinate. Build grounded retrieval over the customer’s docs with the cookbook Search-Ask pattern, a visible citation panel, an enforced refusal, and a 20-question grounding gate — then know when to graduate to agentic retrieval.
The “it works on OUR data” moment
There is a single beat that closes AI evaluations: the SE drops the customer’s own documents into the demo and it answers correctly, with citations, in front of the people who wrote those docs. That beat is also the one most likely to blow up — because the moment you point a fluent model at a real corpus, it will confidently summarize a paragraph that doesn’t say what it claims. Retrieval-Augmented Generation (RAG) is the pattern that wins here: instead of trusting the model’s training-time memory, you retrieve the relevant passages at query time and make the model answer from them, with a citation and a refusal path. This lesson builds that grounded demo and the gate that proves it’s safe to show. Interview angle. “Build a demo over our docs” is the most common SE take-home; grounding is the whole game.
The conceptual bow to put on it when a buyer asks “how do I know it’s not just hallucinating?”: the model is not generating free text — it is conditioning on chunks you chose. RAG pairs the model’s parametric memory with a non-parametric external memory (your index) accessed at inference time, which buys three things the buyer cares about and fine-tuning can’t: freshness (re-index when a doc changes, don’t re-train), attribution (cite the exact chunk), and access control (filter by who may see what — the spine of the secure-data lesson). For a customer-facing demo, attribution is the buying signal: a cited answer is auditable, and auditability is what a skeptical domain expert is actually testing.
Don’t reach for a framework first. The smallest defensible grounded-QA pattern in production is the cookbook Search-Ask method, and it’s the right default for a demo. Search: precompute an embedding for every chunk of the customer’s docs, embed the query, and rank chunks by cosine similarity; then fill the prompt with the top-ranked chunks only while you stay inside a token budget (e.g. 1,500 tokens), so you never blow the model’s limit. Ask: build a system message that includes those retrieved chunks and the question, and explicitly instructs the model to write “I could not find an answer” when the evidence isn’t there. That refusal instruction is not a nicety — it is the line that converts the model’s natural tendency to invent into a trusted behavior.
python
1import numpy as np2from openai import OpenAI3client = OpenAI()45def embed(texts):6 r = client.embeddings.create(model="text-embedding-3-small", input=texts)7 return np.array([d.embedding for d in r.data])89# SEARCH: rank the customer's chunks by relevance, fill UP TO a token budget.10def search(query, chunks, chunk_vecs, budget=1500):11 qv = embed([query])[0]12 sims = chunk_vecs @ qv / (np.linalg.norm(chunk_vecs, axis=1) * np.linalg.norm(qv))13 ranked = [chunks[i] for i in sims.argsort()[::-1]]14 picked, used = [], 015 for c in ranked:16 t = n_tokens(c["text"])17 if used + t > budget:18 break # token-budget discipline19 picked.append(c); used += t20 return picked2122# ASK: ground the answer in the chunks; REFUSE when the evidence is missing.23SYS = ('Answer the question using ONLY the sources below. Cite the source id [n] '24 'after each claim. If the answer is not in the sources, reply exactly: '25 '"I could not find an answer in the documents."')2627def ask(query, chunks, chunk_vecs):28 picked = search(query, chunks, chunk_vecs)29 sources = "\n".join(f'[{c["id"]}] {c["text"]}' for c in picked)30 msgs = [{"role": "system", "content": SYS},31 {"role": "user", "content": f"Sources:\n{sources}\n\nQuestion: {query}"}]32 out = client.chat.completions.create(model="gpt-4o", messages=msgs)33 return out.choices[0].message.content, [c["id"] for c in picked]
Why does this beat a vanilla prompt-only bot so decisively in a demo? Because Search-Ask conditions generation on real chunks and provides an explicit refusal template, so the model’s tendency to fabricate is overridden by the instruction to abstain. The research case study makes the delta concrete: the prompt-only bot over a 600-page manual hallucinated 3/10 and lost; the Search-Ask rebuild with a 1,500-token budget and the “I could not find an answer” fallback answered 8/10 and explicitly refused 1 — and the customer closed. The refusal wasn’t a weakness the buyer forgave; it was the feature that made them trust the other answers.
Chunking & parsing: where demo grounding actually breaks
The unglamorous truth is that most “the RAG demo is wrong” incidents are ingestion bugs, not model bugs. Two failure points dominate. Parsing: customer docs are PDFs with tables, multi-column layouts, and scanned pages; a naive text extract turns a pricing table into word soup, and the model then “hallucinates” a number that was never readable in the first place. Use a layout-aware parser so tables and headings survive. Chunking: split too coarse and the answer-bearing sentence is buried with noise; split too fine and you sever the context that makes a chunk meaningful (a number divorced from the row label it belongs to). Recursive, structure-aware splitting around ~256–512 tokens with a little overlap is the pragmatic default for a demo.
The senior diagnostic move when a demo answer is wrong: check retrieval before you blame the model. Print the chunks that were actually retrieved for the failing question. Nine times out of ten the answer-bearing chunk never made the top-k — a retrieval miss wearing a hallucination costume — and the fix is upstream (better parsing, different chunk size, a hybrid lexical+dense search for exact part numbers and codes) rather than a prompt tweak or a bigger model. This is also a strong thing to narrate in a technical deep-dive: it shows you understand that in RAG, retrieval — not the model — usually caps quality.
code
1RAG demo "wrong answer" -> the real root cause (debug retrieval FIRST)23 symptom in demo likely cause fix4 ------------------------------ --------------------- --------------------------5 invents a number from a table PDF table -> word soup layout-aware parser6 misses an exact part number dense-only retrieval add BM25 (hybrid + RRF)7 cites the wrong paragraph chunk too coarse/noisy smaller, structure-aware8 number split from its label chunk too fine overlap / parent-doc9 confidently answers off-corpus no refusal instruction enforce "I don't know"1011 Print the retrieved chunks before blaming the model. Most "hallucinations"12 are retrieval misses you can fix upstream.
The citation panel is the trust signal — ship it
The difference between a demo that reads as a chatbot and one that reads as an audit tool is a visible citation panel. Pass each chunk to the model with a stable id, require a citation per claim, render the cited source passages underneath every answer, and — crucially — verify the cited ids actually exist before you display them. Models occasionally cite an id that wasn’t in the context or that doesn’t support the claim; catching that in code (and flagging or retrying) is what keeps the panel honest. For a buyer evaluating whether they can put this in front of their customers or auditors, the citation panel is the product — it’s the artifact that turns “trust me” into “check it yourself.”
python
1# Verify citations are real before rendering them. A panel that cites a2# nonexistent source is worse than no panel -- it breaks trust on inspection.3def grounded_answer(query, chunks, chunk_vecs):4 text, picked_ids = ask(query, chunks, chunk_vecs)5 valid = set(picked_ids)6 cited = extract_citations(text) # e.g. {3, 7}7 if not cited:8 return "I could not find a supported answer in the documents.", []9 if not cited <= valid: # model cited something not retrieved10 log_metric("bad_citation");11 return "I could not verify that against the documents.", []12 sources = [chunk_for(i) for i in cited]13 return text, sources # answer + the passages to render
Four stacked layers of grounding defense
Grounding isn’t one trick; it’s four layers you stack, each catching what the one below missed. Layer 1 — prompt hardening: system instructions plus the refusal template (“say ‘I don’t know’ when the evidence is missing”), which makes abstention a first-class behavior. Layer 2 — schema enforcement: strict-mode tool calls and structured outputs, so the model can’t invent argument names or hand your UI an unparseable blob. Layer 3 — retrieval grounding with prompt-level noise filters, the most useful being Chain-of-Note: the model writes a short relevance note per retrieved chunk (relevant / partial / noise) before composing the answer, so an irrelevant paragraph gets tagged as noise instead of embellished. Layer 4 — evaluation: the offline answer-doc smoke test that gates the demo. The senior framing: these are defenses you layer, not alternatives you choose — the citation panel a buyer sees is the visible output of all four.
code
1Four grounding-defense layers (stack them; the citation panel is the output)23 layer defense what it catches4 ----- ----------------------- ------------------------------------------5 1 prompt + refusal tmpl confident answers when evidence is missing6 2 strict schema / outputs invented arg names, unparseable responses7 3 Chain-of-Note filtering noisy/irrelevant chunks the model embellishes8 4 offline answer-doc eval everything else, before the customer sees it910 Case: a compliance bot fabricated an exemption from one irrelevant paragraph.11 Fix = L3 (tag each para relevant/partial/noise) + an output guardrail that12 refuses below a relevance threshold -> "I could not find a definitive answer."
The research case study for Layer 3 is worth carrying into a deep-dive. A Q&A bot over a bank’s compliance docs told a live demo that a customer “is exempt from FATF rule 7 because all transactions under $1,000 are exempt” — fabricated. The post-mortem: the retriever returned an irrelevant paragraph the model then embellished. The fix wasn’t a bigger model; it was a Chain-of-Note pass that tags each paragraph relevant/partial/noise before answering, plus an output guardrail that refuses below a relevance threshold. After the fix the bot reads the same irrelevant paragraph, tags it noise, and replies “I could not find a definitive answer about that exemption in the documents provided.” The refusal is the win — it’s what a compliance buyer needs to see before they’ll trust the answers that aren’t refusals.
When to graduate: agentic retrieval
Search-Ask is a single retrieve-then-answer pass — perfect for most demos and most questions. Graduate to agentic RAG only when the customer’s real questions demand it: multi-hop (“compare the SLA in our enterprise contract to the one in the addendum”), conditional (retrieve only when the question needs a doc, answer directly otherwise), or tool-composed retrieval (search the docs, then call a pricing API). Here you expose the retriever as a tool an agent can call, so the model decides when and how often to search. The cost is real — more latency, more failure surface, harder to debug — so the senior instinct is default to Search-Ask, escalate to agentic retrieval only when a documented question type forces it. The case study mirrors this: the team built the prototype as a single-pass tool-calling agent for speed, and only added multi-step retrieval once the workflow genuinely needed it.
The named-company shape worth quoting: enterprise-search products like Glean wrap the same primitives (hybrid index, permission filter, rerank, grounded answer with citations) but live or die on multi-tenancy and per-user permissions rather than cleverness — every result must be filtered to what that user may see, at retrieval time. For a sales/support demo you rarely need that machinery on day one, but knowing it’s the production endpoint is what lets you answer the buyer’s “how does this scale to our whole company?” without overbuilding the demo. Interview angle. “When would you use an agent for retrieval instead of plain RAG?” → multi-hop / conditional / tool-composed questions, named explicitly — and note the latency and debuggability cost you’re accepting.
Prove it’s grounded: the grounding gate
Grounding needs the same gate as Lesson 1, sharpened for retrieval. Run your 20 known-good questions and grade each answer as faithful (supported by a real cited chunk), partial, or hallucinated — and add the two refusal cases on purpose: an out-of-corpus question (does it correctly say “I don’t know”?) and a near-miss where the right chunk is hard to retrieve. The disciplined version of this is a verification loop in the spirit of ChainPoll: produce a draft answer, generate verification questions, answer them against the source chunks, and reconcile — but even the manual faithful/partial/hallucinated rubric catches the deal-killers. The payoff beyond safety: this graded set is the artifact you hand the buyer to re-run on their own data, which converts a “does it really work?” debate into a science exercise you control.
Interview prep
The SE take-home is frequently “build a clickable POC + an architecture diagram + a deck + a 45-minute mock presentation,” and a grounded-RAG-over-docs demo is the most common shape. Evaluators scan for preparation, understanding of the problem, clean working code, and a clear narrative. Lead each answer below with the buyer outcome, then the mechanism and the number.
01“How do you stop hallucination in a docs demo?” → ground in retrieved chunks, cite per claim, enforce ‘I don’t know,’ and verify citations exist — not ‘a better model.’
02“What’s the minimum viable RAG for a demo?” → cookbook Search-Ask: cosine-ranked chunks to a token budget + an explicit refusal template; graduate to agentic only when needed.
03“The demo answered wrong — how do you debug it?” → print the retrieved chunks first; most ‘hallucinations’ are retrieval misses (bad parse, chunk size, dense-only) fixed upstream.
04“Why cite sources?” → attribution makes the answer auditable; the citation panel is the buying signal for a skeptical expert and the artifact that turns ‘trust me’ into ‘check it.’
05“RAG vs fine-tuning on their docs?” → RAG for freshness, attribution, and access control; fine-tuning bakes facts in (stale on change, no citation, no per-user filter).
06“When do you use an agent for retrieval?” → multi-hop, conditional, or tool-composed questions — and accept the added latency and debuggability cost.
07“How does this scale to their whole company?” → enterprise-search shape: hybrid index + per-user permission filter at retrieval time + rerank + grounded cited answer.
08“How do you know the grounding is good?” → 20-question faithful/partial/hallucinated gate including deliberate out-of-corpus refusals; hand the buyer the set to re-run.
Going deeper, the technical deep-dive will push on the unglamorous edges: “walk me through ingestion for our messy PDFs” (layout-aware parse → structure-aware chunk → contextualize → embed + index), “our queries have exact part numbers — will dense search find them?” (no — add BM25 in a hybrid + RRF), and “how fresh is an answer after we update a doc?” (re-index that doc; that’s the RAG-over-fine-tuning advantage). The strongest signal is connecting each choice to the failure it prevents and a rough number (token budget, chunk size, top-k). Treat the diagram as ingest → retrieve(filtered) → ground+cite → gate, and narrate the three concerns separately.
Your docs demo confidently states an enterprise SLA of 99.99% — but that figure appears nowhere in the customer’s documents. What’s the most likely root cause and the right first fix?
AThe model is outdated — upgrade to the newest modelBThe answer-bearing chunk never made the top-k (or the table parsed to noise) — debug retrieval and add an enforced refusalCRaise the temperature so the model is more careful
A buyer asks during the demo: “How is this different from just fine-tuning a model on our documents?” Best answer?
AFine-tuning is always better because the knowledge is baked into the modelBThey’re equivalent; RAG is just a cheaper way to fine-tuneCRAG retrieves your docs at query time, so it stays fresh on updates, cites the exact source, and can filter by user permissions
You’re building the minimum viable grounded Q&A demo over the customer’s handbook for tomorrow. Which approach is the right default?
AA multi-agent system with a planner, several retriever agents, and a criticBCookbook Search-Ask: cosine-rank chunks into a token budget, then answer from them with an explicit “I could not find an answer” fallbackCStuff the entire handbook into one large-context prompt and ask
The customer’s users frequently search by exact SKU and error codes. Your dense-embedding retrieval keeps missing them in the demo. Best fix?
AIncrease top-k to 50 and hope the right chunk appearsBSwitch entirely to keyword search and drop embeddingsCAdd lexical BM25 alongside dense retrieval and fuse with RRF (hybrid search) so exact codes are matched as strings
Your demo renders a citation panel. During a dry run you notice it sometimes cites “[12]” when only sources [1]–[6] were retrieved. What does a production-quality demo do here?
AHide the citation panel so the discrepancy isn’t visibleBVerify cited ids against the retrieved set before rendering; if a citation is invalid, refuse or retry instead of showing itCTrust the model — it usually cites correctly, so display whatever it returns
Could you build a grounded docs demo — Search-Ask, a verified citation panel, an enforced refusal — debug it by inspecting retrieval, and prove it with a grounding gate?
New to itGetting thereConfident
Takeaways
Grounding is the “works on YOUR data” moment — condition on retrieved chunks, cite, and refuse when unsupported.
Default to cookbook Search-Ask (cosine rank → token budget → explicit “I don’t know”); graduate to agentic retrieval only when needed.
Most demo “hallucinations” are retrieval misses — debug ingestion (parse, chunk, hybrid search) before blaming the model.
The citation panel is the trust signal; verify cited ids exist before rendering them.
Prove grounding with a 20-question faithful/partial/hallucinated gate including deliberate out-of-corpus refusals — and hand the set to the buyer.
Next: agents for sales and support workflows — when the demo needs to take actions and route work, and how to pick the right control surface.