Lesson 1 of 5 · 48 min
Building a convincing LLM demo, fast
The solutions-engineer’s job is the technical win, and the demo is the audition. Wrap an LLM API into something a buyer trusts in an afternoon — start from the cookbook, lock the output with strict-mode tool calls, ship the smallest credible chat shell, and gate it with a 20-question smoke test before you ever click “share screen.”
The demo is the audition, not the trial
question_answering_using_embeddings is twenty minutes. The same applies across the stack: there is a named default for the retrieval pattern, the tool-calling reliability knob, and the chat shell, and each default has already absorbed the failure modes of the layer below it. Deviating from a default costs you the deviation in debugging time — which, the night before a customer call, is time you do not have.
How To Prepare for A Customer or Interview DemoWe The Sales EngineersWhy an LLM demo breaks differently than a normal demo
Key idea
Lock the output: strict-mode tool calls turn an LLM into an API component
Company/Account object and a person is a Contact) and an enrichment provider that fills the gaps the CRM is missing. The reliability knob that matters is hidden in plain sight: setting strict: true on a tool guarantees the model’s arguments adhere to your JSON schema instead of being best-effort. Without it, the model occasionally invents an argument name or returns a malformed object, and your demo throws a parse error on stage.1from openai import OpenAI2client = OpenAI()34# A demo tool that looks like the buyer's stack: read the CRM account, fall back5# to enrichment for missing fields. strict:true LOCKS the arg schema, so the model6# can never hand you a malformed or invented-key object on stage.7tools = [{8 "type": "function",9 "function": {10 "name": "lookup_account",11 "description": "Fetch a CRM Company/Account by domain; enrich missing facts.",12 "strict": True, # <-- the reliability knob13 "parameters": {14 "type": "object",15 "properties": {"domain": {"type": "string"}},16 "required": ["domain"],17 "additionalProperties": False, # strict mode requires this18 },19 },20}]2122def lookup_account(domain, crm, enrich):23 account = crm.get_company(domain) # HubSpot/Salesforce object24 if not account.get("employee_count"): # CRM is missing a field25 account.update(enrich.company(domain)) # enrichment fills the gap (below)26 return account2728def run(user_msg, crm, enrich):29 r = client.chat.completions.create(30 model="gpt-4o", messages=[{"role": "user", "content": user_msg}], tools=tools)31 call = r.choices[0].message.tool_calls[0]32 args = json.loads(call.function.arguments) # guaranteed-valid JSON, every time33 return lookup_account(args["domain"], crm, enrich) # your REAL integration runs hereCommon mistake
“If the demo answers fluently, it’s working.”
The chat shell: ship the smallest credible UI
st.chat_message + st.chat_input) gives a complete Python chat app in ~20 lines — ideal for an internal rehearsal or a screen-share. Gradio is a few lines for a shareable Hugging Face Space when the buyer wants to click it themselves. The Vercel AI SDK (useChat) is the move when you need a hosted, branded URL that streams tokens and looks like a real product — it manages chat state and streams from the provider so you write almost no glue. The senior framing: choose the shell by how the demo is delivered (rehearsal vs. shared link vs. embedded), not by what you happen to know.1# A credible demo chat shell in ~20 lines (Streamlit). The CITATION line under2# each answer is what makes it read as an audit tool, not a toy.3import streamlit as st45if "msgs" not in st.session_state:6 st.session_state.msgs = []78for m in st.session_state.msgs:9 with st.chat_message(m["role"]):10 st.markdown(m["content"])1112if q := st.chat_input("Ask about the account..."):13 st.session_state.msgs.append({"role": "user", "content": q})14 answer, sources = grounded_answer(q) # your retrieve+ground call (L2)15 with st.chat_message("assistant"):16 st.markdown(answer)17 if sources:18 st.caption("Sources: " + ", ".join(sources)) # the trust signal19 else:20 st.caption("No supporting source found — I can’t answer that from the docs.")21 st.session_state.msgs.append({"role": "assistant", "content": answer})Ship the eval first: the 20-question smoke test
1The 20-question smoke test (run it the night before, not in the meeting)23 # question expected verdict4 -- ------------------------------------ ----------------- -------------5 1 "What's our refund window?" 30 days FAITHFUL6 2 "Is account ACME on the pro plan?" yes FAITHFUL7 ...8 17 "Does the API support webhooks?" not in docs REFUSED (good)9 18 "What's the SLA for enterprise?" 99.9% HALLUCINATED <- FIX10 ...1112 Gate: 0 hallucinations before the call. Partials get a citation or a refusal.13 A demo that scores 18 faithful / 2 refused beats one that "felt great" once.Key idea
The 60-minute build clock
strict: true. 25–40: wrap it with one output guardrail — an LLM-as-judge that asks “is every claim in this draft present in the retrieved chunks?” and refuses if not. 40–50: the chat shell with a citation panel. 50–60: the 20-question smoke test, fixing every hallucination. The shape that matters: the last ten minutes (the smoke test) is the highest-leverage minute of the hour, and the UI is only ten.1The 60-minute flagship build (allocate by LEVERAGE, not by polish)23 0-10 ingest: chunk + embed + index; 3 sanity-query checks4 10-25 Search-Ask: token budget + "I could not find an answer"; strict tool5 25-40 one output guardrail: LLM-judge "every claim in the chunks?" -> refuse6 40-50 chat shell + citation panel (the trust signal)7 50-60 20-question smoke test -> fix every hallucination <- highest leverage89 Bad path: skip the guardrail + smoke test -> the bot fabricates on stage ->10 "demo-grade AI" -> the deal dies. Good path: same build + the last 20 min.The afternoon build, end to end
strict: true and structured outputs for any rendered data. (3) Add the grounding contract — answer from context, cite, and refuse when unsupported (the next lesson makes this real over the customer’s docs). (4) Drop it in the smallest shell that fits the delivery mode, branded to the customer, with a citation line under every answer. (5) Run the 20-question smoke test and fix every hallucination. The whole loop is an afternoon, and it produces a demo that survives the one thing a chatbot can’t: a skeptical expert asking the question you didn’t plant.The demo is the audition; the pilot is the trial. You win the audition not by being the most fluent in the room, but by being the most trustworthy — the one whose bot cites, refuses, and was measured before anyone clicked “share screen.”
Interview prep
- 01“How do you build a demo fast without it being flaky?” → start from the cookbook, constrain output (strict tools + schemas), add a refusal path, smoke-test 20 Qs before presenting.
- 02“How do you stop the bot from confidently making things up on stage?” → ground every answer in retrieved context, cite the source, and design refusal as a first-class output — not “use a bigger model.”
- 03“What chat shell would you use?” → choose by delivery mode: Streamlit for rehearsal/screen-share, Gradio for a shareable Space, Vercel AI SDK for a hosted branded URL.
- 04“How do you know the demo is ready?” → a 20-question offline smoke test graded faithful/partial/hallucinated, with a zero-hallucination gate — then reuse it as a regression gate.
- 05“The customer asks something you didn’t rehearse and the bot whiffs — what now?” → ‘That’s a great, specific question; I’ll get you a precise answer by end of day’ — never guess; commit to precision.
- 06“Why brand the demo for us?” → a problem-summary slide + the buyer’s logo and vocabulary turns ‘generic AI’ into ‘a tool built for us’ and maps every feature to a stated pain.
- 07“Walk me through what happens when you call the model with a tool.” → model emits a strict-schema tool request → your code runs the real integration → model writes the grounded final answer.
- 08“What separates a great SE demo from a bad one?” → outcomes over features, clean execution (tabs ready, multiple dry runs), honest about limits, and reliability the buyer can see.
Common mistake
The red-flag answer: “I’d just use the most advanced model so it doesn’t make mistakes.”
Checkpoint
It’s the night before a flagship demo. You have a working Q&A bot over the customer’s docs. What’s the single highest-leverage thing to do next?
Checkpoint
Your demo bot calls a lookup_account tool, but ~1 in 20 times it returns a malformed arguments object and your UI throws a parse error on screen. Best fix?
Checkpoint
A buyer wants to click the prototype themselves over the next week and share it with two colleagues, with their branding. Which shell fits best?
Checkpoint
Mid-demo, the prospect asks a detailed question about a feature you haven’t verified. The bot would happily generate an answer. What’s the strongest move?
Checkpoint
You’re starting a brand-new grounded-Q&A demo for tomorrow. What’s the best first step?
Could you stand up a credible, grounded LLM demo in an afternoon — constrained output, a refusal path, a branded shell, and a smoke test — and present it?
Takeaways
- The demo is the audition; it’s judged on visible reliability, not fluency — engineer the unhappy path.
- Start from the cookbook; lock actions with strict-mode tool calls and rendered data with structured outputs.
- Ship the smallest credible shell for the delivery mode (Streamlit / Gradio / AI SDK), branded to the customer, with a citation line.
- Build the eval first: a 20-question smoke test graded faithful/partial/hallucinated, zero-hallucination gate.
- Never guess on stage — commit to a precise follow-up; honesty beats a fabricated answer the expert can catch.
Next: grounding the demo in the customer’s own documents — RAG that cites and refuses, so the bot answers from their truth, not the model’s memory.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.