Lesson 1 of 5 · 48 min

Building a convincing LLM demo, fast

The solutions-engineer’s job is the technical win, and the demo is the audition. Wrap an LLM API into something a buyer trusts in an afternoon — start from the cookbook, lock the output with strict-mode tool calls, ship the smallest credible chat shell, and gate it with a 20-question smoke test before you ever click “share screen.”

The demo is the audition, not the trial

A solutions engineer (SE) gets paid for the technical win — the moment a buyer believes your product solves their problem. In an AI sale that belief is manufactured in a live demo, and the demo is judged on one axis above all others: does it look like it can be trusted with real work? The research is blunt about this — a polished demo is no longer a differentiator (every vendor has one), so the win comes from a demo that is grounded, honest about its limits, and reliable under questioning. This lesson is the afternoon build: from an API key to a credible chat demo, with the reliability scaffolding that separates a $100K pilot from a “thanks but no thanks.” Interview angle. Nearly every SE loop ends in a build-and-present round — what you learn here is also the thing you’ll be graded on in the interview.
Start with the instinct that saves you a day per demo: search the cookbook before you write code. The OpenAI Cookbook ships 400+ runnable recipes — prompt engineering, RAG, agents, full demo apps — and the senior move is to treat it as a parts bin, not a tutorial. Building a grounded Q&A bot from scratch is a day of undifferentiated setup; adapting question_answering_using_embeddings is twenty minutes. The same applies across the stack: there is a named default for the retrieval pattern, the tool-calling reliability knob, and the chat shell, and each default has already absorbed the failure modes of the layer below it. Deviating from a default costs you the deviation in debugging time — which, the night before a customer call, is time you do not have.
How To Prepare for A Customer or Interview DemoWe The Sales Engineers

Why an LLM demo breaks differently than a normal demo

A normal software demo fails in obvious, recoverable ways — a button 404s, a page is slow. An LLM demo fails silently and credibly: it produces a fluent, confident, wrong answer that a domain expert in the room can falsify on the spot. That is the single deal-killer of AI demos, because a buyer can verify one false claim and lose trust in everything. The research case study is canonical: a logistics-SaaS SE demoed Q&A over a 600-page operations manual with a vanilla prompt-only bot; it hallucinated 3 of 10 demo questions and lost the deal. The rebuild — grounded retrieval plus an explicit “I could not find an answer” fallback — answered 8 correctly and refused 1, and the customer closed. Same model, same docs; the difference was the demo’s relationship to the truth.
So the SE’s mental model is the inverse of a developer’s. A developer optimizes the happy path; an SE engineers the unhappy path to be visible and graceful, because that is what gets graded — by the buyer and by the interview panel. Concretely, three properties make an LLM demo trustworthy, and the rest of this lesson builds each: the output is structured (it parses, every time, so the UI never shows a stray-token crash); the bot can say “I don’t know” (refusal is a first-class, designed behavior, not a bug); and the whole thing is measured before the meeting (a smoke test, not a vibe check). Miss any one and the demo is a liability the moment someone asks a question you didn’t rehearse.

Lock the output: strict-mode tool calls turn an LLM into an API component

The fastest way to make a demo feel like software (not a chatbot) is to make the model do things — look up an account, draft an email, fetch a price — via tool calling (a.k.a. function calling). The model emits a structured request naming a tool and its arguments; your code runs the real integration and feeds the result back; the model writes the final answer grounded in that result. In a GTM demo the integration on the other side of the tool is the buyer’s actual stack: a CRM (HubSpot/Salesforce, where an account is a Company/Account object and a person is a Contact) and an enrichment provider that fills the gaps the CRM is missing. The reliability knob that matters is hidden in plain sight: setting strict: true on a tool guarantees the model’s arguments adhere to your JSON schema instead of being best-effort. Without it, the model occasionally invents an argument name or returns a malformed object, and your demo throws a parse error on stage.
python
1from openai import OpenAI2client = OpenAI()34# A demo tool that looks like the buyer's stack: read the CRM account, fall back5# to enrichment for missing fields. strict:true LOCKS the arg schema, so the model6# can never hand you a malformed or invented-key object on stage.7tools = [{8    "type": "function",9    "function": {10        "name": "lookup_account",11        "description": "Fetch a CRM Company/Account by domain; enrich missing facts.",12        "strict": True,                       # <-- the reliability knob13        "parameters": {14            "type": "object",15            "properties": {"domain": {"type": "string"}},16            "required": ["domain"],17            "additionalProperties": False,    # strict mode requires this18        },19    },20}]2122def lookup_account(domain, crm, enrich):23    account = crm.get_company(domain)                 # HubSpot/Salesforce object24    if not account.get("employee_count"):            # CRM is missing a field25        account.update(enrich.company(domain))       # enrichment fills the gap (below)26    return account2728def run(user_msg, crm, enrich):29    r = client.chat.completions.create(30        model="gpt-4o", messages=[{"role": "user", "content": user_msg}], tools=tools)31    call = r.choices[0].message.tool_calls[0]32    args = json.loads(call.function.arguments)       # guaranteed-valid JSON, every time33    return lookup_account(args["domain"], crm, enrich)   # your REAL integration runs here
For outputs that are not a tool call but still need to be machine-readable — a structured summary card, a lead score, a JSON object your UI renders — reach for structured outputs (schema-constrained decoding), which enforces the schema end-to-end. The pairing is the whole trick: strict tool calls for actions, structured outputs for rendered data, and free-form text only for the prose the human reads. Once outputs are constrained, the class of demo bug where “the parser crashed on a weird response” disappears — which is exactly the bug that, on stage, makes a buyer wonder what else is flaky. Interview angle. “How do you make an LLM reliable enough to put in front of a customer?” → constrain the output (strict tools + schemas), add a refusal path, and measure it — not “use a bigger model.”

The chat shell: ship the smallest credible UI

The UI is a wrapper, not a feature — so pick the shell that gets you to a credible demo fastest, by delivery mode. Three defaults cover almost every SE situation. Streamlit (st.chat_message + st.chat_input) gives a complete Python chat app in ~20 lines — ideal for an internal rehearsal or a screen-share. Gradio is a few lines for a shareable Hugging Face Space when the buyer wants to click it themselves. The Vercel AI SDK (useChat) is the move when you need a hosted, branded URL that streams tokens and looks like a real product — it manages chat state and streams from the provider so you write almost no glue. The senior framing: choose the shell by how the demo is delivered (rehearsal vs. shared link vs. embedded), not by what you happen to know.
python
1# A credible demo chat shell in ~20 lines (Streamlit). The CITATION line under2# each answer is what makes it read as an audit tool, not a toy.3import streamlit as st45if "msgs" not in st.session_state:6    st.session_state.msgs = []78for m in st.session_state.msgs:9    with st.chat_message(m["role"]):10        st.markdown(m["content"])1112if q := st.chat_input("Ask about the account..."):13    st.session_state.msgs.append({"role": "user", "content": q})14    answer, sources = grounded_answer(q)           # your retrieve+ground call (L2)15    with st.chat_message("assistant"):16        st.markdown(answer)17        if sources:18            st.caption("Sources: " + ", ".join(sources))   # the trust signal19        else:20            st.caption("No supporting source found — I can’t answer that from the docs.")21    st.session_state.msgs.append({"role": "assistant", "content": answer})
A small but high-leverage detail: brand the demo for the customer. The research on winning SE demos repeatedly flags a fictional-but-specific company, the buyer’s logo and vocabulary, and a problem-summary slide as slide one. It costs ten minutes and changes the buyer’s read from “generic AI thing” to “a tool built for us.” The counterpart failure mode is just as documented: a candidate who demoed on “an old CRM sandbox” with a messy screen and unready tabs bombed — not on the product, but on the impression of un-readiness. The shell is cheap; the polish around it (clean tabs, branded data, a one-line opener) is what the room actually scores.

Ship the eval first: the 20-question smoke test

Here is the single highest-leverage habit in the whole track, and the one most SEs skip: before any customer-facing demo, run a 20-question offline smoke test. Hand-curate 20 questions with known-good answers (the real ones a buyer will ask), run each through the full pipeline, and grade every answer as faithful / partial / hallucinated. Fix the hallucinations before the call. This is the cheapest insurance you can buy: it catches the vast majority of grounding bugs while you can still fix them, and it converts “I think the demo works” into “I know it answers 18/20 and refuses 2.” In the canonical 60-minute build, this is described as the most valuable ten minutes of the hour.
code
1The 20-question smoke test (run it the night before, not in the meeting)23  #   question                              expected           verdict4  --  ------------------------------------  -----------------  -------------5  1   "What's our refund window?"           30 days            FAITHFUL6  2   "Is account ACME on the pro plan?"    yes                FAITHFUL7  ...8  17  "Does the API support webhooks?"      not in docs        REFUSED (good)9  18  "What's the SLA for enterprise?"      99.9%              HALLUCINATED  <- FIX10  ...1112  Gate: 0 hallucinations before the call. Partials get a citation or a refusal.13  A demo that scores 18 faithful / 2 refused beats one that "felt great" once.
Why grade before building the pretty UI? Because the UI is a wrapper and the eval is the product’s actual quality — and a beautiful chat shell backed by an unmeasured bot is the most common GTM-prototype failure mode there is. The discipline scales: the same 20-question set becomes your regression gate when you tweak the prompt or swap the model the morning of the demo, and it becomes the artifact you can hand the buyer later (“here’s the eval set, re-run it on your data”) — which, the research notes, is what turns a performance debate into a science exercise and wins the technical-win. Interview angle. “How would you know your demo is ready?” → “It passes a 20-question smoke test on the buyer’s real questions, zero hallucinations” is a senior answer; “it looked good when I tried it” is a red flag.

The 60-minute build clock

The research distills the flagship build into a 60-minute clock — a useful forcing function because it allocates time by leverage, not by what’s fun to polish. Picture a compliance-officer demo for a fictional Acme Bank. Minutes 0–10: ingest the docs (drop into a folder, chunk, embed into a vector index, verify retrieval with three sanity queries). 10–25: implement Search-Ask with a token budget and the “I could not find an answer” fallback, one tool wired with strict: true. 25–40: wrap it with one output guardrail — an LLM-as-judge that asks “is every claim in this draft present in the retrieved chunks?” and refuses if not. 40–50: the chat shell with a citation panel. 50–60: the 20-question smoke test, fixing every hallucination. The shape that matters: the last ten minutes (the smoke test) is the highest-leverage minute of the hour, and the UI is only ten.
code
1The 60-minute flagship build (allocate by LEVERAGE, not by polish)23  0-10   ingest: chunk + embed + index; 3 sanity-query checks4  10-25  Search-Ask: token budget + "I could not find an answer"; strict tool5  25-40  one output guardrail: LLM-judge "every claim in the chunks?" -> refuse6  40-50  chat shell + citation panel (the trust signal)7  50-60  20-question smoke test -> fix every hallucination   <- highest leverage89  Bad path: skip the guardrail + smoke test -> the bot fabricates on stage ->10  "demo-grade AI" -> the deal dies. Good path: same build + the last 20 min.
The research contrasts two SEs running the same prototype to make the stakes vivid. One skips the guardrail and the smoke test; in the call the bot states a fabricated exemption (“Acme is exempt from FATF rule 7 because all transactions under $1,000 are exempt”), the customer ends the meeting early and writes the company off as “demo-grade AI.” The other ships the identical prototype with the citation panel and the output guardrail; on a borderline query it answers “I cannot find a definitive answer about that exemption in the documents provided, but section 4.2 may be relevant,” and the customer leaves impressed. Same model, same docs — the difference between a $100K pilot and a “thanks but no thanks” email is the last twenty minutes of the build.

The afternoon build, end to end

Put it together as a repeatable runbook you can execute the day before any flagship meeting. (1) Pull the closest cookbook recipe — don’t start from a blank file. (2) Wire one or two real-feeling tools with strict: true and structured outputs for any rendered data. (3) Add the grounding contract — answer from context, cite, and refuse when unsupported (the next lesson makes this real over the customer’s docs). (4) Drop it in the smallest shell that fits the delivery mode, branded to the customer, with a citation line under every answer. (5) Run the 20-question smoke test and fix every hallucination. The whole loop is an afternoon, and it produces a demo that survives the one thing a chatbot can’t: a skeptical expert asking the question you didn’t plant.
The demo is the audition; the pilot is the trial. You win the audition not by being the most fluent in the room, but by being the most trustworthy — the one whose bot cites, refuses, and was measured before anyone clicked “share screen.”

Interview prep

The SE interview’s gating round is build-and-present a customer demo — at Datadog a live tech demo “pitched as if to a prospect,” at Salesforce a case study handed out ~2 weeks ahead, followed by a discovery call and a 30–45 min presentation graded on Substance, Structure, Relevance, Delivery. Treat the build skills in this lesson as the thing being scored. Answer each below leading with the buyer outcome, then the mechanism.
  1. 01“How do you build a demo fast without it being flaky?” → start from the cookbook, constrain output (strict tools + schemas), add a refusal path, smoke-test 20 Qs before presenting.
  2. 02“How do you stop the bot from confidently making things up on stage?” → ground every answer in retrieved context, cite the source, and design refusal as a first-class output — not “use a bigger model.”
  3. 03“What chat shell would you use?” → choose by delivery mode: Streamlit for rehearsal/screen-share, Gradio for a shareable Space, Vercel AI SDK for a hosted branded URL.
  4. 04“How do you know the demo is ready?” → a 20-question offline smoke test graded faithful/partial/hallucinated, with a zero-hallucination gate — then reuse it as a regression gate.
  5. 05“The customer asks something you didn’t rehearse and the bot whiffs — what now?” → ‘That’s a great, specific question; I’ll get you a precise answer by end of day’ — never guess; commit to precision.
  6. 06“Why brand the demo for us?” → a problem-summary slide + the buyer’s logo and vocabulary turns ‘generic AI’ into ‘a tool built for us’ and maps every feature to a stated pain.
  7. 07“Walk me through what happens when you call the model with a tool.” → model emits a strict-schema tool request → your code runs the real integration → model writes the grounded final answer.
  8. 08“What separates a great SE demo from a bad one?” → outcomes over features, clean execution (tabs ready, multiple dry runs), honest about limits, and reliability the buyer can see.
Going deeper, the panel will probe execution and honesty more than cleverness. Expect “tell me about a demo that went wrong” (they want recovery, not perfection), “how would you compress this to 10 minutes for an exec?” (cut to one flow + the value, drop the plumbing), and a live derail mid-demo where they ask something off-script to see whether you keep clicking and guessing or stay composed and commit to follow-up. The cross-source finding is consistent: when content is strong, execution polish multiplies the signal; when content is weak, no polish saves you. Rehearse the build to a stopwatch 5+ times and prepare a tight 30-second opener.
docsOpenAI Cookbook — runnable recipes for demos (RAG, agents, structured output)OpenAIdocsQuestion answering using embeddings-based search (the Search-Ask recipe)OpenAI CookbookdocsFunction calling — strict mode for schema-reliable tool callsOpenAIvideoWhy AI evals are the hottest new skill for product buildersHamel Husain & Shreya Shankar

Checkpoint

It’s the night before a flagship demo. You have a working Q&A bot over the customer’s docs. What’s the single highest-leverage thing to do next?

ARun a 20-question smoke test on the real questions the buyer will ask, graded faithful/partial/hallucinated, and fix every hallucinationBUpgrade to the most advanced model so it makes fewer mistakesCPolish the UI styling and animations so it looks impressive
Sign up free to answer and see why

Checkpoint

Your demo bot calls a lookup_account tool, but ~1 in 20 times it returns a malformed arguments object and your UI throws a parse error on screen. Best fix?

AWrap the parse in a try/except and show a generic errorBLower the temperature to 0 so the output is deterministicCSet strict: true on the tool (with additionalProperties:false) so arguments are guaranteed to adhere to the JSON schema
Sign up free to answer and see why

Checkpoint

A buyer wants to click the prototype themselves over the next week and share it with two colleagues, with their branding. Which shell fits best?

AA local Streamlit app you run on your laptop during callsBA hosted, branded URL built with the Vercel AI SDK (useChat) that streams tokens and manages chat stateCA Jupyter notebook you send them to run
Sign up free to answer and see why

Checkpoint

Mid-demo, the prospect asks a detailed question about a feature you haven’t verified. The bot would happily generate an answer. What’s the strongest move?

ALet the bot answer — it’s usually right and confidence sellsBChange the subject quickly so no one notices the gapCSay you want to give them a precise answer and will follow up with the product team by end of day, then continue
Sign up free to answer and see why

Checkpoint

You’re starting a brand-new grounded-Q&A demo for tomorrow. What’s the best first step?

ADesign a custom retrieval and prompt architecture from scratch to show depthBAdapt the closest OpenAI Cookbook recipe (e.g. Search-Ask) as your starting point, then customizeCPick the trendiest agent framework and learn it tonight
Sign up free to answer and see why

Could you stand up a credible, grounded LLM demo in an afternoon — constrained output, a refusal path, a branded shell, and a smoke test — and present it?

New to itGetting thereConfident

Takeaways

  • The demo is the audition; it’s judged on visible reliability, not fluency — engineer the unhappy path.
  • Start from the cookbook; lock actions with strict-mode tool calls and rendered data with structured outputs.
  • Ship the smallest credible shell for the delivery mode (Streamlit / Gradio / AI SDK), branded to the customer, with a citation line.
  • Build the eval first: a 20-question smoke test graded faithful/partial/hallucinated, zero-hallucination gate.
  • Never guess on stage — commit to a precise follow-up; honesty beats a fabricated answer the expert can catch.

Next: grounding the demo in the customer’s own documents — RAG that cites and refuses, so the bot answers from their truth, not the model’s memory.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.