Lesson 1 of 6 · 48 min

AI product-sense framework

AI product sense is classical product sense executed under probability — every decision is a distribution, not a binary. The six lenses that capture the delta, the five-step case-round playbook, why “underserved” now means “tolerates probabilistic quality,” and the interview that probes all of it.

Why this is not a new discipline

A senior mental model up front: AI product sense is not a new discipline — it is classical product sense executed under probability. The same five case-round steps apply (clarify, motivate, segment, find the problem, design a v1), but every answer must absorb model failure modes: hallucination, distribution shift, miscalibration, and cost variance. Aakash Gupta puts the contrast sharply — regular product sense optimizes exact flows and controlled outputs; AI product sense optimizes the fit between need, model behavior, and economics. This lesson is the ground floor for the whole track, and the part of the case round where most candidates default to a non-AI answer.
Marty Cagan (SVPG) frames the underlying skill as deep product knowledge — a compass for navigating product risks, not a substitute for testing them. For AI products that compass must point through uncertainty. Cagan is explicit that the role does not shrink: he says virtually all product managers will need to be AI product managers, and that the role becomes more essential, not less, with generative AI. The candidate who treats “AI PM” as “PM who picks a model” has already missed the point — the job is choosing where to point a probabilistic capability at a real user problem, and pricing the cost of being wrong.
Becoming an AI PM | Aman Khan (Arize AI, ex-Spotify, Apple, Cruise)Lenny's Podcast

The six lenses — three of them have no classical equivalent

Aakash Gupta’s six-lens model is the most useful map of what changes. Run an AI case-round prompt through all six: User Reality (what does the user actually need?), Model Behavior (what can this model reliably do, not just demo?), Economics (marginal cost per call is now a first-class constraint), System Design (where does the LLM sit inside a deterministic pipeline?), Trust & Liability (what is the harm if it is wrong?), and Go-to-Market (which segment tolerates probabilistic quality?). The senior framing: three of those lenses — Economics, System Design, and Trust & Liability — have no clean equivalent in classic PM frameworks. They are the AI delta.
code
1THE SIX LENSES -- and which are the AI delta23  Lens             Classical equiv?   The question it forces4  ---------------  -----------------  -------------------------------------------5  User Reality     yes                what job, what frequency, what severity6  Model Behavior   partial            reliable capability vs demo-magic7  Economics        NO (delta)         marginal cost/call; gross margin per use8  System Design    NO (delta)         LLM inside a deterministic pipeline?9  Trust & Liability NO (delta)        cost of a wrong answer (annoying vs lawsuit)10  Go-to-Market     partial            which segment tolerates probabilistic quality1112  Senior move: name the three delta lenses out loud. Most candidates only run13  User Reality + GTM and produce a non-AI answer wearing an AI costume.
The Trust & Liability lens earns its place with a single contrast Aakash uses: an annoying AI reply and a lawsuit-prone medical-advice reply have wildly different cost tails. AI products can fail catastrophically — so severity in your problem-identification step must price the worst case, not the average case. Interview angle. When the prompt is “design an AI feature for X,” naming the cost-of-error tail early (“a wrong answer here is embarrassing / costly / dangerous”) is one of the fastest ways to signal senior judgment, because it forces every later decision — eval rigor, human-in-the-loop, refusal behavior — to follow from it.

The five-step playbook, upgraded for probability

Ben Erez’s definitive product-sense structure still holds in AI cases — but each step takes an AI-specific upgrade. The two highest-leverage moves sit in segmentation and problem identification. In AI, “underserved” usually means tolerates probabilistic quality: a beta-friendly prosumer, a B2B user with indemnification, a power user inside a feedback loop. And problem severity must explicitly price the tail, not just frequency × annoyance.
code
1THE FIVE-STEP CASE PLAYBOOK, AI-UPGRADED   (~45-55 min round)23  Step                 Box     Classic goal            AI upgrade4  -------------------  ------  ----------------------  --------------------------5  Clear communication  3-5m    state assumptions       flag every assumption the6                                                       model makes stochastic7  Product motivation   3-5m    why it matters          why the AI variant beats a8                                                       non-AI baseline9  Segmentation         8-10m   reach vs underserved    "underserved" = tolerates10                                                       probabilistic quality11  Problem ID           8-10m   frequency x severity    severity prices the TAIL12                                                       (cost of a wrong answer)13  Solution / v1        8-10m   impact vs effort        pick model class + eval14                                                       harness + graceful failure

Economics & System Design — the two lenses PMs underweight

Of the three delta lenses, candidates handle Trust & Liability passably and ignore the other two. Economics is the one classic product sense never had to price: every call has a marginal cost, so a feature that is free to compute in a SaaS app is now a per-use COGS line. The senior instinct is to reason in gross margin per use — a consumer feature at $0.04/call across 10M daily uses is $400k/day of inference, which can flip a “great idea” into an unaffordable one. This is why Intercom routes between GPT-4 and GPT-3.5 per scenario and why Notion serves a smaller fine-tuned model on the hot path: the economics lens forces the question “can this be margin-positive at scale?” before the model is chosen.
System Design is the lens that decides where the probabilistic component sits. The senior pattern is almost never “the LLM does everything” — it’s the LLM as one stochastic step inside an otherwise deterministic pipeline: deterministic retrieval and tools fetch the facts, the model reasons over them, and deterministic validators check the output. Perplexity’s citations, a RAG system’s ACL filter, and an agent’s tool calls are all this pattern — push the verifiable work to deterministic code and reserve the model for the genuinely fuzzy step. Interview angle. “Where does the model sit in your design?” is a system-design probe hiding in a product-sense round; “the model reasons, the tools and validators do the deterministic work” is the answer that signals you’ve built this before.
code
1THE THREE DELTA LENSES, AS QUESTIONS YOU MUST ANSWER OUT LOUD23  Economics      "Is this margin-positive at scale?"4                 -> reason in gross-margin-per-use; route models by cost-of-error5                 -> $0.04/call x 10M/day = $400k/day -- the idea may not survive67  System Design  "Where does the stochastic step sit?"8                 -> LLM = one fuzzy step inside a deterministic pipeline9                 -> tools fetch facts; validators check output; model reasons1011  Trust & Liability  "What's the cost of a wrong answer?"12                 -> price the TAIL (annoying vs costly vs dangerous)13                 -> the tail sets eval rigor, HITL, and refusal behavior

Innovation is how you apply the capability, not the capability itself

Tal Raviv and Aman Khan frame AI product sense as the gap between frontier model capability and the specific user problem you choose to point it at — and they are blunt that benchmarks will not give you that intuition. Their one-line definition is worth memorizing: AI product sense is “the ability to correctly anticipate what will be truly impactful for users and also feasible with AI.” It develops from genuine opinions about model tradeoffs, earned by using the models on tasks you actually care about — not from reading a leaderboard.
This is why the strongest answers pick a non-obvious wedge and name a real model capability and its limitation in the same breath. Aravind Srinivas (Perplexity) is the canonical example: rather than “an AI that answers questions,” the product call was that every answer must come from a source with domain authority, turning an answer from something that “comes across like an opinion” into a source of truth. The capability (generation) is commodity; the product sense (citations as the trust boundary) is the wedge. Interview angle. “Design an AI product for X” is really testing whether you can name the wedge — the specific, defensible way you apply a commodity capability — not whether you can list features.

Case studies: the same lens set, five different bets

Five named decisions show the lenses in action. Notion Q&A (Ivan Zhao, Nov 2023) launched explicitly as an “isn’t perfect” beta, anchored by a concrete enterprise win — Remote collapsed 10-minute lookups to seconds — a Trust & Liability call made through language (the beta label sets the prior before the first error). Perplexity made citations the wedge (Trust & Liability + GTM), aiming to be “almost never wrong, so you can trust what it says.” Intercom dynamically routes between GPT-4 and GPT-3.5 per scenario (Economics + System Design) — the textbook AI prioritization move. Klarna mis-judged the lens set: its 2024 claim that a bot did the work of 700 full-time agents reversed in May 2025 when it resumed human hiring after quality fell — a catastrophic tail mis-classified as a routine retrievable failure. Microsoft Recall shipped default-on screen capture and was pulled to opt-in after backlash — a trust surface miscalibrated before launch.
AI product sense is not the same as product sense — regular product sense is about exact flows and controlled outputs; AI product sense is probabilistic, optimizing the fit between need, model behavior, and economics. — Aakash Gupta
Notice what the winners share: each one picked a trust posture explicitly and made the cost-of-error survivable before the first hallucination landed. Notion paid in expectation-setting (beta), Perplexity in coverage (citations narrow what it answers), Intercom in routing (cheap model where the error is cheap). The losers — Klarna, Recall — skipped the Trust & Liability lens entirely and let the tail find them. That is the through-line of the whole track: in AI, the recovery is the product.

Interview prep

AI product-sense rounds reward a specific user, a specific model capability-plus-limitation, and a specific cost-of-error — in that order. Be able to answer each of these in 60-90 seconds, leading with the lens, then the product implication.
  1. 01“How is AI product sense different from product sense?” → same five steps, but every decision is a distribution; three new lenses — economics, system design, trust/liability — have no classical equivalent.
  2. 02“Design an AI product for X.” → pick a non-obvious wedge + a segment that tolerates probabilistic quality; name one capability AND one limitation; tie to one input + one output metric.
  3. 03“Why not just use a bigger / better model?” → product sense is the gap between capability and the specific feasible user problem; the wedge is the defensible application, not the model.
  4. 04“What’s the worst that happens if the model is wrong here?” → price the tail (annoying vs costly vs dangerous); the cost-of-error sets eval rigor, HITL, and refusal behavior.
  5. 05“Which segment do you ship v1 to?” → the one whose tolerance for imperfection you can verify — beta prosumer, indemnified B2B, or a power user in a feedback loop.
  6. 06“Walk me through how you’d evaluate feasibility.” → run the six lenses; reliable capability (not demo-magic) on Model Behavior, marginal cost on Economics, harm on Trust/Liability.
  7. 07“What would you NOT ship in v1?” → name one trade-off you reject and why — the open-ended capability you’re deliberately not exposing yet.
  8. 08“Give an example of strong AI product sense in the wild.” → Perplexity citations as the trust boundary, or Notion’s beta framing — capability is commodity, the wedge is the product call.
Follow-ups dig into the deltas. After your design, expect “how would you know it’s working?” (forces a metric split — covered in Lesson 3), “what if it hallucinates?” (forces the trust pass — Lesson 2), and “cut v1 in half” (forces prioritization — Lesson 3). Exponent’s 2026 guidance is the meta-warning to internalize: AI tools have made a structured answer trivial to generate, so structure is now the floor, not the ceiling — your differentiation has to come from a non-obvious user, a real tradeoff, and named cost-of-error, not from reciting CIRCLES.
articleAI Product Sense Is Not the Same as Product Sense — the six-lens roadmapAakash GuptaarticleHow to build AI product senseTal Raviv & Aman Khan (Lenny's Newsletter)articleThe definitive guide to mastering product sense interviewsBen Erez (Lenny's Newsletter)articleAI Product Management 2 Years In — the role becomes more essential, not lessMarty Cagan (SVPG)

Checkpoint

In a case round, you’re asked to design an AI assistant that drafts replies to patient messages for a clinic. Your first structural move?

AName the cost-of-error tail (a wrong medical reply is dangerous, not just annoying) and let eval rigor, human-in-the-loop, and refusal behavior follow from itBPick the highest-accuracy frontier model so quality is maximizedCList the features (summarize history, draft reply, suggest tone) and prioritize them
Sign up free to answer and see why

Checkpoint

An interviewer says: “you proposed an AI study-helper for students — why wouldn’t a competitor with a better model just win?” Strongest reply?

AWe’d fine-tune on more data so our model stays aheadBWe’d ship faster and add more featuresCThe capability is commodity; the wedge is how we apply it — e.g. grounding every explanation in the student’s own curriculum with citations, so trust and fit are the moat, not raw model quality
Sign up free to answer and see why

Checkpoint

You’re segmenting users for a v1 AI contract-review feature. Which segment choice best reflects AI product sense?

AThe largest segment by reach, to maximize impactBA segment whose tolerance for probabilistic quality you can verify — e.g. an in-house legal team using it as a first-pass triage with a human reviewer, not solo external adviceCWhichever segment the sales team is closest to closing
Sign up free to answer and see why

Checkpoint

Klarna’s AI-only support push (claimed to do the work of 700 agents) reversed in 2025 with a return to human hiring. Which lens failure best explains it?

AEconomics — the model was too expensive per callBGo-to-Market — they targeted the wrong customer segmentCTrust & Liability — a catastrophic quality tail was mis-classified as a routine, retrievable failure, with no designed recovery for when the bot was wrong
Sign up free to answer and see why

Checkpoint

A strong candidate runs CIRCLES flawlessly but never mentions hallucination, eval, or cost-of-error. Per Exponent’s 2026 guidance, how does this read?

AAs a top answer — a clean framework is exactly what interviewers wantBAs a 2024 PM — structured is now the floor; the missing AI-specific layer (failure mode, trust pass, cost-of-error) is where the senior signal livesCAs a fail — using CIRCLES at all is a red flag in AI rounds
Sign up free to answer and see why

Could you run the six lenses on a cold AI prompt, pick a wedge + segment, and field the product-sense questions above?

New to itGetting thereConfident

Takeaways

  • AI product sense = classical product sense under probability; same five steps, every decision a distribution.
  • Six lenses; three are the AI delta — Economics, System Design, Trust & Liability — with no classical equivalent.
  • “Underserved” now means “tolerates probabilistic quality”; segment for a verifiable tolerance.
  • Name the failure mode and price the cost-of-error tail before you design the UX — the recovery is the product.
  • The wedge is how you apply a commodity capability (Perplexity citations), not the model itself.
  • Structure is the floor (Exponent 2026); the senior signal is the AI-specific layer on top of CIRCLES.

Next: designing for trust & failure — trust calibration and graceful failure when the model is wrong.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.