Turn a vague request into a small test with a decision rule.
A founding engineer often receives a proposed solution before the problem is clear. A customer asks for a dashboard, an integration, or an AI assistant. The request is useful evidence, but it does not establish which outcome matters or whether the proposed feature is the cheapest way to reach it. Ask about the last actual instance of the problem.
Separate observation, interpretation, and decision. An observation might be that a support lead spent forty minutes combining three exports yesterday. An interpretation is that fragmented data caused the delay. A decision is to test one combined export. Keeping these separate prevents the first implementation idea from becoming an unquestioned requirement.
Use a small experiment to reduce the uncertainty that blocks the next decision. A prototype can test whether the combined view answers the user's question. A manual service can test whether someone repeatedly needs the outcome. A production integration tests a different claim about reliability and actual use. Do not claim that positive feedback on a mockup proves retention or willingness to pay.
Write the decision rule before seeing results. The threshold can be provisional, but it should say what result would justify more work, a change, or stopping. Include the cost of the experiment and the risk to the user. A short test with three suitable users may be more useful than a large anonymous survey when the uncertainty concerns a specific workflow. It is still a small sample, so describe its limits.
Worked example
A fictional support lead requests a live analytics dashboard. The observed task is preparing a weekly staffing decision using three CSV exports. The proposed experiment is a single combined table showing unresolved volume and age by queue. The team will observe three weekly planning sessions. Continue if the lead can make the staffing decision from the table and uses it in at least two sessions without a manual reconstruction.
This threshold is invented for the exercise. It does not prove market demand. The technical slice imports fixed files and produces one table. Live ingestion, custom charts, and account administration remain outside this experiment because they do not answer the immediate uncertainty.
Keep a claim ledger
A founder can hear several versions of the same request and still misunderstand the job. Record what was observed separately from what the team inferred. This makes disagreement productive: people can challenge an interpretation without denying the user's experience.
Claim
Evidence type
What would change the decision?
Lead assembled three exports yesterday
Observed workflow or artifact
Find that assembly is rare rather than weekly
Assembly is the main delay
Interpretation
Observe most time spent resolving inconsistent definitions
One combined table will help
Proposed solution
User still rebuilds the spreadsheet to calculate missing measures
A live dashboard is necessary
Unproven requirement
A weekly snapshot supports the decision equally well
The ledger is an original planning method for this exercise. The linked employer screens support the relevance of product and engineering judgment; they do not prove this exact discovery method or interview question is used.
Ask for concrete examples without turning the conversation into a feature vote. The last completed workflow can reveal input files, decision deadlines, handoffs, and workarounds. A statement such as it would be nice to have live data leaves those details unresolved. A screenshot or exported sample can ground the discussion if access is appropriate, but a sample alone does not establish frequency or economic value.
Define an observation protocol
The three-session test needs a consistent observation rule. Otherwise the team can redefine success after a friendly response. Use the same staffing task, record whether the table was available before the decision, and distinguish independent use from use that required the engineer to explain every row.
code
1Session record2- Was the staffing decision due this week?3- Was the input complete before the meeting?4- Could the lead identify the queue needing action?5- Was a separate spreadsheet reconstructed?6- What correction or explanation was required?7- Did the lead use the output in the actual decision?
These fields are not a survey score. They document the behavior relevant to the decision rule. If the lead is absent for one session, that session is not necessarily a product failure, but it also is not evidence of successful use. State whether the observation window will extend or the result remains incomplete.
Assistance matters. If an engineer manually repairs definitions during the meeting, the table may still be useful, but the test has not established independent operation. Record assisted and unassisted outcomes separately. Manual service can be the experiment, provided the claim remains demand for the outcome rather than proof of scalable delivery.
Interpret mixed results
Suppose session one uses the table only after twenty minutes of explanation. Session two uses it independently. Session three reconstructs a spreadsheet because queue labels changed. Under a rule requiring independent use in two sessions without reconstruction, the result is not yet a pass. It identifies a promising outcome and a data-definition gap.
The next change should address that gap if it is small enough to test. For example, a mapping table for source queue names may remove the reconstruction step. A complete live ingestion platform would add many assumptions beyond the observed problem. Describe the new experiment and what evidence would justify expanding it.
Do not change the threshold after seeing results merely to preserve momentum. It is reasonable to revise the experiment when a broken instrument or wrong user group invalidates it, but record why and treat the next run as a new test. Otherwise every outcome becomes success and the experiment cannot improve a decision.
Misconceptions and a second case
One misconception is that a loud request establishes broad demand. It establishes a request from that person in that context. Another is that a small behavioral test proves willingness to pay or long-term retention. It can establish that a specific output helped a specific repeated task; pricing and continued use need their own evidence.
Exercise: a prospect asks for a real-time dashboard, but the staffing decision occurs once every Friday. A manual snapshot takes thirty minutes to prepare and arrives Thursday evening. The lead uses it independently in three sessions. What is supported? The snapshot supports that workflow and the live-update requirement is not established for it. The manual preparation cost remains a delivery constraint. Award one point for the supported behavior, one for separating live updates from the outcome, one for retaining the cost limitation, and one for a next test tied to preparation or a genuinely time-sensitive task.
In an interview, say which uncertainty your proposed build reduces. A technically impressive feature can answer the wrong question. A small experiment is valuable when its result can change the next commitment, including a decision not to build the original request.
Exercise and solution
The lead praises the table but still reconstructs a spreadsheet every week. Should the experiment count as success? No. The observed behavior failed the decision rule. Ask what the spreadsheet provides that the table lacks, then decide whether a smaller change can test that gap. Award one point for using behavior rather than praise, one for identifying the missing outcome, and one for limiting the next change.
Interview probe and wrap-up
What would make you stop building a requested feature? A strong answer identifies evidence that the problem is infrequent, the proposed solution does not help, or the cost exceeds the supported value. Follow up with pressure from a large prospect. A weak answer treats every feature request as a specification. The engineer's contribution is a better decision, sometimes expressed as a smaller build.
A lead says live updates are essential, but the only observed decision is weekly. What is established?
AA weekly snapshot proves pricing.BLive streaming is a confirmed requirement for every customer.CA weekly decision exists; the need for live updates remains to be tested.DThe lead has no real problem.
Three sessions require two independent uses. One assisted use, one independent use, one reconstruction. Result?
ANot a pass under the stated rule; investigate the assistance and reconstruction gaps.BProof the entire product should stop.CPass because two sessions opened the table.DPass because praise was positive.
The lead is absent for a scheduled observation. How should it be counted?
AAutomatically as successful retention.BAutomatically as a product-caused failure.CRemove it without recording the reason.DRecord the missing observation and apply a stated extension/incomplete-result rule.
AIt replaces the user's actual task outcome.BAssisted success does not establish independent delivery.CAssistance makes every demand signal worthless.DIt automatically proves automation will fail.
Queue-name mapping caused reconstruction. Best next bounded test?
AAdd unrelated filters to increase activity.BBuild a complete live ingestion platform immediately.CTest a defined mapping correction and observe whether reconstruction stops.DChange success to mean opening the page.
Can you separate observed work, interpretation, and proposed solution, then interpret a mixed experiment without changing its rule after the fact? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.