Lesson 3 of 6 · 49 min

Defining success & first 90 days

Success is a number in the customer’s UI, not in your code. A three-layer Definition of Done, the four-metric FDE scorecard, and a concrete 30/60/90 calendar — week-1 small win, a reusable asset by day 60, a cross-customer improvement by day 90 — drawn from Exponent and Perspective AI.

From “it works” to a number the customer can check

The single rule every FDE playbook converges on: success is a metric in the customer’s UI, not in your code. Salesforce frames the whole model around it — “FDEs don’t arrive with a checklist of predefined deliverables. They align to a customer’s business outcomes and stay accountable until those outcomes are achieved.” The consequence is concrete: charters, success criteria, and weekly status are all written in business-outcome language, not feature-shipment language. If the customer can’t answer “what number moves from what to what,” you don’t have a spec yet.
Exponent names the six factors that drive an AI system’s success criteria — “data volume, refresh frequency, latency, cost, governance, and what kind of error is acceptable.” That last one is the senior nuance: a false positive and a false negative are not equally bad, and which one the customer can tolerate shapes the entire eval. Chip Huyen adds that “manual data inspection is critical” and offers “the highest ratio of value” — you “measure what matters” by looking at real outputs, not by trusting an intuitive metric, because “a convincing summary might not be a good summary.”
So the Definition of Done has three layers, written at kickoff. (1) A single leading indicator the customer will feel — the one number, with a from/to. (2) A thin eval suite that scores it — automated metrics plus human review on a stratified sample plus a live user-feedback signal (Exponent’s combination: “exact match, BLEU/ROUGE for narrow tasks, custom LLM-as-judge with rubric… human review on a sampled stratified set… continuous user-feedback signal”). (3) A runbook for the customer’s ops team — escalation criteria, monitoring, what to do when it’s wrong. Miss any layer and “done” drifts back to subjective.
code
1DEFINITION OF DONE  (write this at kickoff, in the charter)23  LAYER 1  Leading indicator the customer FEELS4           one metric, with from -> to and where it's measured5           e.g. "median dispatch decision time: 90s -> 45s, on the ops dashboard"67  LAYER 2  Thin eval suite that SCORES it8           automated (exact-match / rubric LLM-judge) + human review on a9           stratified sample + a live user-feedback signal1011  LAYER 3  Runbook for the customer's OPS team12           monitoring, escalation criteria, what to do when it's wrong1314  Test: if the customer can't say "what number moves from what to what",15        you don't have a spec -- you have a vibe.

The four-metric FDE scorecard

Perspective AI operationalizes “success” at the function level into four CEO-visible numbers. These are worth memorizing because they reframe FDE work from “did we ship” to “did we create durable value” — and the reusable-asset targets are sharper than most teams expect.
code
1THE FOUR-METRIC FDE SCORECARD  (Perspective AI)23  Metric                What it measures                       Target4  -------------------   ------------------------------------   -----------5  Deal velocity         pilot kickoff -> production deploy      <= 90 days median6  Net Revenue Retention NRR on FDE-touched accounts             130%+7  Productization rate   features shipped from FDE work / eng    >= 1.0 per engagement8  Reusable-asset ratio  code in product repo vs customer repo   70%+ in main by mo. 12910  Read it as: ship fast (<=90d), expand accounts (130% NRR), and turn every11  engagement into >=1 product feature -- or you're running a consultancy.
The DEV playbook frames the same outcome as a unit-economics compression: a typical enterprise sale takes ~15 months to value; an FDE engagement targets “5 months instead of 15.” And it lists the deliverables of a successful exit: “a production system, an evaluation framework (automated evals, monitoring dashboards, escalation criteria), a runbook of operational procedures, and internal champion enablement.” Notice that three of four exit deliverables are the Definition-of-Done layers — the eval, the runbook, the metric — plus the human you leave behind who can run it. Interview angle. “How do you know your AI system is actually working well?” → name the customer-facing metric, the eval that scores it (goldens + rubric + thresholds), and the live feedback signal — in two minutes.

The first 90 days: a concrete calendar

Two provider-grade 30/60/90 calendars cover the ramp — Exponent’s lens is the individual FDE; Perspective AI’s lens is the function. Stack them. The three non-negotiables they share: a small win in week 1 to earn the right to ask harder questions, a reusable asset from a single-customer insight before day 60, and a cross-customer or cross-team improvement before day 90.
code
1FIRST 90 DAYS  (individual FDE, after Exponent + Perspective AI)23  Days 1-30   LEARN & EARN TRUST4              read all customer call transcripts; ship first eval harness;5              shadow customer calls; SHIP ONE SMALL WIN (week 1).6  Days 31-60  OWN A DEPLOYMENT & PRODUCE ONE REUSABLE7              lead kickoff + discovery + integration; hold the pager;8              build ONE reusable integration; present at a customer QBR;9              write the "what should we productize?" memo.10  Days 61-90  CROSS-CUSTOMER LEVERAGE & HANDOFF11              drive a cross-customer improvement; PRODUCTIZE >=1 feature12              from the first engagement so the next customer gets it13              out of the box; propose tooling that lifts team throughput.1415  Anchor milestones: wk1 no-feature demo on real auth -> wk4 thin end-to-end16  pipeline (one happy-path query) -> wk8 productization memo -> wk12 ship it.
Two details make this credible rather than generic. First, the week-1 small win is often a no-feature demo that integrates with the customer’s real authentication — it proves you can operate inside their environment and earns trust before you’ve built anything substantive. This is Nabeel’s “beat the low expectations early” mechanic made into a milestone. Second, the day-60 reusable asset is the moment FDE work either starts feeding product or quietly becomes consultancy — Perspective AI locks it with a written “what should we productize?” memo, because “without a written commitment it never happens.”
How Palantir Built the Ultimate Founder FactoryLenny’s Podcast · Nabeel Qureshi

Presenting the plan to the customer (and to a hiring panel)

You will present a first-90-days plan twice: to the customer at kickoff (to set expectations) and, almost verbatim, in an FDE interview (“what would your first 90 days look like?”). Both audiences want the same thing — evidence you’ll create visible, measured value on a credible cadence, not a learning-plan. Lead with the week-1 win and the day-90 outcome (the metric moving), then fill the middle. The startup question bank phrases the bar as making deployments “observable so issues are easy to detect and diagnose” — your plan should name when the eval, the dashboard, and the runbook exist, not just that they will.
A subtle senior move: tie each phase to a handoff artifact, not a date. Week 1 produces the trust-building demo; week 4 the thin pipeline + first eval; week 8 the productization memo; week 12 the productized feature in the next customer’s environment. Treating each phase as a measurable handoff (the prototype, the runbook, the reference call) rather than a checklist is what makes the plan land as a commitment. Interview angle. Perspective AI’s function-level calendar even front-loads “first eval harness for the core model; one bug fix in prod; read all customer call transcripts” into days 1–10 — naming that kind of concrete day-10 deliverable signals you’ve actually onboarded as an FDE before.
A first-90-days plan is a promise about visible, measured value on a cadence — week-1 win, day-60 reusable asset, day-90 cross-customer lift — not a syllabus of what you’ll learn. The same plan sets customer expectations and passes the interview.
articleHow to Build a Forward-Deployed Engineering Function (the scorecard + 30/60/90)Perspective AIarticleForward Deployed Engineers Are Proving AI Makes Tech More Human (outcomes, not checklists)SalesforcearticleAI Engineering with Chip Huyen (manual inspection, “measure what matters”)Pragmatic Engineer

Checkpoint

At kickoff, the sponsor says “we’ll know it’s working when the team feels more productive.” What’s the strongest FDE response?

AAccept it — qualitative sentiment is a valid success criterion for an AI toolBPin it to a leading indicator with a from/to the customer can see — e.g. “median time on task: 12 min → 7 min on your dashboard” — plus the eval that scores it and the runbook when it’s wrongCPromise a productivity dashboard later and start building the feature now
Sign up free to answer and see why

Checkpoint

You’re asked in an FDE interview to describe your first 90 days. Which plan best signals seniority?

ASpend the first month learning the codebase and product deeply, then start contributing in month twoBShadow the team for 90 days to fully understand the domain before shipping anythingCWeek-1 small win (a demo on the customer’s real auth), own a deployment and build one reusable integration by day 60, drive a cross-customer improvement and productize a feature by day 90
Sign up free to answer and see why

Checkpoint

An interviewer asks “how do you know your AI system is actually working well?” What’s the strongest two-minute answer?

AName the customer-facing metric (from→to), the eval that scores it — goldens + rubric + pass/fail thresholds — and the live user-feedback signal, plus manual inspection of real outputsBPoint to strong demo results and positive stakeholder reactionsCCite the model’s benchmark scores (e.g. MMLU) as evidence of quality
Sign up free to answer and see why

Checkpoint

By day 55 of an engagement you’ve shipped a working deployment but produced nothing reusable and written no productization memo. What does the scorecard say you’re risking?

ANothing — a working deployment is the whole job; reuse is the product team’s concernBYou’re trending toward consultancy-in-disguise: missing the day-60 reusable asset and the written productization commitment that the scorecard makes CEO-visibleCYou should slow down future deployments so reuse happens naturally
Sign up free to answer and see why

Checkpoint

A sponsor pushes for the most ambitious success metric possible — “zero errors, fully automated, no human in the loop.” The data suggests false negatives are far costlier than false positives here. What’s the senior move?

AAgree to zero errors — aiming high motivates the team and impresses the sponsorBReframe success around acceptable error: define the tolerable false-negative rate and a human-in-the-loop threshold, and write that into the eval and runbookCPick higher accuracy overall and ignore the false-positive/false-negative distinction to keep it simple
Sign up free to answer and see why

Interview prep

The success-and-ramp material shows up two ways in FDE rounds: as “how do you know it’s working?” (a success-criteria question) and as “what would your first 90 days look like?” (a ramp question). Both reward the same instinct — translate work into visible, measured customer value on a credible cadence. Lead with the customer-facing number and the eval that scores it; for the ramp, lead with the week-1 win and the day-90 outcome, then fill the middle with reusable-asset and cross-customer milestones.
  1. 01“How do you define success?” → a single leading indicator the customer feels (from→to, on their screen), an eval that scores it, and a runbook when it’s wrong.
  2. 02“How do you know it’s actually working?” → goldens + rubric + thresholds + live feedback, plus manual inspection of real outputs; a convincing summary isn’t a good summary.
  3. 03“What’s your first 90 days?” → wk1 small win → own a deployment + one reusable integration by day 60 → cross-customer improvement + productize ≥1 feature by day 90.
  4. 04“What’s the week-1 win?” → a no-feature demo on the customer’s real auth — proves you can operate in their environment and beats low expectations early.
  5. 05“What metrics matter for the FDE function?” → deal velocity (≤90d), NRR (130%+), productization rate (≥1/engagement), reusable-asset ratio (70%+ in main by month 12).
  6. 06“What do you leave behind?” → a production system, an eval framework, a runbook, and an enabled internal champion.
  7. 07“How do you handle ‘zero errors’ asks?” → reframe around acceptable error type (which of FP/FN the customer can tolerate) and a human-in-the-loop threshold.
  8. 08“How do you make a deployment observable?” → name when the eval, monitoring dashboard, and escalation criteria exist — a date-bound handoff, not a promise.
Going deeper, expect: “what number did you actually move, and by how much?” (have a from→to with a real magnitude ready — “handle time 12→7 min,” not “improved efficiency”); “who decided it was a success?” (the customer, against the pre-agreed metric on their dashboard — not you); and “the metric moved but the customer isn’t happy — what happened?” (you optimized a proxy, not the leading indicator they feel — re-anchor on the outcome, a classic exec-communication recovery). Translating every claim into a customer/KPI number, not personal effort, is the through-line interviewers score.

Could you write a three-layer Definition of Done for a real engagement and present a credible first-90-days plan anchored on customer-facing milestones?

New to itGetting thereConfident

Takeaways

  • Success is a number in the customer’s UI — a leading indicator with from→to, an eval that scores it, and a runbook when it’s wrong.
  • Name the acceptable error type (FP vs FN); “zero errors” is a vibe, not a spec.
  • The four-metric scorecard: deal velocity ≤90d, NRR 130%+, productization ≥1/engagement, 70%+ reusable assets in main by month 12.
  • First 90 days: week-1 small win → reusable asset by day 60 → cross-customer improvement by day 90.
  • The week-1 win is often a no-feature demo on the customer’s real auth — trust before substance.
  • Tie each phase to a handoff artifact (demo, pipeline+eval, productization memo, shipped feature), not a date.

Next: explaining technical tradeoffs to executives — the Pyramid Principle, SCQA, and the five-line exec one-pager.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.