Lesson 6 of 6 · 50 min

Capstone: an integrated assistant

Assemble the whole track into one engagement — deduped data, idempotent connectors, the agent layer, a streamed prototype — shipped on a customer’s clock with eval gates and guardrails. Then the full FDE interview loop: the coding round, the system-design round, and the client-facing behavioral round, scored on the five-dimension rubric.

The whole engagement, end to end

The capstone is the unified playbook an FDE walks into a customer with, in order: (1) meet the ops team for an hour and map the data + integration landscape on a board; (2) ship a 20-sample golden set and an eval-driven notebook (lesson 1); (3) wire the CRM/ticketing/doc connectors through a service that signs every outgoing request with an idempotency key and verifies every incoming webhook (lessons 2 + 4); (4) build the agent loop on function calling + structured outputs, eval-gated in CI (lesson 3); (5) ship a React/TS prototype on Vite + TanStack for the customer to click on their own data (lesson 5); (6) graduate the rough edges into the platform. This lesson runs that build as one integrated assistant — and then prepares you for the interview that hires for it.
The architecture of the integrated assistant is the whole track composed: a data layer (canonical, deduped, provenance-stamped), a connector layer (per-customer-scoped OAuth, idempotent writes, HMAC-verified CDC webhooks), an agent layer (function calling into those idempotent connectors, strict structured outputs, deterministic guardrails with an approval gate), an eval layer (golden set in CI, code-based + LLM-as-judge, production telemetry), and a prototype surface (streamed, secret-safe, with an evals-as-UI page). Everything you’ve built clips together — and the failure modes you avoided in each lesson are the ones that kill the 95% of pilots that die at the seam.
How Palantir Scaled: Why the Best Software Is Built Backwardsa16z Deep Dives

Shipping reliably and fast: the Palantir/OpenAI playbook

Four practices make the integrated build ship on a customer’s clock without accumulating silent debt. Demo-driven development: every cycle ends with a live demo on customer data — scope discipline by construction. Gravel road → paved highway: ship the script/dashboard/cron that solves the workflow now; productize the path only after the route is proven. Eval-driven development: the golden set is written first and gates every change — TDD for the stochastic parts. Hub-and-spoke: every customer attaches to a small set of centralized adapters with bespoke code in a per-customer profile, so the platform team doesn’t lag the field forever. The 12-rule senior playbook from the research compresses to: ship the smallest real workflow in a week, idempotency on every write, schema-adapt at the edge, full-jitter + DLQ, verify webhooks first, per-customer OAuth, eval set before prompt, codify rough edges into the platform, and pair engineering with change management — engineering-only deployments stall at the second team.

Time-to-first-value: the cadence you’re measured on

The research gives concrete, defensible targets across the public case-study corpus — and they’re aggressive because cycle time is the unit of FDE output. Palantir delivered COVID-19 flows in days; Morgan Stanley was a 6–8 week build. The numbers below are what “fast” actually means, and the through-line is unambiguous: speed comes from co-located engineering plus on-customer-data demos plus acceptance of a rough Phase 1 — not from waiting on the customer to change, unless you pair the engineering delivery with a change-management partner.
code
1TIME-TO-FIRST-VALUE TARGETS (Palantir/OpenAI-style engagement)23  Phase                                Conservative   Stretch    Anchor4  ----------------------------------   ------------   --------   ----------------------5  kickoff -> first demo on cust. data   7 days         2 days     Palantir COVID-196  -> first production rollout            30-60 days     14 days    Morgan Stanley 6-8wk7  -> user adoption > 50%                 90 days        60 days    Morgan Stanley 98%/4mo8  -> first Playbook codification         90 days/FDE    30 days    "eat pain, excrete product"910  Speed = co-located eng + demos on real data + a deliberately rough Phase 1.11  It does NOT come from customer-side change unless paired with change management.

Adoption is a technical risk, not a victory lap

A deployment is not “shipped” until the second internal team independently adopts it — first-team adoption usually rides on the FDE’s personal relationships and doesn’t generalize. The cautionary numbers from the research: Morgan Stanley’s 6–8 week build slipped into a 4-month trust-building phase (tone/citation mismatch) before reaching 98% adoption; a healthcare AI platform sat at 12% adoption 90 days post-launch. Three mitigations that work: production telemetry on every model call (prompt hash, response hash, latency, refusal flag) before you call the rollout complete; eval regression in CI on every prompt change; and guardrails as deterministic code wrapping the probabilistic output. Interview angle. “A client’s AI platform is at 12% adoption 90 days post-launch — what do you do?” The strong answer diagnoses with telemetry and user interviews, finds the workflow friction, and pairs the fix with change management — not “prompt it better.”
There’s a role-trap worth naming because it’s the most common way FDEs lose the plot: sliding into customer-success-manager work. The FDE is necessary-but-not-sufficient on the customer-facing side — you must run the exec readout and pair-debug a pdb session the same afternoon — but the moment you stop holding the engineering bar and become a relationship operator, you’ve become “sparkling sales engineering.” McCardel’s warning cuts both ways: companies cargo-cult the FDE title without the doctrine, and individuals cargo-cult the customer-facing part without the code. Hold both.

The FDE interview loop: three rounds, one rubric

Across Palantir, OpenAI, and Anthropic the loop has the same spine — a coding round, a system-design / re-engineering round, and a client-facing behavioral round — and they’re not independent gates: the interviewer reads all five rubric dimensions in every round. Budget prep roughly 35% coding / 35% system design / 30% client-facing. The coding round is practical engineering (rate limiters, parse messy CSVs, OAuth 1.0↔2.0 token exchange, a tool-call dispatcher, SQL deduping duplicate ingest rows), not tree inversions — and Palantir adds a re-engineering round (find bugs in 100–500 lines while ignoring red herrings). System design is anchored in real customer workflows (Foundry pipelines surviving schema drift, HIPAA RAG, MCP servers for CRM, Claude-at-100M-users) — ship a walking skeleton end-to-end first, then layer scale/security/observability. The client round role-plays the real job: deliver a slip to a CTO, align a VP on a contested metric, diagnose a “broken” pilot.
code
1THE FDE INTERVIEW LOOP (Palantir / OpenAI / Anthropic)23  Round              What it tests                         Strong signal4  ----------------   -----------------------------------   ----------------------------------5  Coding             practical eng under constraints       narrate continuously; brute force6                     (rate limiter, parse CSV, OAuth,      first then optimize; SQL window7                     tool dispatcher, SQL dedup)           functions; ~10% scoping then code8  System design      real customer workflows, not          walking skeleton FIRST; 3 options9                     "design Twitter"                      w/ explicit tradeoff; name what10                                                           breaks in prod (drift, injection)11  Re-engineering     read 100-500 lines, find bugs,        top-to-bottom read, ignore red12   (Palantir)        ignore red herrings                   herrings, surface root cause13  Client-facing      ambiguous, customer-facing            Acknowledge -> Diagnose -> Own;14                     situations under pressure             "I" + a risk accepted + a date1516  THE 5-DIMENSION RUBRIC (scored in EVERY round):17   ambiguity tolerance | customer empathy | shipping speed |18   production accountability | translation to non-technical stakeholders19   (+ at Anthropic: mission fit / ethical reasoning)
Palantir adds a round that doesn’t exist in standard SWE loops: re-engineering / problem decomposition — read 100–500 lines of unfamiliar code and find the bugs while ignoring red herrings (a reported case was a double-counting HashMap bug across a 250+ line codebase). Practice reading unfamiliar code top-to-bottom with a notepad, surfacing the root cause rather than cosmetically refactoring. The same round sometimes appears as an open decomposition prompt (“help elderly people with poor vision cook for themselves”) — they’re testing whether you can structure an ambiguous problem, not whether you know an algorithm.
Two scaffolds dominate the literature; memorize both. The 5-step deployment frame for design/ambiguity: clarify/scope → map stakeholders → identify data/constraints → propose tradeoffs → surface failure modes. The 3-step client frame “Acknowledge, Diagnose, Own”: acknowledge the customer’s problem with empathy, diagnose the root cause out loud, own the next step with a specific action, a date, and an owner. The single decisive question across every employer is “how do you know it’s working?” — pre-stage the answer (golden dataset + task-specific rubric + online eval + audit logs), because it’s asked in all three rounds. And ownership language separates strong from weak everywhere: “I” not “we,” a specific risk accepted and the logic behind it — not “I helped with” or “check the logs.”
The strongest verbatim example from the research, a Palantir client simulation: “A client VP demands a Foundry dashboard reporting ‘mission success rate’ tomorrow, but the metric depends on a missing join key and inconsistent event timestamps across two systems. How do you align stakeholders on a defensible definition and ship something that won’t be reversed next week?” The strong answer: force a one-page metric contract (definition, inclusion rules, time window, known gaps), get explicit sign-off, ship an MVP with caveats in the UI, and commit to a follow-up that closes the data gaps with dated owners. Notice that single paragraph hits all five rubric dimensions at once — stakeholder alignment, scope-first, MVP-with-caveats, a dated next step, and translating data reality to a non-technical VP.
The interviewer is often a real FDE who has seen real production failures and customer escalations. Generic answers read as remote; a one-page metric contract, a walking-skeleton week-one plan, an Acknowledge-Diagnose-Own rehearsal read as credible. Your prep is not memorization — it is the muscle memory of having actually done these on a real deployment.
The loops share a skeleton but diverge in emphasis, which is how you allocate prep. Palantir-leaning: more Python OOD, SQL on real-world data, and the re-engineering round. OpenAI-leaning: production realism in coding (latency, edge cases, instrumentation) and customer storytelling tied to a specific use case. Anthropic-leaning: the longest, most values-laden loop — progressive CodeSignal coding, MCP/RAG architecture, and explicit familiarity with the published safety writing, plus values-in-conflict and tough-feedback scenarios. The non-obvious common ground that generalizes everywhere: ownership language, scope-first answers, and the Acknowledge-Diagnose-Own frame.

Interview prep

This is the FDE loop itself. Rehearse the two scaffolds, pre-stage the “how do you know it’s working?” answer, and anchor every story to a specific deployment you owned — generic answers read as remote; a one-page metric contract and a walking-skeleton week-one plan read as credible. Research each employer’s customer base (Palantir-Gov leans defense; commercial Palantir leans pharma/finance/manufacturing; OpenAI leans consumer-grade reliability; Anthropic leans AI safety) and be honest about mission fit before you loop.
  1. 01“Why FDE, not a regular SWE role?” → I want to own the customer outcome end-to-end, embedded, shipping on their clock — name a specific deployment you owned.
  2. 02“Tell me about a deployment that went badly.” → own it in first person, name the specific change it produced; never blame external factors.
  3. 03“A VP wants a dashboard tomorrow on a contested metric.” → one-page metric contract + explicit sign-off + MVP with caveats + dated follow-up owners.
  4. 04“How do you know it’s working?” → golden dataset + task-specific rubric + online eval + audit logs — the same answer in all three rounds.
  5. 05“Design an assistant over the customer’s CRM + tickets + docs.” → walking skeleton first (data → agent → answer), then scoped OAuth, idempotent writes, verified webhooks, per-doc ACLs.
  6. 06“Deliver a 3-week slip to the CTO.” → early, with options and empathy; acknowledge → diagnose → own with a recovery plan and dates.
  7. 07“The pilot ‘doesn’t work.’ Diagnose it.” → it’s almost always the wrappers, not the model; telemetry + the golden set localize it, then fix + change-manage.
  8. 08“12% adoption at 90 days — what now?” → diagnose with telemetry + user interviews, fix the workflow friction, pair the fix with change management to the second team.
Follow-ups cluster on ownership and production realism: “what’s a technical decision you reversed, and what did you learn?” (intellectual honesty — name the reversal plainly), “describe your first 30/60/90 days” (1–30 learn + a small win, 31–60 own a deployment + a reusable integration, 61–90 a cross-customer improvement), “if the customer pulled the budget, where would you cut?” (keep the one workflow that proves value), and “how would you split the work between your team and theirs?” (you own adapters/agent/evals; they own canonical ownership, ACLs, IdP). At Anthropic specifically, expect values-in-conflict and tough-feedback scenarios plus familiarity with the published safety writing — mission engagement is scored, not assumed.
articleA Day in the Life of a Palantir Forward Deployed Software Engineer (the role, first-hand)PalantirarticlePalantir Forward Deployed Engineer Interview Guide (verbatim client-sim + SQL prompts)DataInterviewarticleUnderstanding Forward Deployed Engineering (culture, R&D-not-COGS, the cargo-cult warning)Barry McCardelarticleWhat are Forward Deployed Engineers, and why are they suddenly in demandThe Pragmatic EngineerarticleTrading Margin for Moat: Why the Forward Deployed Engineer Is Backa16z

Checkpoint

You’re asked to design an integrated assistant over a customer’s Salesforce, ServiceNow, and SharePoint. How do you open the system-design round?

AStart by detailing the vector DB choice, queue topology, and storage layoutBScope with clarifying questions, then ship a thin walking skeleton (data → agent → answer) end-to-end, then layer scoped OAuth, idempotent writes, verified webhooks, and per-doc ACLsCAsk which model they want to use and design around its context window
Sign up free to answer and see why

Checkpoint

The integrated assistant works in your demo but “doesn’t work” once it hits the customer’s production environment. Where do you look first?

AThe model — it must be hallucinating on their data, so swap to a bigger modelBThe wraparound systems — CRM lookup, ticketing validation, doc retrieval, SSO gating — using telemetry and the golden set to localize the failureCThe prompt — rewrite it until production behaves
Sign up free to answer and see why

Checkpoint

A client VP demands a dashboard tomorrow for a metric that depends on a missing join key and inconsistent timestamps. Strongest client-round answer?

AForce a one-page metric contract with sign-off, ship an MVP with caveats in the UI, and commit to closing the data gaps with dated ownersBExplain that the metric is impossible until the data is fixedCBuild whatever join you can overnight and present the number without caveats
Sign up free to answer and see why

Checkpoint

You’ve now written the same bespoke CRM-field-reconciliation hack on three customer accounts. What does the FDE doctrine say to do?

AKeep it in each customer’s profile — every customer is uniqueBCodify it into the platform repo (“eat pain, excrete product”) so future customers inherit it via a shared adapterCDelete two of them and standardize on the third customer’s version verbatim
Sign up free to answer and see why

Checkpoint

You cleared the coding and system-design rounds strongly. In the hiring-manager round you’re asked about working with this employer’s customers. What most determines the outcome now?

ARestating your technical wins in more detail to reinforce the strong roundsBHonest mission/values fit and customer empathy — a mismatch here is a soft reject that strong technical rounds cannot overcomeCNegotiating compensation to signal confidence
Sign up free to answer and see why

Could you build the integrated assistant end-to-end on a customer’s clock — and run the full FDE loop (coding, system design, client simulation) against the five-dimension rubric?

New to itGetting thereConfident

Takeaways

  • The integrated assistant is the whole track composed: data → connectors → agent → evals → prototype, shipped on the customer’s clock.
  • Ship via demo-driven + gravel-road + eval-driven + hub-and-spoke; codify repeated hacks back into the platform.
  • Adoption is a technical risk — it’s not shipped until the second team adopts; pair engineering with change management.
  • The FDE loop is three rounds (coding, system design, client) scored on one five-dimension rubric in every round.
  • Memorize the 5-step deployment frame and the 3-step Acknowledge-Diagnose-Own frame; pre-stage “how do you know it’s working?”.
  • Mission/values fit and ownership language are decisive — a mismatch is a soft reject that technical skill can’t overcome.

You’ve built the full-stack, enterprise-integration FDE toolkit — from messy data to a shipped assistant to the interview loop. Go deploy forward.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.