Lesson 6 of 6 · 50 min

Capstone: private RAG for a regulated customer

Put it together: a worked, defensible private-RAG design for a regulated healthcare customer in their VPC — the request lifecycle end to end, the discovery questions that frame it, the 30/60/90 deployment motion, and the full FDE system-design interview rubric you’ll be scored against.

The brief

The canonical FDE prompt, near-verbatim from the interview banks: “Design a private, VPC-deployed RAG system for a healthcare customer with HIPAA constraints and 50M documents.” A published variant adds the scale envelope: 1,000+ users, 10-100 q/s peak, 10-100M chunks, up to 1M new docs/day, P50 ≤ 1.5s, P95 ≤ 3s, freshness < 5 minutes. This capstone walks that build end to end, integrating every prior lesson, and then lays out the rubric you are scored against. The single most important move comes first and costs nothing: spend the first half of the answer on clarifying questions before you draw a box.
The clarifying questions are not stalling — they are how a forward-deployed engineer frames the whole design, and interviewers grade them directly. Ask: Who are the users and what is the PHI boundary? (clinicians vs billing vs patients — different ACLs). What is the compliance regime and residency? (HIPAA BAA chain; in-jurisdiction control plane?). What is the per-source freshness SLA? (a lab result is seconds-fresh; a policy doc is days). What is the auth boundary? (which IdP, which groups map to which corpora). What is the deployment posture they can afford? (PrivateLink + in-VPC gateway is the default; air-gap only on a real trigger). Each answer collapses a branch of the design. Interview angle. Drawing the architecture before asking these is the fastest way to be scored as “jumps to a solution.”
Building Sidekick: An AI Assistant People Actually Use (Felipe Leusin)Remix

The worked architecture — two planes, eight layers

Posture first: SaaS-managed model via PrivateLink, gateway in the customer’s VPC, customer-held KMS, vector store in their account. This clears HIPAA readiness (no third-party BAA for the model vendor when the data plane is in-account), deploys in ~2-4 weeks, and leaves the door open to self-hosted open-weights later because the model is a swappable endpoint. Now the two planes, with each layer carrying the lesson it came from.
code
1PRIVATE RAG for a HIPAA customer -- the deployable design23  CONTROL PLANE (decides + records)        DATA PLANE (does the work)4  ------------------------------------     ------------------------------------5  Identity  : Entra ID SSO, JIT users      Ingest : EHR/SharePoint -> layout6              identity -> agent + conns              parse -> chunk -> Presidio7  Policy    : gateway in VPC; SpiceDB                redact (Stage 1) -> embed8              ReBAC, post-filter on hits     Retrieve: hybrid BM25+vector, ReBAC9  Audit     : append-only signed log,                pre/post-filter, rerank,10              doc_ids+chunk_ids, authz +             Stage-2 PII re-check11              guardrail verdict, versions    Generate: Bedrock via PrivateLink,12  Guardrails: Presidio + NeMo + Bedrock              answer ONLY from retrieved13              GR; system>context>user                vetted corpus14  Observe   : self-hosted Langfuse in       Output : Stage-3 gateway tokenize;15              VPC; freshness SLO dashboard           groundedness guardrail1617  Egress check: every arrow stays in-account or on the private backbone.18  KMS = customer. DNS = private. Model = swappable. Logs = replayable.

Scale & latency under the SLO

For the 50M-chunk / 100 q/s envelope, the scale levers are familiar from the RAG-systems track but constrained by the security design. RAM: 50M chunks × 1,536-d × 4 bytes ≈ 300 GB raw before the HNSW graph — so dimension is a first-order cost decision (Matryoshka truncation or product quantization to fit budget). Tenancy: partition the index by department for hard isolation and to dodge the filtered-ANN recall cliff when a user sees a small slice. Freshness: a CDC ingestion path with a monitored source-mtime-vs-embed-mtime gap, because the customer’s <5-minute SLA is far tighter than the model’s native freshness assumption. Latency budget: tokenize → retrieve (ANN + ReBAC check) → rerank top candidates → generate, with the expensive reranker only on the shortlist to hold P95 ≤ 3s.
code
1LATENCY budget to hold P50 <= 1.5s / P95 <= 3s (illustrative, per stage)23  Stage                 Typical    Lever to protect it4  -------------------   --------   --------------------------------------5  tokenize + embed q    ~20-50ms   cache stable system prefix6  ANN retrieve (k=50)   ~10-40ms   HNSW efSearch tuned; partitioned index7  ReBAC post-filter     ~1-5ms     microsecond CheckPermission per hit8  PII re-check chunks   ~10-30ms   narrow scan, only surviving chunks9  rerank shortlist      ~50-150ms  cross-encoder on top ~50, NOT all k10  generate (P50/P95)    bulk       short output, stream; route reasoning11                                     model only for the hard tail1213  Output length dominates total latency; the reranker runs ONLY on the14  shortlist. Each guardrail/ACL stage is cheap because it scopes its input.
A subtle scale interaction the rubric rewards you for naming: the security layers actually help latency when placed correctly. The ReBAC post-filter and the Stage-2 PII re-check both operate on the small surviving candidate set, not the corpus, so they cost single-digit-to-tens of milliseconds; the reranker — the expensive stage — runs only on the shortlist. Put the cheap, scoping controls early and the expensive synthesis late, and the composed pipeline holds its P95 while still enforcing every control. The anti-pattern is running a full-corpus PII scan or an unbounded rerank per query; that is how a “secure” design blows its latency budget and gets value-engineered out in week three.

The request lifecycle, end to end

The single most persuasive thing you can do in the interview is walk one request through every layer, naming the artifact each produces. This proves the planes are real, not boxes on a slide.
code
1ONE REQUEST through the private RAG (a clinician asks about a patient)23  1. Auth      Entra ID SSO -> JWT asserts user:dr_kim; agent token scoped4               to dr_kim's privileges (least privilege)5  2. Policy    gateway resolves dr_kim's allowed corpora/matters via SpiceDB6  3. Guard-in  input screened: injection/jailbreak patterns, schema check7  4. Retrieve  hybrid BM25+vector over the dept-partitioned index ->8               top-k; POST-FILTER each hit with SpiceDB CheckPermission9  5. PII re-ck Stage-2 re-scan of surviving chunks (label may have shed)10  6. Generate  Bedrock via PrivateLink; prompt = system >  only11               permitted, re-checked chunks > ; answer from12               retrieved content only13  7. Guard-out groundedness >= SLO or block/route-to-human; Stage-3 tokenize14               any residual PII in the completion15  8. Audit     append signed event: user_id, doc_ids+chunk_ids returned,16               per-doc authz decision, guardrail verdicts, model+index17               versions, trace_id -- replayable for 6+ years1819  Fail any step -> that step's control is the next step's input. The boundary20  is the COMPOSITION; no single layer is the security boundary.
One more architectural decision the interviewer will probe: defend choosing managed Bedrock over self-hosting for this customer. The strong answer is not “Bedrock is better” — it is a reasoned trade: managed gives the compliance ceiling (HIPAA-ready via PrivateLink + in-account data plane) and the time-to-deploy the customer needs now, while keeping the model a swappable endpoint preserves the self-hosted migration for later. That is the Morgan Stanley / Shopify move precisely — buy control over the system around the model (evals, corpus, audit, guardrails) rather than over the model weights, and leave the weights replaceable. Pair the claim with its limit: managed inference adds per-token cost and a vendor in the decision chain, which is exactly why you kept the data plane in their account and the control plane model-agnostic.

The deployment motion: 30 / 60 / 90

The forward-deployed half of the role is judged on whether you can land this on a real customer, not just architect it. Rehearse the motion in the exact shape the interview banks reward. Days 1-30: learn the product, shadow the customer’s clinicians and security team, map the legacy landing zone (if the EHR is a 15-year-old system with no API, plan CDC or scheduled extracts — never a big-bang migration), and ship one small, real win (e.g. a single department’s policy-doc assistant with the full control plane). Days 31-60: lead a scoped pilot — wire the eval set from their tickets, demonstrate one real guardrail block with one audit entry, validate the freshness SLO. Days 61-90: own a measurable outcome (end-to-end task completion, not “the demo worked”), and hand the platform team the runbook, the SLOs, and the quarterly red-team agenda.
Two field scenarios the interview will throw at you, with the strong-answer shape. “Adoption is 12% after 90 days and the customer blames the product.” Diagnose before prescribing, assume the product has partial fault, instrument where users drop off, and design a 30-day pilot motion measured on task completion — do not get defensive. “Your system throws errors on screen during a live exec demo.” Own the moment publicly, triage with a kill switch or rollback, give a path backward, and never blame the customer. The hidden rubric across both: ownership language — “I did X,” with a named artifact and a metric, not “we helped with.”
The role is the engineer who ships and sits with the customer when it breaks. “We built a RAG” is the demo; “here is how I landed it in their VPC, measured it on their tickets, blocked the first leak with an audit entry, and owned the moment the live demo errored” is the deployment. The whole track exists to close that gap.
articleA Day in the Life of a Palantir Forward Deployed Software EngineerPalantirarticleDesign enterprise RAG search system (OpenAI interview question + scale envelope)PracHubpaperSeven Failure Points When Engineering a RAG SystemBarnett et al. (arXiv)

Checkpoint

You are 30 seconds into the HIPAA-RAG prompt. What is the strongest first move?

AAsk clarifying questions — users and PHI boundary, compliance/residency, per-source freshness SLA, auth boundary, affordable posture — then frame the design around the answersBImmediately draw the full architecture to show breadthCState that you would use the largest available model for accuracy
Sign up free to answer and see why

Checkpoint

For the 50M-chunk envelope with users who each see a small slice of departments, which combination correctly handles isolation and recall?

AOne shared flat index with a post-filter ACL on every queryBPartition the index by department, pre-filter to the user’s allowed set, and post-filter the hits with ReBAC — avoiding the filtered-ANN recall cliff and enforcing isolationCBake department tags into metadata and trust them as the only gate
Sign up free to answer and see why

Checkpoint

Asked to “walk one request end to end,” which sequence best demonstrates a defensible private RAG?

AEmbed the query, search the vector DB, send top-k to the LLM, return the answerBAuthenticate the user, then let the model decide which documents it is allowed to readCSSO/JWT → gateway resolves permissions → input guardrail → retrieve + post-filter ReBAC → PII re-check → generate from permitted chunks only → output guardrail + tokenize → append signed audit event
Sign up free to answer and see why

Checkpoint

A healthcare client’s adoption is 12% after 90 days and they blame the product. What is the strongest forward-deployed response?

ADiagnose before prescribing: assume partial product fault, instrument where users drop off, and design a 30-day pilot motion measured on end-to-end task completionBExplain that the product is fine and the users need more trainingCRecommend a larger model to improve answer quality
Sign up free to answer and see why

Checkpoint

The interviewer asks for one security trade-off you made on this deployment. Which answer best fits the strong-candidate rubric?

A“I never had to make a security trade-off — the design was secure by default.”B“We added differential-privacy noise to embeddings to defend against inversion of the PHI corpus, accepting a measured recall hit we recovered with a reranker — chosen because a breach of raw vectors was the higher risk.”C“We turned on every guardrail at maximum strictness regardless of impact.”
Sign up free to answer and see why

Interview prep

The FDE system-design interview scores three dimensions roughly equally — technical depth, real-world deployment thinking, and client-facing communication — and the capstone question touches all three. Run the self-checklist before you call an answer done: did I open with clarifying questions (scale, compliance, SLO)? did I name the posture and the two planes? did I put authorization in deterministic retrieval-time code? did I walk one request end to end? did I name one RBAC pattern, one guardrail layer, the stale-index SLO, and ≥5 of the seven failure points? did I bring one named, defended security trade-off? did I close with a 30/60/90 motion and an “I did” story with a metric?
  1. 01“Design a private VPC RAG for a HIPAA customer, 50M docs.” → clarify first; PrivateLink + in-VPC gateway posture; two planes; ACL at retrieval; redact 3 stages; signed audit; freshness SLO.
  2. 02“What’s your P50/P95 and how do you hold it?” → per-stage budget (tokenize/retrieve+ReBAC/rerank/generate); reranker only on the shortlist; cache stable prefixes.
  3. 03“How do you hit <5-minute freshness on 1M new docs/day?” → CDC ingestion, incremental re-embed, monitored source-vs-embed mtime gap, freshness-aware fallbacks.
  4. 04“Walk one request end to end.” → SSO → policy resolve → input guard → retrieve+post-filter ReBAC → PII re-check → generate from permitted chunks → output guard+tokenize → signed audit.
  5. 05“What are the failure modes?” → ≥5 of the seven FP plus permission leak, stale index, indirect injection, inversion — each mapped to its control.
  6. 06“One security trade-off you made?” → named + defended + its cost mitigated (e.g. DP-noised embeddings vs reranker recovery); never “I never made one.”
  7. 07“The legacy system has no API — how do you ingest?” → CDC or scheduled extracts; screen-scrape last resort; never big-bang.
  8. 08“What’s your 30/60/90?” → learn + shadow + one small win; lead a pilot with their eval set and a real guardrail block; own a measured outcome + hand off the runbook.
Going deeper. The deepest follow-ups fuse the dimensions. “Rank your eval metrics by business value vs feedback speed.” (groundedness/refusal are fast feedback; task-completion and a 99%-style SLA are closest to business value — say which you’d gate on). “Production is inconsistent but staging is fine — debug live.” (parallel data-diff and config-diff; bring a debugging tree, name embedding/prompt version drift). “Defend choosing Bedrock over self-hosting for this customer.” (compliance ceiling + time-to-deploy now, model-as-swappable-endpoint preserves the migration — buying control over the system around the model, the Morgan Stanley/Shopify pattern). Close every answer by handing the customer something concrete — a clearer SLO, a structured change request — not just a diagram.

Could you design and defend a private RAG for a regulated customer end to end — posture, two planes, one request lifecycle, the 30/60/90 motion, and a named security trade-off?

New to itGetting thereConfident

Takeaways

  • Clarify scope, compliance, and SLOs before drawing anything; the clarifying questions are graded.
  • Default posture for a regulated customer: SaaS-managed model via PrivateLink, gateway + data plane in their VPC, model swappable.
  • Walk one request through the composed controls — SSO → policy → guard-in → retrieve+ReBAC → PII re-check → generate → guard-out → signed audit.
  • No single layer is the security boundary; the boundary is the composition of the six chained controls.
  • The capstone is architecture + deployment motion (30/60/90) + eval/audit as production infra + customer-comms restraint.
  • Bring a named, defended security trade-off and an “I did” story with a metric — never “I never made trade-offs.”

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.