Put it together: a worked, defensible private-RAG design for a regulated healthcare customer in their VPC — the request lifecycle end to end, the discovery questions that frame it, the 30/60/90 deployment motion, and the full FDE system-design interview rubric you’ll be scored against.
The brief
The canonical FDE prompt, near-verbatim from the interview banks: “Design a private, VPC-deployed RAG system for a healthcare customer with HIPAA constraints and 50M documents.” A published variant adds the scale envelope: 1,000+ users, 10-100 q/s peak, 10-100M chunks, up to 1M new docs/day, P50 ≤ 1.5s, P95 ≤ 3s, freshness < 5 minutes. This capstone walks that build end to end, integrating every prior lesson, and then lays out the rubric you are scored against. The single most important move comes first and costs nothing: spend the first half of the answer on clarifying questions before you draw a box.
The clarifying questions are not stalling — they are how a forward-deployed engineer frames the whole design, and interviewers grade them directly. Ask: Who are the users and what is the PHI boundary? (clinicians vs billing vs patients — different ACLs). What is the compliance regime and residency? (HIPAA BAA chain; in-jurisdiction control plane?). What is the per-source freshness SLA? (a lab result is seconds-fresh; a policy doc is days). What is the auth boundary? (which IdP, which groups map to which corpora). What is the deployment posture they can afford? (PrivateLink + in-VPC gateway is the default; air-gap only on a real trigger). Each answer collapses a branch of the design. Interview angle. Drawing the architecture before asking these is the fastest way to be scored as “jumps to a solution.”
The worked architecture — two planes, eight layers
Posture first: SaaS-managed model via PrivateLink, gateway in the customer’s VPC, customer-held KMS, vector store in their account. This clears HIPAA readiness (no third-party BAA for the model vendor when the data plane is in-account), deploys in ~2-4 weeks, and leaves the door open to self-hosted open-weights later because the model is a swappable endpoint. Now the two planes, with each layer carrying the lesson it came from.
code
1PRIVATE RAG for a HIPAA customer -- the deployable design23 CONTROL PLANE (decides + records) DATA PLANE (does the work)4 ------------------------------------ ------------------------------------5 Identity : Entra ID SSO, JIT users Ingest : EHR/SharePoint -> layout6 identity -> agent + conns parse -> chunk -> Presidio7 Policy : gateway in VPC; SpiceDB redact (Stage 1) -> embed8 ReBAC, post-filter on hits Retrieve: hybrid BM25+vector, ReBAC9 Audit : append-only signed log, pre/post-filter, rerank,10 doc_ids+chunk_ids, authz + Stage-2 PII re-check11 guardrail verdict, versions Generate: Bedrock via PrivateLink,12 Guardrails: Presidio + NeMo + Bedrock answer ONLY from retrieved13 GR; system>context>user vetted corpus14 Observe : self-hosted Langfuse in Output : Stage-3 gateway tokenize;15 VPC; freshness SLO dashboard groundedness guardrail1617 Egress check: every arrow stays in-account or on the private backbone.18 KMS = customer. DNS = private. Model = swappable. Logs = replayable.
Scale & latency under the SLO
For the 50M-chunk / 100 q/s envelope, the scale levers are familiar from the RAG-systems track but constrained by the security design. RAM: 50M chunks × 1,536-d × 4 bytes ≈ 300 GB raw before the HNSW graph — so dimension is a first-order cost decision (Matryoshka truncation or product quantization to fit budget). Tenancy: partition the index by department for hard isolation and to dodge the filtered-ANN recall cliff when a user sees a small slice. Freshness: a CDC ingestion path with a monitored source-mtime-vs-embed-mtime gap, because the customer’s <5-minute SLA is far tighter than the model’s native freshness assumption. Latency budget: tokenize → retrieve (ANN + ReBAC check) → rerank top candidates → generate, with the expensive reranker only on the shortlist to hold P95 ≤ 3s.
code
1LATENCY budget to hold P50 <= 1.5s / P95 <= 3s (illustrative, per stage)23 Stage Typical Lever to protect it4 ------------------- -------- --------------------------------------5 tokenize + embed q ~20-50ms cache stable system prefix6 ANN retrieve (k=50) ~10-40ms HNSW efSearch tuned; partitioned index7 ReBAC post-filter ~1-5ms microsecond CheckPermission per hit8 PII re-check chunks ~10-30ms narrow scan, only surviving chunks9 rerank shortlist ~50-150ms cross-encoder on top ~50, NOT all k10 generate (P50/P95) bulk short output, stream; route reasoning11 model only for the hard tail1213 Output length dominates total latency; the reranker runs ONLY on the14 shortlist. Each guardrail/ACL stage is cheap because it scopes its input.
A subtle scale interaction the rubric rewards you for naming: the security layers actually help latency when placed correctly. The ReBAC post-filter and the Stage-2 PII re-check both operate on the small surviving candidate set, not the corpus, so they cost single-digit-to-tens of milliseconds; the reranker — the expensive stage — runs only on the shortlist. Put the cheap, scoping controls early and the expensive synthesis late, and the composed pipeline holds its P95 while still enforcing every control. The anti-pattern is running a full-corpus PII scan or an unbounded rerank per query; that is how a “secure” design blows its latency budget and gets value-engineered out in week three.
The request lifecycle, end to end
The single most persuasive thing you can do in the interview is walk one request through every layer, naming the artifact each produces. This proves the planes are real, not boxes on a slide.
code
1ONE REQUEST through the private RAG (a clinician asks about a patient)23 1. Auth Entra ID SSO -> JWT asserts user:dr_kim; agent token scoped4 to dr_kim's privileges (least privilege)5 2. Policy gateway resolves dr_kim's allowed corpora/matters via SpiceDB6 3. Guard-in input screened: injection/jailbreak patterns, schema check7 4. Retrieve hybrid BM25+vector over the dept-partitioned index ->8 top-k; POST-FILTER each hit with SpiceDB CheckPermission9 5. PII re-ck Stage-2 re-scan of surviving chunks (label may have shed)10 6. Generate Bedrock via PrivateLink; prompt = system > only11 permitted, re-checked chunks > ; answer from12 retrieved content only13 7. Guard-out groundedness >= SLO or block/route-to-human; Stage-3 tokenize14 any residual PII in the completion15 8. Audit append signed event: user_id, doc_ids+chunk_ids returned,16 per-doc authz decision, guardrail verdicts, model+index17 versions, trace_id -- replayable for 6+ years1819 Fail any step -> that step's control is the next step's input. The boundary20 is the COMPOSITION; no single layer is the security boundary.
One more architectural decision the interviewer will probe: defend choosing managed Bedrock over self-hosting for this customer. The strong answer is not “Bedrock is better” — it is a reasoned trade: managed gives the compliance ceiling (HIPAA-ready via PrivateLink + in-account data plane) and the time-to-deploy the customer needs now, while keeping the model a swappable endpoint preserves the self-hosted migration for later. That is the Morgan Stanley / Shopify move precisely — buy control over the system around the model (evals, corpus, audit, guardrails) rather than over the model weights, and leave the weights replaceable. Pair the claim with its limit: managed inference adds per-token cost and a vendor in the decision chain, which is exactly why you kept the data plane in their account and the control plane model-agnostic.
The deployment motion: 30 / 60 / 90
The forward-deployed half of the role is judged on whether you can land this on a real customer, not just architect it. Rehearse the motion in the exact shape the interview banks reward. Days 1-30: learn the product, shadow the customer’s clinicians and security team, map the legacy landing zone (if the EHR is a 15-year-old system with no API, plan CDC or scheduled extracts — never a big-bang migration), and ship one small, real win (e.g. a single department’s policy-doc assistant with the full control plane). Days 31-60: lead a scoped pilot — wire the eval set from their tickets, demonstrate one real guardrail block with one audit entry, validate the freshness SLO. Days 61-90: own a measurable outcome (end-to-end task completion, not “the demo worked”), and hand the platform team the runbook, the SLOs, and the quarterly red-team agenda.
Two field scenarios the interview will throw at you, with the strong-answer shape. “Adoption is 12% after 90 days and the customer blames the product.” Diagnose before prescribing, assume the product has partial fault, instrument where users drop off, and design a 30-day pilot motion measured on task completion — do not get defensive. “Your system throws errors on screen during a live exec demo.” Own the moment publicly, triage with a kill switch or rollback, give a path backward, and never blame the customer. The hidden rubric across both: ownership language — “I did X,” with a named artifact and a metric, not “we helped with.”
The role is the engineer who ships and sits with the customer when it breaks. “We built a RAG” is the demo; “here is how I landed it in their VPC, measured it on their tickets, blocked the first leak with an audit entry, and owned the moment the live demo errored” is the deployment. The whole track exists to close that gap.
You are 30 seconds into the HIPAA-RAG prompt. What is the strongest first move?
AAsk clarifying questions — users and PHI boundary, compliance/residency, per-source freshness SLA, auth boundary, affordable posture — then frame the design around the answersBImmediately draw the full architecture to show breadthCState that you would use the largest available model for accuracy
For the 50M-chunk envelope with users who each see a small slice of departments, which combination correctly handles isolation and recall?
AOne shared flat index with a post-filter ACL on every queryBPartition the index by department, pre-filter to the user’s allowed set, and post-filter the hits with ReBAC — avoiding the filtered-ANN recall cliff and enforcing isolationCBake department tags into metadata and trust them as the only gate
Asked to “walk one request end to end,” which sequence best demonstrates a defensible private RAG?
AEmbed the query, search the vector DB, send top-k to the LLM, return the answerBAuthenticate the user, then let the model decide which documents it is allowed to readCSSO/JWT → gateway resolves permissions → input guardrail → retrieve + post-filter ReBAC → PII re-check → generate from permitted chunks only → output guardrail + tokenize → append signed audit event
A healthcare client’s adoption is 12% after 90 days and they blame the product. What is the strongest forward-deployed response?
ADiagnose before prescribing: assume partial product fault, instrument where users drop off, and design a 30-day pilot motion measured on end-to-end task completionBExplain that the product is fine and the users need more trainingCRecommend a larger model to improve answer quality
The interviewer asks for one security trade-off you made on this deployment. Which answer best fits the strong-candidate rubric?
A“I never had to make a security trade-off — the design was secure by default.”B“We added differential-privacy noise to embeddings to defend against inversion of the PHI corpus, accepting a measured recall hit we recovered with a reranker — chosen because a breach of raw vectors was the higher risk.”C“We turned on every guardrail at maximum strictness regardless of impact.”
The FDE system-design interview scores three dimensions roughly equally — technical depth, real-world deployment thinking, and client-facing communication — and the capstone question touches all three. Run the self-checklist before you call an answer done: did I open with clarifying questions (scale, compliance, SLO)? did I name the posture and the two planes? did I put authorization in deterministic retrieval-time code? did I walk one request end to end? did I name one RBAC pattern, one guardrail layer, the stale-index SLO, and ≥5 of the seven failure points? did I bring one named, defended security trade-off? did I close with a 30/60/90 motion and an “I did” story with a metric?
01“Design a private VPC RAG for a HIPAA customer, 50M docs.” → clarify first; PrivateLink + in-VPC gateway posture; two planes; ACL at retrieval; redact 3 stages; signed audit; freshness SLO.
02“What’s your P50/P95 and how do you hold it?” → per-stage budget (tokenize/retrieve+ReBAC/rerank/generate); reranker only on the shortlist; cache stable prefixes.
03“How do you hit <5-minute freshness on 1M new docs/day?” → CDC ingestion, incremental re-embed, monitored source-vs-embed mtime gap, freshness-aware fallbacks.
04“Walk one request end to end.” → SSO → policy resolve → input guard → retrieve+post-filter ReBAC → PII re-check → generate from permitted chunks → output guard+tokenize → signed audit.
05“What are the failure modes?” → ≥5 of the seven FP plus permission leak, stale index, indirect injection, inversion — each mapped to its control.
06“One security trade-off you made?” → named + defended + its cost mitigated (e.g. DP-noised embeddings vs reranker recovery); never “I never made one.”
07“The legacy system has no API — how do you ingest?” → CDC or scheduled extracts; screen-scrape last resort; never big-bang.
08“What’s your 30/60/90?” → learn + shadow + one small win; lead a pilot with their eval set and a real guardrail block; own a measured outcome + hand off the runbook.
Going deeper. The deepest follow-ups fuse the dimensions. “Rank your eval metrics by business value vs feedback speed.” (groundedness/refusal are fast feedback; task-completion and a 99%-style SLA are closest to business value — say which you’d gate on). “Production is inconsistent but staging is fine — debug live.” (parallel data-diff and config-diff; bring a debugging tree, name embedding/prompt version drift). “Defend choosing Bedrock over self-hosting for this customer.” (compliance ceiling + time-to-deploy now, model-as-swappable-endpoint preserves the migration — buying control over the system around the model, the Morgan Stanley/Shopify pattern). Close every answer by handing the customer something concrete — a clearer SLO, a structured change request — not just a diagram.
Could you design and defend a private RAG for a regulated customer end to end — posture, two planes, one request lifecycle, the 30/60/90 motion, and a named security trade-off?
New to itGetting thereConfident
Takeaways
Clarify scope, compliance, and SLOs before drawing anything; the clarifying questions are graded.
Default posture for a regulated customer: SaaS-managed model via PrivateLink, gateway + data plane in their VPC, model swappable.
Walk one request through the composed controls — SSO → policy → guard-in → retrieve+ReBAC → PII re-check → generate → guard-out → signed audit.
No single layer is the security boundary; the boundary is the composition of the six chained controls.
The capstone is architecture + deployment motion (30/60/90) + eval/audit as production infra + customer-comms restraint.
Bring a named, defended security trade-off and an “I did” story with a metric — never “I never made trade-offs.”