Lesson 1 of 6 · 48 min

Enterprise RAG architecture

The forward-deployed deliverable is not “an LLM in a VPC” — it is a control plane (auth, policy, audit, guardrails) wrapped around a data plane (ingest, retrieve, generate), both landed in the customer’s environment. Why you ship a gateway and a corpus, not a model; the layered reference architecture; and how Morgan Stanley actually built it.

What you’re actually deploying

The senior mental model up front: a forward-deployed RAG engagement is not “stand up a vector DB and call an LLM.” It is landing a control plane — auth, authorization, audit, guardrails, observability — wrapped around a data plane — ingest, retrieve, generate — inside the customer’s environment, with their keys, their buckets, their DNS, their model. The data plane does the work; the control plane decides whether the work is allowed to leave and logs what happened when it did. Most failed enterprise deployments shipped a clever data plane and bolted security on after — which is exactly backwards, and the part a system-design interview will push on hardest.
Why does the consumer-grade shape fail in the enterprise? Because the moment your buyer is a bank, a hospital, or a defense contractor, “the model gave a good answer” is table stakes and the real contract is accountability: every retrieval and every generation must be attributable to a named user, governed by a policy you can point at, and replayable by an auditor six months later. A demo proves the model can answer; the deployment proves the system can be trusted, governed, and audited. That shift — from “can it answer” to “can it be defended” — is the entire job.
The clearest way to see the decomposition is to ask, for any request, who produces a signed, attributable artifact? Identity (the IdP) attests who is asking. The policy plane decides what they may see. Retrieval returns chunks filtered to that policy. The generator synthesizes only from those chunks. Guardrails screen both the inbound prompt and the outbound answer. And the audit layer appends an append-only, signed record keyed to the user identity. That chain — not the model — is what turns a chatbot into a SOC 2 Type II-auditable system. Interview angle. When asked to “design an enterprise RAG,” your first move is to draw two planes and name the control-plane components; candidates who draw only the retrieval pipeline have already signaled they’ve built a demo, not a deployment.
How Morgan Stanley deploys AI that actually worksScale AI

The layered reference architecture

Across documented enterprise deployments the architecture stabilizes into eight layers, each with one job and one customer-VPC implication. Memorize this table — it is the skeleton you draw on the whiteboard, and every later lesson zooms into one row of it.
code
1ENTERPRISE RAG -- the eight layers, each producing a signed, attributable artifact23  Layer        Component                         Job                         VPC implication4  ----------   -------------------------------   -------------------------   ------------------------5  Identity     SSO + IdP (Okta, Entra ID)        authenticate the end user   JIT user, not shared acct6  Policy       gateway / SpiceDB / OPA           RBAC + ABAC + ReBAC on      gateway sees internal ACLs7                                                  tools AND documents          -> must run IN the VPC8  Ingest       loader -> parse -> chunk -> embed ETL S3/SharePoint/Notion    runs in VPC; writes vectors9  Retrieval    hybrid (BM25 + vector) + rerank   pull top-k relevant chunks  ACL filter injected here10  Generation   LLM (Bedrock / Azure OpenAI /     synthesize from context     PrivateLink / private11               open-weights on EKS)                                           endpoint -- never public12  Eval+Observe RAGAS + Langfuse + judges         offline + online scoring    self-hosted collector13  Guardrails   Presidio + NeMo + Bedrock GR      PII, injection, jailbreak   all 3 layers, never one14  Audit        SIEM egress (Splunk/Sentinel)     append-only signed log      WORM bucket, user-keyed1516  The mechanism is accountability: every layer emits an artifact keyed to user identity.
The non-obvious senior point hides in the Policy row: the gateway must run inside the customer VPC, because the authorization decision needs to see the customer’s internal dataset ACLs (who is in which SharePoint group, which matter a lawyer is staffed on) — data that never leaves their network. A SaaS control plane that physically sits in-region but is operated from outside fails strict supervisory tests (German BfDI, French CNIL, Swiss FINMA all distinguish “data in jurisdiction” from “control plane in jurisdiction”). So the policy plane is the one component you can almost never outsource — it is co-located with the data by definition.
Read the table top-to-bottom and a second pattern appears: the order is also the order in which a request is governed. Identity establishes who; policy narrows what; ingest and retrieval produce the candidate context; generation synthesizes; eval/observe and guardrails screen; audit records. A request can fail at any layer, and each layer’s output is the next layer’s input — which is why a weakness early (say, retrieval that does not see identity) cannot be patched late (a guardrail that scans the answer). This is the single idea the whole track returns to: the security boundary is the composition of the layers, not any one of them.

Put authorization in the gateway, not the model

When the model can call dozens of tools and read a private corpus, the worst possible place to put authorization is in the prompt (“only use documents the user is allowed to see”). The model is a probabilistic text generator; a prompt-side instruction is an aspiration, not a control, and a user who phrases their request the right way will route around it. The architectural fix, drawn straight from humanlayer’s widely-cited 12-Factor Agents, is that “effective agents are primarily comprised of software” — the LLM merely emits structured outputs that deterministic code executes. The gateway is that deterministic code: the LLM never touches the database directly, it requests a tool call the gateway decides whether to honor.
This inverts the typical RAG posture, where vector search lives client-side and the model ingests raw chunks. In the enterprise shape, retrieval is a tool behind the gateway, and the gateway enforces RBAC on the call, redacts PII on the way out, and logs the whole interaction. MintMCP — a productized version of this pattern launched Feb 2026 — frames it as an MCP Gateway (fronting connectors, “made enterprise-ready”) plus an Agent Monitor giving “real-time tracing of every tool call, command, and file access across all agent activity.” You do not have to buy it, but its decomposition is the reference: a policy plane in front of the model and the data, not behind them. Interview angle. “Where do you enforce permissions in a RAG system?” The strong answer is “at retrieval, in deterministic code the model can’t bypass — never as a prompt instruction.” Saying “I’d tell the model to respect permissions” is a red flag that ends the security portion of the round.
Why Your RAG System Is Broken, and How to Fix It (Jason Liu)The TWIML AI Podcast

Case study: Morgan Stanley’s locked-down RAG

Morgan Stanley’s “AI @ Morgan Stanley Assistant” (announced with OpenAI in March 2023) is the canonical template, and it maps cleanly onto the eight layers. The single most important architectural decision: GPT-4 “generates responses exclusively from internal Morgan Stanley content” over a corpus of hundreds of thousands of internal research pages, with “appropriate guardrails.” By restricting the reachable search space to a vetted corpus, they sidestep the entire category of “the model said something embarrassing from the open web” — the failure mode is engineered out by the retrieval boundary, not patched by the prompt.
The system disaggregates into exactly the planes above: a RAG pipeline over the vetted corpus (data plane), the advisor-facing assistant plus a separate “Debrief” meeting-summarizer (application), and — critically — an eval loop where they built “hundreds of thousands of questions” indexed against the corpus, ran the model over them, and had financial advisors “flag any incorrect, sensitive, or otherwise problematic responses for investigation.” That flagging loop is the control plane’s feedback mechanism: every flagged answer points at a specific corpus chunk they can re-embed or re-chunk. The lesson for a forward-deployed engineer: build the customer’s first eval set out of their own support tickets and a frozen corpus, not the vendor’s defaults — and make the corpus boundary the primary safety control.
Restrict the reachable search space to trustworthy corpora, then layer fine-grained ACLs on top. Morgan Stanley’s whole safety story is “the model can only see our vetted content” — the retrieval boundary does the work the prompt cannot.
There is a productized analogue worth naming because it crystallizes the pattern: MintMCP (launched Feb 2026) ships an MCP Gateway fronting 100+ hosted connectors plus an Agent Monitor advertising “real-time tracing of every tool call, command, and file access,” “role-based access control for your entire organization,” “PII detection and secret scanning,” and “SOC 2 Type II” with “data residency options.” Its own deployment note is instructive for an FDE: it is “SaaS-first ... though VPC and self-hosted options are available upon request.” You do not have to adopt it, but its decomposition is the reference — policy plane and audit plane in front of the model — and scoping a pilot without a gateway-shaped abstraction tends to entrench the wrong one (raw prompt plus a bag of tools) that is very hard to unwind later.

The three failure modes that define the work

Across published field reports, enterprise RAG fails in three recurring shapes — and the rest of this track is organized around defending against each. Naming them precisely is half of a senior diagnosis.
  1. 01Permission leak — the retriever returns a document the user may not read, because the vector DB has no idea who is asking (Lesson 2). The single most common production security bug.
  2. 02Stale index — “a stale index turns the model into a confident generator of out-of-date answers, the failure mode RAG was supposed to eliminate.” Silent: nothing errors (Lessons 4 & 6).
  3. 03Consumer-endpoint leak — sensitive data pasted into a public SaaS LLM with no tenant control, no audit, no redaction. The Samsung failure; the reason private deployments exist (Lessons 3 & 5).
The Samsung incident (March 2023) is the cautionary standard for the third shape: an engineer pasted internal source code into ChatGPT to debug it, the data was retained vendor-side, and within weeks Samsung banned generative AI on staff devices. The architectural reading is not “AI is dangerous” — it is that a consumer endpoint has no row-level tenant control, no audit trail of input content, and no PII redaction layer. Anything sensitive must terminate in a VPC you control. Interview angle. If asked “why not just use ChatGPT Enterprise for this regulated customer?”, the strong answer names the three missing controls and the regulatory residency tests — not a vague “it’s not secure.”

Forward-deployed posture: pick it before you pick a vendor

The deployment posture shapes the vendor list, the compliance ceiling, and your own engineering burden — so a senior FDE chooses it in discovery, before drawing any boxes. The mechanism that orders the whole spectrum is who owns the keys, the buckets, and the indexes: the more the customer owns, the lower the vendor blast radius, but the higher the SRE and eval burden you take on.
code
1DEPLOYMENT POSTURE -- pick this in discovery, it shapes everything downstream23  Posture                         Compliance   Vendor    FDE       Time to    Customer4                                  ceiling      lock-in   control   deploy     RAG effort5  -----------------------------   ----------   -------   -------   --------   ----------6  Public SaaS, SSO-bound          moderate     heavy     low       days       lowest7  (ChatGPT Enterprise)            (SOC 2)8  Cloud LLM via PrivateLink       high         medium    high      2-4 wks    medium9  + gateway in VPC (Bedrock)      (HIPAA-rdy)10  BYO-VPC, open-weights LLM       highest      low       highest   6-12 wks   high11  (Llama on EKS/AKS)              (sovereign)12  Air-gapped (no internet)        max absolute none      total     3-6+ mo    highest1314  Rule: pick the LOWEST lock-in posture the customer can afford in engineering hours,15        then build the eval + audit plane FIRST, before committing to a model.
A divergence worth surfacing on a customer call: Morgan Stanley and Shopify both run on cloud-managed frontier models (OpenAI hosted, respectively) yet invest heavily in proprietary eval frameworks and curated corpora. They are buying control over the system around the model, not the model itself. Conversely, Klarna deployed a hosted-OpenAI assistant that handled 2.3M conversations in month one (the work of ~700 agents) and by mid-2025 was re-balancing toward private models — because the cost-and-quality surface of unbounded hosted inference became a forced variable. The general rule: leave architectural room for the customer to migrate from hosted Bedrock to self-hosted Llama without re-platforming — which you get for free if the control plane never depended on the model vendor.
repo12-Factor Agents — principles for reliable LLM applicationshumanlayerarticleMorgan Stanley — AI evals to shape the future of financial servicesOpenAIarticleRAG data storage for enterprise AI: a design guide (stale-index failure)Scality

Checkpoint

A regulated customer asks you to “deploy your RAG product into our VPC.” What is the strongest framing of the actual deliverable?

AA vector database and an LLM endpoint running on their cloud accountBA control plane (identity, policy, audit, guardrails) wrapped around a data plane (ingest, retrieve, generate), with the model as a swappable config choiceCA fine-tuned model trained on the customer’s documents so the knowledge lives in the weights
Sign up free to answer and see why

Checkpoint

In your architecture, where must the authorization decision (which documents this user may see) be computed?

AIn the system prompt — instruct the model to only use documents the user is allowed to seeBIn the embedding metadata at ingest time, baked in onceCIn deterministic gateway code at retrieval time, which the model cannot bypass
Sign up free to answer and see why

Checkpoint

A CISO objects: “Why can’t we just use ChatGPT Enterprise for this?” Which response best reflects the architectural reality?

AIt names the three missing controls — no row-level tenant control over the corpus, no audit trail of input content, no PII redaction layer — plus the data-residency tests it failsBChatGPT Enterprise is simply not secure and should never be used in any enterpriseCA bigger or newer model would solve the customer’s concerns
Sign up free to answer and see why

Checkpoint

You are scoping a deployment for a customer who may later need to move off a hosted model to self-hosted open-weights. What design choice protects that path best?

AStandardize on the hosted model’s proprietary features so the integration is as tight as possibleBTreat the model as a swappable endpoint behind the gateway, so identity, policy, retrieval, and audit never depend on the model vendorCAvoid the question — model choice never changes once a deployment ships
Sign up free to answer and see why

Checkpoint

Which posture choice would you recommend for a customer needing HIPAA readiness, a 3-week timeline, and minimal SRE burden on their side?

AFully air-gapped, open-weights LLM on customer GPUsBPublic SaaS LLM with SSOCCloud-managed LLM via PrivateLink with the gateway in their VPC
Sign up free to answer and see why

Interview prep

The FDE system-design round on enterprise RAG scores three dimensions roughly equally: technical depth, real-world deployment thinking, and client-facing communication. The architecture question is where you set the frame for the whole interview. Open with clarifying questions (who are the users, what is the compliance regime, what is the freshness SLA, what is the auth boundary) before any boxes, then draw the two planes and reason about tradeoffs out loud. Lead with the mechanism, then the production implication, then a named case.
  1. 01“Design a private, VPC-deployed RAG for a healthcare customer, HIPAA, 50M docs.” → name the posture (PrivateLink + gateway in VPC), then the two planes; PHI redaction at ingest + output, embeddings encrypted at rest, ACL re-check at retrieval, append-only audit.
  2. 02“What is the actual deliverable in a customer VPC?” → a control plane (identity/policy/audit/guardrails) around a data plane; the model is a config choice.
  3. 03“Where do you enforce permissions?” → in deterministic gateway code at retrieval time, never as a prompt instruction or frozen embedding metadata.
  4. 04“Why not ChatGPT Enterprise?” → no row-level tenant control over the corpus, no audit of input content, no redaction layer; plus residency (data-in-jurisdiction ≠ control-plane-in-jurisdiction).
  5. 05“What are the main failure modes of enterprise RAG?” → permission leak, stale index, consumer-endpoint leak (Samsung) — and the control each maps to.
  6. 06“How do you keep the customer from getting locked into a model?” → model as swappable endpoint behind a model-agnostic control plane; migration is a config change.
  7. 07“What does ‘deployed’ mean to you?” → producing SOC 2 Type II-relevant audit logs a third party could replay — not “the demo answered.”
  8. 08“Walk me through one request end to end.” → IdP attests who → gateway computes policy → retrieval ACL-filters → generator uses only those chunks → guardrails screen in/out → audit appends a signed, user-keyed record.
Going deeper. Expect the follow-ups that separate “read about it” from “shipped one.” “The customer’s legacy system is a 15-year-old on-prem ERP with no API — how do you ingest it?” (change-data-capture or scheduled extracts; screen-scraping as a last resort; never a big-bang migration). “Adoption is 12% after 90 days and they blame the product — what do you do?” (diagnose before prescribing, assume partial product fault, design a 30-day pilot motion measured on end-to-end task completion). “Their system returns errors on screen during a live exec demo — what do you say and do?” (own the moment, triage publicly, offer a path backward — kill switch / rollback — never blame the client). The hidden rubric is ownership language: “I did,” with a named artifact and a metric — not “we helped with.”

Could you whiteboard the two planes, place authorization in the gateway, pick a posture from constraints, and walk one request end to end?

New to itGetting thereConfident

Takeaways

  • Ship a control plane (identity/policy/audit/guardrails) around a data plane (ingest/retrieve/generate) — not “an LLM in a VPC.”
  • Authorization lives in deterministic gateway code at retrieval time; the model emits requests it cannot self-authorize.
  • Make the model a swappable endpoint so the customer can migrate without re-platforming.
  • Morgan Stanley’s safety story is the corpus boundary (answer only from vetted internal content) plus an advisor-flagging eval loop.
  • Three failure modes organize the track: permission leak, stale index, consumer-endpoint leak (Samsung).
  • Pick the deployment posture in discovery; “deployed” means audit logs a third party could replay.

Next: auth, SSO, RBAC & auditability — permission-aware retrieval and the audit trail that survives a regulator.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.