The reference stack that decides whether an LLM solution survives an audit, a Black Friday, or a prompt-injection incident — the seven layers, multi-tenancy, the SLO-first design move, streaming RAG vs nightly ETL, and the architecture round interviewers actually run.
The model is one component, not the system
A senior mental model up front: an enterprise AI solution succeeds or fails on the supporting stack — retrieval, evals, connectors, observability, governance — not on which foundation model you picked. The durable system-design pillars (caching, rate limiting, idempotency, back-pressure, a non-negotiable security layer) are more important once an LLM is in the loop, not less, because the model adds nondeterminism and per-token latency on top. This lesson is the architecture round of a solutions-engineer interview, and the whiteboard you will draw in front of a customer.
Why does the stack dominate the model? Because the LLM is a lossy compression of the public web at training time — it has no row for the customer’s contracts, their ticket history, or yesterday’s policy change, and it is probabilistic, slow per token, and arbitrary-code-adjacent the moment it can call a tool. Every property a customer cares about — freshness, attribution, access control, latency SLO, cost ceiling, auditability — is delivered by a layer around the model, not by the weights. The solutions engineer who opens a design review by debating GPT-vs-Claude has already lost the room; the one who opens with the seven layers and the SLOs has earned it.
The canonical reference is the agents-towards-production playbook, which is explicit that production-grade AI is end-to-end: build → evaluation → deployment → monitoring → security, and that these phases are a non-skippable checklist, not a menu. Treat its phases as the scope-setting artifact in a customer engagement: map the customer’s requirements to the phases, and switching phases reactively (shipping before you have an eval, adding monitoring after the incident) is exactly what produces slow time-to-value on enterprise deals.
A workable enterprise AI architecture has seven layers, and being able to draw them in order — naming what each one buys you — is the single highest-signal move in an architecture round. Each layer maps to one of the interviewer’s scoring axes (problem decomposition, LLM-aware architecture, data/context strategy, reliability, eval/monitoring, security/privacy, cost/latency).
01Edge / API gateway — TLS termination, WAF, OAuth/OIDC, per-token rate limiting, request/response logging. The denial-of-wallet boundary.
02Orchestrator — agent runtime, plan-execute loop, conversation memory, tool registry. Treat it like a microservice mesh: retries, timeouts, circuit breakers.
04Model layer — a routing abstraction over one or more foundation models, with exact-prefix caching, semantic caching, and explicit fallback to a cheaper model.
05Tools / connectors — every external system exposed as an MCP tool or resource (Lesson 4).
The unwritten rule from the forward-deployed-engineer rubric is “propose a walking skeleton before iterating.” Enumerate the data flow end to end, then turn on retrieval, then tools, then eval, then guardrails — each layer maps to one rubric category, and you light them up one at a time. Interview angle. The fastest move from “No Hire” to “Hire” on a whiteboard is to start with the skeleton and explicitly not start with fine-tuning. Jumping to “we’ll fine-tune on their data” as a first move is a documented weak-answer tell: it treats the LLM as the system and skips the layers that actually make it enterprise-grade.
code
1THE WALKING SKELETON (turn layers on one at a time)23 request -> [gateway] -> [orchestrator] -> [model] -> response (1) baseline4 |5 +-> [retrieval] ground the answer (2) + RAG6 +-> [tools/MCP] take actions (3) + connectors7 +-> [evals/guardrails] gate quality (4) + safety8 +-> [observability] trace every span (5) + telemetry910 Each layer maps to ONE architecture-rubric axis. Name the axis as you add it.11 Weak answer: open with "fine-tune the model." Strong: open with the skeleton.
Multi-tenancy is what separates enterprise work from a demo
A demo serves one tenant; an enterprise solution serves hundreds with contractual data-isolation requirements, and the model you choose for tenancy drives cost, blast radius, and which compliance questions you can answer “yes” to. Three viable models, in increasing isolation and cost:
code
1MULTI-TENANCY MODELS (pick the cheapest that meets the contract)23 Model Isolation Cost When to use4 ---------------------------- --------- ------ --------------------------5 (a) Shared everything, low lowest internal tools, low-reg6 per-tenant namespace + RLS buyers; hardest to scale7 (b) Shared model, medium medium DEFAULT for enterprise:8 isolated retrieval (RBAC) strict metadata filters9 metadata filter per query on every vector query10 (c) Per-tenant replicas + high highest only when a contract11 dedicated model endpoints forces data isolation1213 Confluent's reference: enforce RBAC at the RETRIEVAL layer, not the prompt.
Default to (b) and escalate to (c) only when a contract forces it. The load-bearing detail, straight from Confluent’s enterprise-RAG reference architecture: zero-trust retrieval and strict RBAC metadata filtering must be enforced at the retrieval layer, not the prompt layer. A permission check written into the system prompt is not a permission check — the model can be talked out of it. A metadata predicate on the vector query (filter to tenant_id and the caller’s acl before similarity even runs) is a real boundary. Interview angle. “How do you stop tenant A from retrieving tenant B’s documents?” → filter-then-retrieve at the index with the caller’s identity, never “the prompt tells it not to.” Saying the latter is a red flag that you have not run multi-tenant retrieval in production.
Design backwards from the SLO, not forwards from the model
The one design choice that pays off disproportionately is gating model selection on prompt size and SLO. A request that needs 8K of context to answer a hard, multi-step question should never sit on the same hot path as a short classification — route them to different models and different caches. Four numbers decide the architecture, and you should state them out loud before drawing a box: time-to-first-token, total latency, accuracy, and $/request. They pick the model and the cache strategy; the model does not pick them.
Two levers do most of the work. Semantic caching of prior similar queries can drop average $/request by an order of magnitude on repetitive enterprise traffic (the same fifty support questions asked a thousand ways). And back-pressure — the message-queue/topic pattern from the system-design primer — is what keeps a viral moment or a runaway agent loop from melting your inference budget; without it, an LLM endpoint is a denial-of-wallet target where each unbounded request costs real money. Interview angle. “Your AI feature’s bill 10×’d overnight with flat traffic — what happened and how is it bounded?” → an agent loop or retry storm with no per-token rate limit and no cap on output tokens; the fix is gateway rate limiting, max-tokens, a cost budget per session, and back-pressure on the queue.
code
1WHAT DECIDES THE DESIGN (state these before you draw a box)23 Lever Target Architectural consequence4 -------------- --------------- ----------------------------------5 TTFT e.g. < 500 ms shrink/cACHE the prompt (prefill)6 Total latency e.g. p95 < 2 s smaller model on hot path; stream7 Accuracy eval gate RAG + reranker + eval harness8 $/request cost ceiling semantic cache, model routing,9 token budget, cheaper fallback1011 Route by prompt size + SLO: short/cheap classification and long/hard12 reasoning do NOT belong on the same model or the same cache.
Streaming RAG beats nightly ETL — and the freshness SLA drives cost
The chronic failure of enterprise RAG is the “stale brain”: a nightly batch ETL re-indexes documents, so the assistant confidently answers from yesterday’s policy after it changed at 9am. Confluent’s reference architecture replaces batch document ETL with change-data-capture (CDC) + Apache Flink for semantic chunking and PII redaction, then streams into a horizontally scalable vector store (Pinecone, Milvus, Qdrant) with real-time upserts so vectors stay synchronized with source systems in near real time. Ragie’s production-RAG guide adds the complementary point that fine-tuning chunking, reranking, and the orchestration layer are the dominant levers for retrieval quality — not the embedding model.
The tradeoff is honest and you should name it: streaming ingestion is the right architecture for enterprise freshness, but it is operationally heavier than a cron job and requires Flink (or equivalent) expertise. The senior move is to let the freshness SLA decide. A knowledge base that changes weekly is fine on a nightly batch and cheap; a system where employees edit source docs all day needs CDC and per-tenant isolation; a system that must reflect a change within seconds needs real-time upserts. Interview angle. When asked to design enterprise search or a support assistant, your first question back should be “how fresh must an answer be after the source changes?” — it instantly sizes the ingestion architecture and shows you have built one, not read about one.
A scaling failure mode that only appears in production: the re-embedding clock. Change the embedding model — or the chunker, or the contextualization prompt — and every stored vector now represents different text and is no longer comparable to the rest, so recall sags a few points and stays there with no error in the logs. At ~1 TB corpora a full re-embed runs on the order of $12k/month, which is why mature teams re-index behind a dual index, validate the new one on a held-out labelled set, and cut over atomically. Mixing two embedding generations in one index is an invisible recall leak that you find from a customer complaint weeks later.
Build vs buy: the named-vendor tradeoffs
Customers rarely build all seven layers from scratch, and a solutions engineer is expected to know the realistic combinations and their tradeoffs. A typical enterprise architecture combines agents-towards-production patterns for the build, Evidently for offline + online eval, Honeycomb for tracing, and either Runlayer or a Speakeasy/Tyk-built gateway for MCP governance. The consistent tension across every layer is the same: open protocols vs vendorized platforms solve the same problem from opposite directions. Open-protocol-first stacks win on portability and vendor optionality; vendor-covered stacks win on time-to-compliance — auditability and IdP integration on day one in a regulated buyer.
code
1LAYER -> a realistic build-vs-buy menu (and what you trade)23 Layer Buy (fast compliance) Build/compose (portable)4 -------------- ----------------------- --------------------------5 Eval Evidently platform agents-towards-production +6 your own harness7 Observability Honeycomb (Agent Timeline) OpenTelemetry + Phoenix/Langfuse8 MCP governance Runlayer (IdP, audit, Speakeasy/Tyk gateway, you add9 18k+ servers) auth + audit + policy10 RAG ingestion managed vector DB Confluent CDC + Flink (heavier)1112 Rule: buy the layer where time-to-compliance is the binding constraint;13 compose the layer where portability / optionality matters more.
Interview angle. “Would you build or buy the eval / observability / connector layer?” The weak answer picks one; the strong answer names the axis — Runlayer and Evidently are the most opinionated choices and win on time-to-compliance, while a Speakeasy/Tyk gateway and an OpenTelemetry-based observability stack are the most flexible and win on portability. Weight the evidence by stage, too: eval and observability claims are engineering-grade (Evidently’s >80% judge-human agreement, Intercom’s 2-second win), while a vendor’s scale claim (Runlayer’s 18,000+ servers) is marketing-grade — useful for a buyer conversation, not a benchmark.
In an architecture round you’re given: “Design a customer-facing medical-billing AI that turns patient notes into insurer claims.” What’s the strongest opening move?
APick the best foundation model and fine-tune it on a corpus of past claimsBLay out the walking skeleton (gateway → orchestrator → retrieval → tools → evals/guardrails → observability), then state the SLOs and turn layers on one at a timeCEstimate the token cost per claim and optimize the prompt length first
A regulated customer asks how you guarantee tenant A can never retrieve tenant B’s documents in a shared-model deployment. Best answer?
AInstruct the model in the system prompt to only use the current tenant’s dataBGive each tenant a dedicated fine-tuned model so the data never mixesCEnforce RBAC at the retrieval layer: filter to tenant_id and the caller’s ACL before similarity runs, so foreign documents are never candidates
A support assistant over a knowledge base that employees edit all day keeps citing policies that changed hours ago. Which architecture change actually fixes it?
AIncrease the nightly batch re-index to run twice a dayBMove to CDC + streaming ingestion (e.g. Flink) with real-time upserts into the vector store, sized to the freshness SLACSwap to a larger-context model so it can read more documents at once
Your AI feature’s monthly bill 10×’d overnight with flat user traffic. Most likely cause and the right safeguard?
AThe provider raised prices; switch to a cheaper modelBAn agent loop or retry storm with no token cap; bound it with gateway rate limits, max-tokens, a per-session cost budget, and back-pressure on the queueCUsers discovered the feature; this is normal growth
A customer wants the lowest-cost design for a high-volume support assistant where most questions repeat. Which lever should you reach for first?
AFine-tune a small model so each call is cheaperBBuy more GPU capacity to lower per-token costCAdd semantic caching of similar prior queries (plus model routing and a token budget), cutting average $/request by an order of magnitude on repetitive traffic
The solutions-engineer architecture round whiteboards an enterprise AI system from a one-line prompt (“design our Claude chat service,” “design a system that processes 10k uploads/month and handles provider downtime”) and scores you on seven axes: problem decomposition, LLM-aware architecture, data/context strategy, reliability, eval/monitoring, security/privacy, cost/latency. Lead with the skeleton, name the SLOs, and attach a failure mode to every layer. Strong candidates ground outputs with retrieval/citations, design for rate limits/retries/fallbacks, and ship a monitoring plan; weak ones treat the LLM as source of truth and jump to fine-tuning.
01“Walk the architecture of an enterprise AI solution.” → seven layers in order (gateway → orchestrator → retrieval → model → tools → evals/guardrails → observability); name what each buys.
02“Where does the security boundary live?” → retrieval RBAC filter + IdP-checked tool gateway, never the prompt; filter-then-retrieve with the caller’s ACL.
03“What decides model choice?” → the SLOs (TTFT, p95, accuracy, $/req) and prompt size; route short/cheap and long/hard to different models — the SLO picks the model.
04“Multi-tenancy options?” → shared+RLS, shared-model+isolated-retrieval (default), per-tenant replicas (contract-forced); escalate only when isolation is contractual.
05“How do you keep the knowledge base fresh?” → CDC + streaming upserts sized to the freshness SLA; nightly batch is the “stale brain” unless edits are rare.
06“Your AI bill 10×’d at flat traffic — why?” → runaway agent loop / retry storm; bound with rate limits, max-tokens, per-session budget, back-pressure (denial-of-wallet).
07“What breaks when you change the embedding model?” → re-embedding clock — vectors no longer comparable, silent recall sag; dual-index, validate, cut over atomically.
08“Build vs buy for the gateway / eval / observability?” → buy the governance layer for time-to-compliance, compose open primitives for portability; name the tradeoff explicitly.
To go deeper, expect the follow-up that ends most loops: “what breaks, and how do you detect it?” for each layer you drew — the panel wants a failure mode and a telemetry signal per box, not a happy path. Other probes: “your retrieval recall looks fine but answers are wrong” (generation/grounding problem, not retrieval — instrument the boundary); “the customer is in healthcare and air-gapped” (VPC/on-prem deployment, data residency, no calls to a public API — Lesson 4’s connector governance); and “how would you stage this rollout?” (walking skeleton behind a flag, eval gate, canary, then widen). In every case, name the layer and the metric before the fix.
Could you whiteboard the seven layers from a one-line prompt, name the SLOs that drive model choice, and attach a failure mode to each layer?
New to itGetting thereConfident
Takeaways
The supporting stack, not the model, decides an enterprise AI build — draw the seven layers in order.
Open the architecture round with a walking skeleton; never open with “fine-tune.”
Security lives at the retrieval RBAC filter and the tool gateway, not in the prompt.
Design backwards from the SLOs (TTFT, p95, accuracy, $/req); the SLO picks the model.
Streaming CDC ingestion beats nightly ETL; the freshness SLA sizes the pipeline.
Bound the spend (rate limits, max-tokens, back-pressure) — LLM endpoints are denial-of-wallet targets.
Next: how you prove the thing is any good — evals for customer-facing AI, and the grader bug that turned 42% into 95%.