Isolation means every egress from the agent runtime terminates in something the customer can route and audit — their keys, their bucket, their DNS, their model. PrivateLink and private endpoints, the air-gap spectrum, what each posture actually buys you under GDPR/HIPAA/ITAR, and the three-year TCO reality.
The line that ends the CISO objection
The single most-explaining sentence in this space is AWS’s: an interface VPC endpoint lets you “access Amazon Bedrock as if it were in your VPC, without the use of an internet gateway, NAT device, VPN connection, or Direct Connect ... Instances in your VPC don’t need public IP addresses.” That one fact resolves the most common regulated-CISO objection — “but the model still touches the public internet” — because with PrivateLink it does not. The senior framing of the whole lesson: isolation is achieved when every egress from the agent runtime terminates in something the customer can route and audit — private DNS, customer-managed KMS, customer-controlled vector store, model endpoint in the customer account.
This is why “deploy the agent inside the customer’s VPC” is an under-specification that gets engineers in trouble. Deploy it inside the customer’s network with their keys, their bucket, their DNS, and their model. The diagnostic discipline: draw the sequence diagram of fetch → parse → embed → retrieve → generate → log for every customer PoC, and if any arrow crosses a public endpoint without a customer-controlled hop, draw a red box around it. Each red box is a residency, audit, or egress objection waiting to surface in security review — better you find it on the whiteboard than the customer’s pen-tester finds it in week six.
The private-network pattern is portable across hyperscalers, which matters because you will land on whichever cloud the customer already runs. On AWS, a VPC interface endpoint (PrivateLink) to Bedrock plus customer-managed KMS keeps both traffic and keys inside the account. On Azure, Azure OpenAI Service supports network isolation via private endpoints (reached through Azure Bastion, no public IP). On GCP, Vertex AI / Gemini Enterprise supports VPC Network Peering. The control objectives are identical; only the primitive names differ — which is exactly why your control plane should not hard-code one vendor’s networking.
code
1PRIVATE NETWORK primitive by cloud -- same objective, different name23 Cloud Model access primitive Keys Reached via4 ----- ------------------------------- ------------------ ------------------5 AWS PrivateLink interface endpoint customer KMS (CMK) private DNS, no IGW6 to Bedrock7 Azure Private endpoint to Azure customer-managed Azure Bastion,8 OpenAI Service key (CMK) no public IP9 GCP VPC Network Peering to Vertex CMEK peered VPC10 AI / Gemini Enterprise1112 Objective everywhere: no public IP on the runtime, traffic stays on the cloud13 backbone, customer holds the keys, DNS resolves privately. Don't hard-code one.
The deployment spectrum: SaaS → VPC → air-gapped
Real enterprise deployment posture is a spectrum of how much the customer owns, from SSO-bound SaaS to a fully air-gapped network with no internet at all. Squirro’s framing captures the practical distinction: air-gapped is best when the data must never leave the physical perimeter; VPC is the compromise when you trust the hyperscaler network. Each step up the spectrum buys compliance ceiling and control at the cost of velocity and SRE burden you take on.
code
1DEPLOYMENT spectrum -- reach of customer control vs cost of owning it23 Posture Network Typical customer4 ------------------------------ ------------------- ------------------------5 Public SaaS, SSO-bound managed, public low-sensitivity / SMB6 SaaS + PrivateLink + gateway private backbone, most regulated commercial7 in VPC controlled egress (legal, health, finance)8 BYO-VPC, managed model via customer account, customer owns data plane9 private endpoint + CMK private (sovereignty needs)10 Fully self-hosted open-weights no managed model; highest control; you own11 (Llama/Mistral/Qwen on EKS) controlled egress evals, SRE, scale-out12 Air-gapped (no internet) no egress; updates defense / gov / pharma /13 via approved media classified1415 In ALL of these the LLM is the same KIND of artifact -- open-weights on customer16 GPUs or a permitted commercial model deployed inside the perimeter under licence.
A crucial point juniors miss: more open-weights does not mean more secure. Air-gapping a Llama bundle does nothing to stop a permission leak in your vector DB — the ReBAC checks from Lesson 2 are still required. Isolation defends against egress and residency threats; it is orthogonal to authorization threats. A perfectly air-gapped system with no document-level ACLs leaks every privileged memo to every employee, just without sending it over the internet. Interview angle. “If we air-gap the whole thing, are we secure?” The strong answer separates the threat models: air-gap closes egress, but you still need retrieval-time authorization, redaction, and audit inside the perimeter.
When air-gap is required, not optional
Air-gapping is operationally expensive — long update funnels, signed model and patch bundles transferred by approved media, no hosted telemetry — so it is justified only by specific triggers, not by general nervousness. Four scenarios force it: (1) privileged or confidential legal content where a third-party SaaS chain-of-custody creates waiver risk; (2) classified or export-controlled material under ITAR or equivalent (no non-US-person access); (3) strict residency regimes where supervisory authorities (German BfDI, French CNIL, Swiss FINMA, Nordic health authorities) require both data and control plane within jurisdiction; (4) internal policy predating the AI question that prohibits any third-party data egress.
Inside the air-gapped perimeter, every hosted service is replaced with a local equivalent: Pinecone/hosted-Weaviate become self-hosted Weaviate, Qdrant, or pgvector; Cohere-hosted rerankers become local BGE or Jina cross-encoders; hosted self-eval APIs become a same-pool (or deliberately different) local model. The hardware footprint is now tractable — NVIDIA published a community blueprint for an RTX PRO 6000 Blackwell-based on-prem RAG that runs fully offline, and a single-node Llama-3.1-class retriever-plus-generator fits a 2-4 GPU rack appliance in a standard data-center row. The point for an FDE: air-gapped is demanding but feasible in 2026, and the architecture is the same eight layers — just with every external dependency pulled inside.
code
1REPLACING hosted services for an air-gapped / sovereign deployment23 Hosted (SaaS) -> In-perimeter equivalent4 ---------------------------- --------------------------------5 Pinecone / hosted Weaviate -> self-hosted Weaviate / Qdrant / pgvector6 Cohere / Voyage reranker (API) -> local BGE / Jina cross-encoder on GPU7 hosted LLM (Bedrock/OpenAI) -> Llama / Mistral / Qwen on customer GPUs8 hosted eval / judge API -> same-pool or deliberately-different local model9 hosted observability cloud -> self-hosted Langfuse/Phoenix; export aggregates10 managed KMS (external) -> customer HSM / in-account KMS1112 The eight layers are unchanged -- only the dependency boundary moves inward.13 Updates arrive as signed model + patch bundles over an approved media funnel.
What each posture buys under real regulation
Forward-deployed engineers must map posture to the specific regime in scope, because the customer’s lawyers will. The recurring, non-obvious distinction: regulators increasingly separate “data resident in jurisdiction” from “control plane in jurisdiction” — a US-hosted SaaS control plane that physically stores data in the EU does not satisfy German or French supervisory authorities. This is why the policy/gateway plane so often has to be self-operated inside the customer’s environment.
code
1WHAT on-prem / private buys you, by regime23 Regime Primary RAG requirement What on-prem/private buys4 ------------- ---------------------------------- --------------------------5 GDPR (EU) lawful basis; subject rights; data + control plane in6 transfer restrictions jurisdiction; easier erasure7 HIPAA PHI protection; BAA chain no 3rd-party BAA for the AI8 vendor when in-house9 ITAR no non-US-person access to US-person-only env; no10 controlled data vendor-support exposure11 EU AI Act risk classification; traceability; full audit log; no external12 human oversight (from Aug 2 2026) deps in the decision chain13 FCA / UK FS operational resilience; 3rd-party removes a category of14 risk third-party risk entirely1516 Key nuance: "data in jurisdiction" != "control plane in jurisdiction" -- a17 US-operated SaaS control plane fails strict EU supervisory tests.
PII and the audit trail interact with deployment here too: EU AI Act Article 12 traceability is satisfied cleanly by an on-prem deployment with an immutable, append-only log, whereas a US SaaS control plane that cannot guarantee in-jurisdiction operation fails the test regardless of where the bytes sit. So the deployment posture and the audit design (Lesson 2) are not independent choices — a regulated customer’s residency requirement often forces a self-operated control plane, which in turn shapes where the audit chokepoint lives.
Draw the sequence diagram — fetch, parse, embed, retrieve, generate, log — for every PoC, and red-box any arrow that crosses a public endpoint without a customer-controlled hop. The red boxes are your security-review findings, surfaced on a whiteboard in week one instead of by the customer’s pen-tester in week six.
The TCO reality (and the migration room)
The on-prem-vs-SaaS debate is usually framed as a sticker-price decision, but the honest three-year model converges closer than the RFP suggests. On-prem costs are visible: GPU compute ($30-100k+ per LLM-capable server, refreshed every 3-5 years), the vector store, index + audit storage, power/cooling, and an FTE-fraction for ops and security. SaaS hides the GPU cost but adds per-token pricing, ongoing vendor security review, DPA legal review, audit-cooperation time, and a contractual risk premium. At 50+ users and meaningful content volume, three-year TCO converges — so the deciding factor is compliance posture, not price.
The forward-deployed move is to architect so the customer can slide along this spectrum without re-platforming. Klarna is the cautionary tale: a hosted-OpenAI assistant that handled 2.3M conversations in month one, but by mid-2025 the cost-and-quality surface of unbounded hosted inference drove a redesign toward private models. Had the control plane been model-agnostic and the data plane already in their account, that migration would have been a config change rather than a rebuild. Interview angle. “How would you let this customer move from hosted Bedrock to self-hosted Llama later?” → keep the model a swappable endpoint behind the gateway, keep keys/buckets/index in the customer account from day one, so the open-weights migration is incremental, not a forklift.
A regulated CISO says “your design still sends our queries over the public internet to the model.” Using AWS, what is the precise architectural answer?
AA PrivateLink interface endpoint to Bedrock keeps traffic on the AWS backbone with no internet gateway or public IP, and customer KMS holds the keysBEncrypt the traffic with TLS so it is safe in transit over the internetCMove to a fully air-gapped deployment immediately
A customer insists on a fully air-gapped deployment “to be maximally secure.” They have no ITAR data, no strict-residency mandate, and need to launch in a month. What is the senior response?
AAgree and start the air-gap build immediatelyBProbe for a real trigger; absent one, recommend SaaS-managed model via PrivateLink with the gateway in their VPC, and note air-gap doesn’t fix authorization anywayCTell them air-gap is unnecessary because the cloud is always secure
You air-gap a customer’s entire RAG stack with an open-weights model. Which statement is correct?
AThe system is now fully secure because no data can leave the networkBOpen-weights models are inherently more secure than managed modelsCAir-gap closes egress, but you still need retrieval-time authorization, PII redaction, and audit inside the perimeter
An EU customer’s counsel says a US-hosted SaaS control plane storing data in an EU region is insufficient. Why might they be right?
AEU regulators only care about encryption strength, which US providers may lackBStrict EU supervisory authorities distinguish “data in jurisdiction” from “control plane in jurisdiction”; a US-operated control plane can fail even with EU data residencyCThere is no real distinction; data residency always satisfies GDPR
A customer wants hosted Bedrock now but may need to move to self-hosted Llama within a year for cost reasons. What do you put in place on day one?
AKeep keys, buckets, and the vector index in the customer’s account, with the model as a swappable endpoint behind the gatewayBUse Bedrock-specific features heavily to maximize current performanceCDefer the question — model choice never changes after launch
Deployment questions test whether you can reason about the customer’s network, keys, and regulator — not just recite “put it in a VPC.” The strong pattern: name the posture from the customer’s constraints, draw the egress diagram and red-box any public hop, separate the egress threat model from the authorization threat model, and map posture to the specific regime. Quote the PrivateLink fact verbatim — it is the single most useful line for clearing a security review.
01“The model still touches the internet — true?” → no, with PrivateLink/private endpoint/VPC peering: no public IP, traffic on the cloud backbone, customer holds keys.
02“What does ‘deploy in our VPC’ actually require?” → their keys, their bucket, their DNS, their model; every egress terminates in a customer-controlled, auditable hop.
03“When is air-gap required?” → privileged/legal waiver risk, ITAR/export control, strict residency (both data and control plane), or a no-egress policy — not general nervousness.
04“Is air-gap automatically most secure?” → it maximizes the egress/residency ceiling but doesn’t fix authorization, redaction, or stale index; orthogonal threat models.
05“What does on-prem buy under HIPAA / ITAR / EU AI Act?” → no 3rd-party BAA; US-person-only env; clean Article 12 traceability with an immutable log.
06“Data residency vs control-plane residency?” → regulators distinguish them; a US-operated control plane can fail even with in-region data, forcing a self-operated gateway.
07“On-prem vs SaaS cost?” → 3-year TCO converges at 50+ users; the deciding factor is compliance posture, not sticker price.
08“How do they migrate to self-hosted later?” → model as swappable endpoint, data plane in the customer account from day one — incremental, not a re-platform.
Going deeper. Expect the field-deployment follow-ups. “Their landing zone is a 15-year-old on-prem ERP with no API — how do you ingest into the VPC RAG?” (change-data-capture or scheduled batch extracts into the in-VPC pipeline; screen-scraping only as a last resort; never a big-bang migration). “How do you patch an air-gapped model?” (signed model and patch bundles transferred via approved media on a defined funnel; you own the update cadence). “The customer wants hosted telemetry but is air-gapped — reconcile that.” (self-host the observability collector inside the perimeter and export only aggregated, non-sensitive metrics). The hidden rubric is whether you reconcile constraints and hand the customer a concrete plan rather than a purist absolute.
Could you draw the egress diagram, pick a posture from triggers and TCO, and map it to HIPAA/ITAR/EU AI Act?
New to itGetting thereConfident
Takeaways
Isolation = every egress terminates in a customer-controlled, auditable hop: their keys, bucket, DNS, model.
PrivateLink / private endpoint / VPC peering put the managed model on the private backbone with no public IP — the answer to “it still touches the internet.”
The posture spectrum (SaaS → VPC → air-gapped) trades compliance ceiling for velocity and SRE burden; pick from real triggers.
Air-gap closes egress only — authorization, redaction, and audit are still required inside the perimeter.
Regulators distinguish data-in-jurisdiction from control-plane-in-jurisdiction; a US-operated control plane can fail EU tests.
3-year TCO converges at 50+ users; keep the model swappable and the data plane in-account so migration isn’t a re-platform.
Next: evals & observability in the field — measuring and monitoring the deployment on-site, where you can’t see the customer’s data.