Lesson 3 of 6 · 46 min

Private/VPC deployment patterns

Isolation means every egress from the agent runtime terminates in something the customer can route and audit — their keys, their bucket, their DNS, their model. PrivateLink and private endpoints, the air-gap spectrum, what each posture actually buys you under GDPR/HIPAA/ITAR, and the three-year TCO reality.

The line that ends the CISO objection

The single most-explaining sentence in this space is AWS’s: an interface VPC endpoint lets you “access Amazon Bedrock as if it were in your VPC, without the use of an internet gateway, NAT device, VPN connection, or Direct Connect ... Instances in your VPC don’t need public IP addresses.” That one fact resolves the most common regulated-CISO objection — “but the model still touches the public internet” — because with PrivateLink it does not. The senior framing of the whole lesson: isolation is achieved when every egress from the agent runtime terminates in something the customer can route and audit — private DNS, customer-managed KMS, customer-controlled vector store, model endpoint in the customer account.
This is why “deploy the agent inside the customer’s VPC” is an under-specification that gets engineers in trouble. Deploy it inside the customer’s network with their keys, their bucket, their DNS, and their model. The diagnostic discipline: draw the sequence diagram of fetch → parse → embed → retrieve → generate → log for every customer PoC, and if any arrow crosses a public endpoint without a customer-controlled hop, draw a red box around it. Each red box is a residency, audit, or egress objection waiting to surface in security review — better you find it on the whiteboard than the customer’s pen-tester finds it in week six.
Deploy an AI assistant on-premise, air-gapped, VPC, or SaaSTabnine

The same shape on every cloud

The private-network pattern is portable across hyperscalers, which matters because you will land on whichever cloud the customer already runs. On AWS, a VPC interface endpoint (PrivateLink) to Bedrock plus customer-managed KMS keeps both traffic and keys inside the account. On Azure, Azure OpenAI Service supports network isolation via private endpoints (reached through Azure Bastion, no public IP). On GCP, Vertex AI / Gemini Enterprise supports VPC Network Peering. The control objectives are identical; only the primitive names differ — which is exactly why your control plane should not hard-code one vendor’s networking.
code
1PRIVATE NETWORK primitive by cloud -- same objective, different name23  Cloud   Model access primitive            Keys                 Reached via4  -----   -------------------------------   ------------------   ------------------5  AWS     PrivateLink interface endpoint    customer KMS (CMK)   private DNS, no IGW6          to Bedrock7  Azure   Private endpoint to Azure         customer-managed     Azure Bastion,8          OpenAI Service                    key (CMK)            no public IP9  GCP     VPC Network Peering to Vertex     CMEK                 peered VPC10          AI / Gemini Enterprise1112  Objective everywhere: no public IP on the runtime, traffic stays on the cloud13  backbone, customer holds the keys, DNS resolves privately. Don't hard-code one.

The deployment spectrum: SaaS → VPC → air-gapped

Real enterprise deployment posture is a spectrum of how much the customer owns, from SSO-bound SaaS to a fully air-gapped network with no internet at all. Squirro’s framing captures the practical distinction: air-gapped is best when the data must never leave the physical perimeter; VPC is the compromise when you trust the hyperscaler network. Each step up the spectrum buys compliance ceiling and control at the cost of velocity and SRE burden you take on.
code
1DEPLOYMENT spectrum -- reach of customer control vs cost of owning it23  Posture                          Network               Typical customer4  ------------------------------   -------------------   ------------------------5  Public SaaS, SSO-bound           managed, public       low-sensitivity / SMB6  SaaS + PrivateLink + gateway     private backbone,     most regulated commercial7  in VPC                           controlled egress     (legal, health, finance)8  BYO-VPC, managed model via       customer account,     customer owns data plane9  private endpoint + CMK           private               (sovereignty needs)10  Fully self-hosted open-weights   no managed model;     highest control; you own11  (Llama/Mistral/Qwen on EKS)      controlled egress     evals, SRE, scale-out12  Air-gapped (no internet)         no egress; updates    defense / gov / pharma /13                                   via approved media    classified1415  In ALL of these the LLM is the same KIND of artifact -- open-weights on customer16  GPUs or a permitted commercial model deployed inside the perimeter under licence.
A crucial point juniors miss: more open-weights does not mean more secure. Air-gapping a Llama bundle does nothing to stop a permission leak in your vector DB — the ReBAC checks from Lesson 2 are still required. Isolation defends against egress and residency threats; it is orthogonal to authorization threats. A perfectly air-gapped system with no document-level ACLs leaks every privileged memo to every employee, just without sending it over the internet. Interview angle. “If we air-gap the whole thing, are we secure?” The strong answer separates the threat models: air-gap closes egress, but you still need retrieval-time authorization, redaction, and audit inside the perimeter.

When air-gap is required, not optional

Air-gapping is operationally expensive — long update funnels, signed model and patch bundles transferred by approved media, no hosted telemetry — so it is justified only by specific triggers, not by general nervousness. Four scenarios force it: (1) privileged or confidential legal content where a third-party SaaS chain-of-custody creates waiver risk; (2) classified or export-controlled material under ITAR or equivalent (no non-US-person access); (3) strict residency regimes where supervisory authorities (German BfDI, French CNIL, Swiss FINMA, Nordic health authorities) require both data and control plane within jurisdiction; (4) internal policy predating the AI question that prohibits any third-party data egress.
Inside the air-gapped perimeter, every hosted service is replaced with a local equivalent: Pinecone/hosted-Weaviate become self-hosted Weaviate, Qdrant, or pgvector; Cohere-hosted rerankers become local BGE or Jina cross-encoders; hosted self-eval APIs become a same-pool (or deliberately different) local model. The hardware footprint is now tractable — NVIDIA published a community blueprint for an RTX PRO 6000 Blackwell-based on-prem RAG that runs fully offline, and a single-node Llama-3.1-class retriever-plus-generator fits a 2-4 GPU rack appliance in a standard data-center row. The point for an FDE: air-gapped is demanding but feasible in 2026, and the architecture is the same eight layers — just with every external dependency pulled inside.
code
1REPLACING hosted services for an air-gapped / sovereign deployment23  Hosted (SaaS)                  ->  In-perimeter equivalent4  ----------------------------       --------------------------------5  Pinecone / hosted Weaviate     ->  self-hosted Weaviate / Qdrant / pgvector6  Cohere / Voyage reranker (API) ->  local BGE / Jina cross-encoder on GPU7  hosted LLM (Bedrock/OpenAI)    ->  Llama / Mistral / Qwen on customer GPUs8  hosted eval / judge API        ->  same-pool or deliberately-different local model9  hosted observability cloud     ->  self-hosted Langfuse/Phoenix; export aggregates10  managed KMS (external)         ->  customer HSM / in-account KMS1112  The eight layers are unchanged -- only the dependency boundary moves inward.13  Updates arrive as signed model + patch bundles over an approved media funnel.

What each posture buys under real regulation

Forward-deployed engineers must map posture to the specific regime in scope, because the customer’s lawyers will. The recurring, non-obvious distinction: regulators increasingly separate “data resident in jurisdiction” from “control plane in jurisdiction” — a US-hosted SaaS control plane that physically stores data in the EU does not satisfy German or French supervisory authorities. This is why the policy/gateway plane so often has to be self-operated inside the customer’s environment.
code
1WHAT on-prem / private buys you, by regime23  Regime          Primary RAG requirement              What on-prem/private buys4  -------------   ----------------------------------   --------------------------5  GDPR (EU)       lawful basis; subject rights;        data + control plane in6                  transfer restrictions                jurisdiction; easier erasure7  HIPAA           PHI protection; BAA chain            no 3rd-party BAA for the AI8                                                       vendor when in-house9  ITAR            no non-US-person access to           US-person-only env; no10                  controlled data                      vendor-support exposure11  EU AI Act       risk classification; traceability;   full audit log; no external12                  human oversight (from Aug 2 2026)    deps in the decision chain13  FCA / UK FS     operational resilience; 3rd-party    removes a category of14                  risk                                 third-party risk entirely1516  Key nuance: "data in jurisdiction" != "control plane in jurisdiction" -- a17  US-operated SaaS control plane fails strict EU supervisory tests.
PII and the audit trail interact with deployment here too: EU AI Act Article 12 traceability is satisfied cleanly by an on-prem deployment with an immutable, append-only log, whereas a US SaaS control plane that cannot guarantee in-jurisdiction operation fails the test regardless of where the bytes sit. So the deployment posture and the audit design (Lesson 2) are not independent choices — a regulated customer’s residency requirement often forces a self-operated control plane, which in turn shapes where the audit chokepoint lives.
Draw the sequence diagram — fetch, parse, embed, retrieve, generate, log — for every PoC, and red-box any arrow that crosses a public endpoint without a customer-controlled hop. The red boxes are your security-review findings, surfaced on a whiteboard in week one instead of by the customer’s pen-tester in week six.

The TCO reality (and the migration room)

The on-prem-vs-SaaS debate is usually framed as a sticker-price decision, but the honest three-year model converges closer than the RFP suggests. On-prem costs are visible: GPU compute ($30-100k+ per LLM-capable server, refreshed every 3-5 years), the vector store, index + audit storage, power/cooling, and an FTE-fraction for ops and security. SaaS hides the GPU cost but adds per-token pricing, ongoing vendor security review, DPA legal review, audit-cooperation time, and a contractual risk premium. At 50+ users and meaningful content volume, three-year TCO converges — so the deciding factor is compliance posture, not price.
The forward-deployed move is to architect so the customer can slide along this spectrum without re-platforming. Klarna is the cautionary tale: a hosted-OpenAI assistant that handled 2.3M conversations in month one, but by mid-2025 the cost-and-quality surface of unbounded hosted inference drove a redesign toward private models. Had the control plane been model-agnostic and the data plane already in their account, that migration would have been a config change rather than a rebuild. Interview angle. “How would you let this customer move from hosted Bedrock to self-hosted Llama later?” → keep the model a swappable endpoint behind the gateway, keep keys/buckets/index in the customer account from day one, so the open-weights migration is incremental, not a forklift.
docsUse interface VPC endpoints (PrivateLink) for Amazon BedrockAWSarticleOn-Premise RAG: Deployment Guide for Regulated SectorsEdtekarticleFrom Air-Gapped AI to VPC DeploymentsSquirro

Checkpoint

A regulated CISO says “your design still sends our queries over the public internet to the model.” Using AWS, what is the precise architectural answer?

AA PrivateLink interface endpoint to Bedrock keeps traffic on the AWS backbone with no internet gateway or public IP, and customer KMS holds the keysBEncrypt the traffic with TLS so it is safe in transit over the internetCMove to a fully air-gapped deployment immediately
Sign up free to answer and see why

Checkpoint

A customer insists on a fully air-gapped deployment “to be maximally secure.” They have no ITAR data, no strict-residency mandate, and need to launch in a month. What is the senior response?

AAgree and start the air-gap build immediatelyBProbe for a real trigger; absent one, recommend SaaS-managed model via PrivateLink with the gateway in their VPC, and note air-gap doesn’t fix authorization anywayCTell them air-gap is unnecessary because the cloud is always secure
Sign up free to answer and see why

Checkpoint

You air-gap a customer’s entire RAG stack with an open-weights model. Which statement is correct?

AThe system is now fully secure because no data can leave the networkBOpen-weights models are inherently more secure than managed modelsCAir-gap closes egress, but you still need retrieval-time authorization, PII redaction, and audit inside the perimeter
Sign up free to answer and see why

Checkpoint

An EU customer’s counsel says a US-hosted SaaS control plane storing data in an EU region is insufficient. Why might they be right?

AEU regulators only care about encryption strength, which US providers may lackBStrict EU supervisory authorities distinguish “data in jurisdiction” from “control plane in jurisdiction”; a US-operated control plane can fail even with EU data residencyCThere is no real distinction; data residency always satisfies GDPR
Sign up free to answer and see why

Checkpoint

A customer wants hosted Bedrock now but may need to move to self-hosted Llama within a year for cost reasons. What do you put in place on day one?

AKeep keys, buckets, and the vector index in the customer’s account, with the model as a swappable endpoint behind the gatewayBUse Bedrock-specific features heavily to maximize current performanceCDefer the question — model choice never changes after launch
Sign up free to answer and see why

Interview prep

Deployment questions test whether you can reason about the customer’s network, keys, and regulator — not just recite “put it in a VPC.” The strong pattern: name the posture from the customer’s constraints, draw the egress diagram and red-box any public hop, separate the egress threat model from the authorization threat model, and map posture to the specific regime. Quote the PrivateLink fact verbatim — it is the single most useful line for clearing a security review.
  1. 01“The model still touches the internet — true?” → no, with PrivateLink/private endpoint/VPC peering: no public IP, traffic on the cloud backbone, customer holds keys.
  2. 02“What does ‘deploy in our VPC’ actually require?” → their keys, their bucket, their DNS, their model; every egress terminates in a customer-controlled, auditable hop.
  3. 03“When is air-gap required?” → privileged/legal waiver risk, ITAR/export control, strict residency (both data and control plane), or a no-egress policy — not general nervousness.
  4. 04“Is air-gap automatically most secure?” → it maximizes the egress/residency ceiling but doesn’t fix authorization, redaction, or stale index; orthogonal threat models.
  5. 05“What does on-prem buy under HIPAA / ITAR / EU AI Act?” → no 3rd-party BAA; US-person-only env; clean Article 12 traceability with an immutable log.
  6. 06“Data residency vs control-plane residency?” → regulators distinguish them; a US-operated control plane can fail even with in-region data, forcing a self-operated gateway.
  7. 07“On-prem vs SaaS cost?” → 3-year TCO converges at 50+ users; the deciding factor is compliance posture, not sticker price.
  8. 08“How do they migrate to self-hosted later?” → model as swappable endpoint, data plane in the customer account from day one — incremental, not a re-platform.
Going deeper. Expect the field-deployment follow-ups. “Their landing zone is a 15-year-old on-prem ERP with no API — how do you ingest into the VPC RAG?” (change-data-capture or scheduled batch extracts into the in-VPC pipeline; screen-scraping only as a last resort; never a big-bang migration). “How do you patch an air-gapped model?” (signed model and patch bundles transferred via approved media on a defined funnel; you own the update cadence). “The customer wants hosted telemetry but is air-gapped — reconcile that.” (self-host the observability collector inside the perimeter and export only aggregated, non-sensitive metrics). The hidden rubric is whether you reconcile constraints and hand the customer a concrete plan rather than a purist absolute.

Could you draw the egress diagram, pick a posture from triggers and TCO, and map it to HIPAA/ITAR/EU AI Act?

New to itGetting thereConfident

Takeaways

  • Isolation = every egress terminates in a customer-controlled, auditable hop: their keys, bucket, DNS, model.
  • PrivateLink / private endpoint / VPC peering put the managed model on the private backbone with no public IP — the answer to “it still touches the internet.”
  • The posture spectrum (SaaS → VPC → air-gapped) trades compliance ceiling for velocity and SRE burden; pick from real triggers.
  • Air-gap closes egress only — authorization, redaction, and audit are still required inside the perimeter.
  • Regulators distinguish data-in-jurisdiction from control-plane-in-jurisdiction; a US-operated control plane can fail EU tests.
  • 3-year TCO converges at 50+ users; keep the model swappable and the data plane in-account so migration isn’t a re-platform.

Next: evals & observability in the field — measuring and monitoring the deployment on-site, where you can’t see the customer’s data.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.