Lesson 5 of 6 · 48 min

Guardrails & PII handling

Every prior layer can fail, so guardrails are the last line — and PII has three redaction opportunities, each mandatory. Presidio + NeMo + Bedrock as defense in depth, redaction at ingest/retrieval/output, indirect prompt injection (CVE-2025-32711), embedding inversion, and the permission-leak failure mode confronted directly.

Why guardrails exist at all

Guardrails are the last line because every prior layer can fail: the ACL check can have a stale tag, the deployment can have a misconfigured egress, the retriever can surface a chunk it should not. A guardrail sits in three places — in front of the LLM (input screening), in front of the user (output screening), and around the agent (tool-call policy) — and its job is to turn a potential leak into a decision the system acts on: redact, block, route to human, or alert. The Samsung leak is the canonical reason this matters: an engineer pasted internal source code into ChatGPT, the data was retained, and the architectural remedy is the entire stack from Lessons 1-4 plus a guardrail layer in front of every model call.
The mechanism that separates a real guardrail from a decorative one: every guardrail must terminate in a decision the system can act on. PII hit → redact or block. Faithfulness below SLO → block or fall back to human review. Tool call out of RBAC scope → block and alert. The worst guardrail design is one that logs to a dashboard nobody reads; the second worst fires only for happy-path attack patterns. The rule for every customer pilot: demonstrate at least one real block with one audit-log entry by end of week one, or the guardrails are decorative, not structural.
Prompt Injection, explainedSimon Willison

PII redaction at three stages — all mandatory

The pipeline has three redaction opportunities and missing any one is a leak. Treat all three as required, not optional. Stage 1 — ingest (before embedding): structural redaction at the document-loader layer (regex + NER hybrid, e.g. Microsoft Presidio, plus tag dictionaries), so the document never enters the index until it carries a classification label. This is the “non-negotiable pre-index control” for PHI/PII. Stage 2 — retrieval (re-check): re-run sensitive-data detection on retrieved chunks before concatenating them into the prompt — a narrow, fast scan, justified because OWASP LLM08 warns “embeddings strip context and classification,” so a label can be lost between ingest and retrieval. Stage 3 — output (tokenize/redact): real-time tokenization replacing sensitive spans in both inbound context and outbound completions at the gateway, with originating values held in an internal vault.
code
1PII REDACTION -- three stages, each mandatory, each catches the prior miss23  Stage              Tooling                         Failure if skipped4  ----------------   -----------------------------   ----------------------------5  1. Ingest          Presidio, spaCy NER, regex      PII embedded INTO vectors --6     (pre-index)     per data class, classifiers     leaks on any close semantic7                                                      query, forever8  2. Retrieval       re-scan retrieved chunks         stale/missing classification9     (re-check)      (metadata filters, fast NER)     metadata still gets returned10  3. Output          gateway tokenization (vault-     sensitive content survives11     (gateway)       backed), NeMo/F5/Calypso         the prompt to the user1213  Inversion defense  differential-privacy noise at    stolen vectors reconstruct14  (non-obvious)      embedding time; encrypted        50-70% of original tokens15                     vectors at rest                  (ALGEN-style attacks)1617  Microsoft warns automated PII detection is best-effort -- layer multiple18  recognizers; never trust a single recognizer or a single stage.
The non-obvious fourth row is embedding inversion, and it is a favorite advanced interview probe. ACL 2024-2025 research shows some decoder-style architectures can reconstruct 50-70% of the original input tokens from stolen vectors, and the ALGEN attack trains a transferable black-box inversion model from as few as ~1,000 samples. So your vector store is not a safe place to put raw PII even though it is “just numbers” — a leaked index can be partially reversed back into text. The validated mitigation is differential-privacy noise at embedding generation (at a measurable cost to search recall) plus encrypting vectors at rest. Interview angle. “If our vector DB is breached, what is the worst case?” The strong answer names embedding inversion and RAG poisoning — not just “they get the embeddings, which are meaningless.”
There is a real tension between redaction coverage and retrieval quality, and a senior FDE engineers it on purpose rather than discovering it in production. Aggressive ingest-stage redaction improves safety but eats recall — redact the patient name or the contract counterparty and a query that needs exactly that entity now retrieves nothing. The corollary on the output side: intent-based tokenization protects the downstream user but cannot help during retrieval-stage decisioning, because by then the embedding already exists. The resolution is to make redaction per-data-class (mask SSNs and card numbers hard; preserve searchable entities behind a vault token that keeps the embedding meaningful) and to put the aggressive redaction only where the audit posture demands it. Witness.ai reports ~99.3% true-positive efficacy on session-scoped tokenization at the gateway — the point being that a vault-backed token preserves enough semantics for retrieval while exposing only the token to the LLM.

The guardrail stack: three layers, defense in depth

Guardrails must run on both sides of the model and never rely on a single tool. The reference stack layers three: Layer 1 — PII detection/redaction (Presidio scans inbound prompt and outbound context+response with separate policies per side). Layer 2 — programmable rails between code and LLM (NVIDIA NeMo Guardrails defines flows for what the bot must and must not respond to; an evaluator like faithfulness can be wired in as a rail that blocks below-SLO answers). Layer 3 — vendor content filters at egress (AWS Bedrock Guardrails — denied topics, content filters, sensitive-information filters, word policies; Azure OpenAI and ChatGPT Enterprise ship equivalents; Fiddler reports sub-100ms guardrail responses). The canonical regulated-environment reference architecture is F5 AI Guardrails (powered by Calypso AI) on Red Hat OpenShift AI, which screens prompts for injection/jailbreak and responses for sensitive-data emission and hallucination through a single versioned-policy gateway.
  1. 01Input guardrails — schema checks + prompt-injection pattern detection before retrieval runs.
  2. 02Retrieval guardrails — metadata filters + similarity thresholds so only allowed, on-topic data is accessed.
  3. 03Output guardrails — groundedness check + format/schema enforcement + policy rules restricting agent actions.
  4. 04Each layer closes a specific bucket: hallucination, unsafe/harmful content, PII/data leakage, prompt injection, workflow/tool abuse.
A senior tension to surface, not hide: stronger guardrails can degrade service quality. NeMo and Bedrock Guardrails force refusals when they detect prompt-injection patterns — including false positives on legitimate internal prompts that mention “ignore the above instructions” (a phrase engineers literally use when testing). Always ship an explicit allow-list for known-good internal patterns. The same tension appears in PII: aggressive ingest redaction improves safety but eats retrieval recall (you redacted the entity the query needed). Engineer the trade-off where the audit posture demands it, and make the redaction policy per-data-class, not a blanket mask. Interview angle. “Is contextual grounding enough to stop injection?” Quote the AWS caveat verbatim: “contextual grounding helps mitigate prompt injection by limiting responses to the retrieved context but does not eliminate all injection vectors.”

Indirect prompt injection — the zero-click vector

The threat that makes guardrails non-optional is indirect prompt injection, where the malicious instruction rides inside retrieved content rather than the user’s message. CVE-2025-32711 (patched June 2025) is the enterprise standard: a zero-click exploit in an enterprise copilot where an attacker emails hidden instructions that tell the agent to scan recent messages for sensitive keywords and append the results to an attacker-controlled URL. The user never interacts with the prompt — the retrieval of the email is itself the trigger. This is why the instruction hierarchy must treat retrieved context as untrusted by default: the schema is system prompt > retrieved context > user input, and untrusted retrieved content must never be able to steer a privilege-bearing tool call.
The documented mitigations are a layered set, not a single filter: parse external content into fixed-size passages, normalize Unicode to strip hidden characters, strip injection payloads at ingestion, and constrain the agent’s tool scope so retrieved content cannot reach a high-privilege action (the least-privilege principle from Lesson 2, applied to tool calls). Delimiting data from instructions in the prompt — wrapping retrieved text in tags the model is told to treat as quoted, never as commands — is the cheapest structural defense. The DEF CON 33 “shadow data” research extends the threat: embeddings and vector stores are an underprotected attack surface in their own right, leaking via inversion and poisoning, which is why guardrails and the ReBAC/redaction layers compose rather than substitute.
code
1THE 7 RAG SECURITY RISKS that ship most often (map each to a control)23  Risk                                Layer it lives at    Control4  ---------------------------------   ------------------   ----------------------5  Malicious content ingestion         ingest               sanitize + classify6  Over-permissioned retrieval         retrieval            ReBAC pre/post-filter7  Exfiltration via model responses    output               output guardrail + audit8  Indirect prompt injection           retrieval/runtime    system>context>user +9                                                            tool-scope limits10  Vector-database poisoning           ingest/store         provenance + signed11                                                            ingest + drift monitor12  Sensitive leakage via embeddings    store                DP noise + encrypt at rest13  Supply-chain (3rd-party components)  build               pin + review deps1415  Maps onto OWASP LLM01 (injection), LLM02 (sensitive disclosure),16  LLM06 (excessive agency), LLM08 (vector/embedding weaknesses).
DEF CON 33 — Exploiting Shadow Data from AI Models and EmbeddingsDEFCONConference

The permission-leak failure mode, head-on

Bring the permission leak from Lesson 2 together with guardrails, because the two shapes need different remedies. Shape one: the model is told what it may see and honors the instruction — until a user constructs a prompt that bypasses it (and they always will). Shape two: the model is fed a search result and the search system failed to filter. Shape two is closed by pre/post-filter ReBAC (Lesson 2). Shape one is closed by an architectural rule, not a smarter prompt: never put unfiltered retrieval in the prompt; always run the RBAC check on the search result before constructing the prompt. The we45 red-team reports show how this leaks in practice — queries that extracted API keys from a single markdown file embedded weeks earlier, or surfaced “non-public board deck” slides via organizational-jargon keywords. The staging index and the wide retriever function as covert data paths even when the prompt is validated.
repoPresidio — PII detection, redaction & anonymization SDKMicrosoftrepoNVIDIA NeMo Guardrails — programmable rails between code and LLMNVIDIAarticleRAG Systems are Leaking Sensitive Data (red-team field report)we45

Checkpoint

During a pilot you add Presidio redaction at ingest and call PII handling done. What is the senior critique?

AIngest redaction is best-effort and embeddings strip classification — you also need a retrieval re-check and output tokenization, plus DP noise/encryption on the vectorsBPresidio is sufficient because it is a Microsoft-supported frameworkCRedaction is unnecessary if the deployment is in a VPC
Sign up free to answer and see why

Checkpoint

An attacker emails a document containing hidden instructions; when the agent later retrieves it, it exfiltrates data to a URL — the user never typed anything malicious. What class of attack is this and the core defense?

ADirect prompt injection; fix by sanitizing the user’s input fieldBIndirect prompt injection (CVE-2025-32711 class); treat retrieved context as untrusted (system > context > user), strip payloads at ingest, and constrain tool scope so retrieved content can’t trigger privileged actionsCA model hallucination; fix with a larger model
Sign up free to answer and see why

Checkpoint

An interviewer asks: “Is contextual grounding enough to prevent prompt injection?” Best answer?

AYes — limiting responses to retrieved context fully eliminates injectionBGrounding is irrelevant to injectionCNo — it helps by limiting responses to retrieved context but does not eliminate all injection vectors; you still need input/output guardrails, tool-scope limits, and untrusted-context handling
Sign up free to answer and see why

Checkpoint

After enabling strict injection guardrails, your own internal QA prompts that say “ignore the previous instructions” start getting refused. What is the right fix?

AShip an explicit allow-list for known-good internal patterns and tune thresholds, accepting that strong guardrails create a quality/false-positive tradeoff to manageBRemove the injection guardrail since it produces false positivesCTell QA to never use that phrase again
Sign up free to answer and see why

Checkpoint

A customer asks: “If our vector database is breached but contains only embeddings, what’s the real exposure?” Strongest answer?

AEmbeddings are just numbers, so a breach reveals nothing usefulBEmbedding inversion can reconstruct much of the original text (50-70%), and the store is also a poisoning target — mitigate with DP noise at embedding time and encryption at restCThe only risk is that the attacker learns how many documents you have
Sign up free to answer and see why

Interview prep

The security section of an FDE RAG interview rewards a structured threat-model walk, not a list of products. Structure it as: (1) the threat model — name concrete attacks (permission leak, indirect injection, embedding inversion, RAG poisoning); (2) PII redaction at three stages; (3) the guardrail stack at three layers with defense in depth; (4) the specific injection countermeasure you chose and its known limitation. Weak candidates stop after “we have RBAC” and skip the threat model entirely. Always pair a control with its caveat — that pairing is the senior signal.
  1. 01“What can go wrong in enterprise RAG?” → permission leak, sensitive-data emission, indirect prompt injection, vector poisoning, embedding inversion, supply-chain — mapped to OWASP LLM01/02/06/08.
  2. 02“Where do you redact PII?” → three mandatory stages: ingest (pre-index), retrieval re-check, output tokenization; plus DP noise/encryption on the vectors.
  3. 03“What’s the guardrail stack?” → input (schema + injection detection), retrieval (filters + thresholds), output (groundedness + schema + policy); defense in depth, both sides, never one tool.
  4. 04“What is indirect prompt injection?” → malicious instructions in retrieved content (CVE-2025-32711); retrieval is the trigger; defend with system>context>user hierarchy + tool-scope limits.
  5. 05“Is contextual grounding enough?” → helps but does not eliminate all injection vectors (AWS caveat); layer guardrails and constrain tools.
  6. 06“Worst case if the vector DB leaks?” → embedding inversion recovers 50-70% of text + poisoning; mitigate with DP noise and encryption at rest.
  7. 07“Do guardrails ever hurt?” → yes — over-refusal on legitimate prompts; ship an allow-list and tune thresholds; redaction can eat recall, so go per-data-class.
  8. 08“How do you prove guardrails work in a pilot?” → demonstrate one real block with one audit-log entry by end of week one, or they’re decorative.
Going deeper. The follow-ups push on realism. “Name a security trade-off you actually made.” (the most-quoted FDE question — bring one named, defended choice, e.g. “we accepted a small recall hit from DP-noised embeddings because the corpus held regulated PII”; never “I never made security trade-offs”). “A guardrail blocks a legitimate answer in front of the customer — what do you do?” (own it, show the audit entry that explains the block, tune the allow-list — don’t silently disable the control). “How do you red-team the permission leak?” (friendly user, 20 phrasings, anything that returns a forbidden chunk is a re-auth bug). The hidden rubric is whether you treat security as a set of reasoned trade-offs with known limits, not a checklist of product names.

Could you redact PII at three stages, stand up the guardrail stack, defend against indirect injection, and close the permission-leak shape architecturally?

New to itGetting thereConfident

Takeaways

  • Guardrails are the last line because every prior layer can fail; each must terminate in a real decision — redact, block, route, alert.
  • Redact PII at all three stages — ingest (pre-index), retrieval re-check, output tokenization — each catches the prior miss.
  • Vectors are not safe for raw PII: embedding inversion recovers 50-70% of text; add DP noise and encryption at rest.
  • Run the guardrail stack on both sides with defense in depth (Presidio + NeMo + Bedrock/F5); never trust a single tool.
  • Indirect prompt injection (CVE-2025-32711) rides in retrieved content; enforce system>context>user and constrain tool scope.
  • Close the permission leak architecturally: ReBAC on the retrieval result, never unfiltered retrieval in the prompt; red-team it quarterly.

Next: the capstone — design and defend a private RAG for a regulated customer, end to end.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.