RBAC on the chat surface is easy and worth almost nothing; RBAC on the retrieved document is the enterprise contract, and it is where the leaks happen. Pre- vs post-filter ReBAC (AuthZed/Pinecone), why pre-baked ACLs go stale, SSO down to the agent identity, and the audit schema a regulator actually wants.
Where the leaks actually happen
RBAC on the chat surface — who can open the assistant — is trivial and protects almost nothing. The real enterprise contract is RBAC on the retrieved document: did the system return a chunk this specific user is permitted to read? In most stacks the answer is silently no, because the retriever that selects documents has no awareness of the user’s identity. The embedding model maps the query, top-k similarity search runs, chunks flow to the prompt — and none of those steps ever queried the directory. The retrieval engine is the security boundary nobody is governing.
Make the failure concrete with a real disclosed case. A legal team deployed an internal RAG so lawyers could find precedents; three months in, a routine compliance review found that junior analysts without matter-level authorization could query and receive pre-motion strategy memos and engagement letters tied to senior-led matters. Firewalls, SIEM, and DLP produced zero alerts — the access pattern was an authenticated employee, on an approved endpoint, during business hours, and the leaked content was natural-language privileged advice, not a 16-digit account number with a signature DLP could match. This is the cleanest demonstration that perimeter security cannot see RAG leakage; only retrieval-layer instrumentation catches it.
It gets worse in a lab. In Amine Raji’s reproducible red-team of a 100%-local ChromaDB + LangChain RAG stack (full lab code on GitHub), cross-tenant exfiltration succeeded on 20 out of 20 queries and knowledge-base poisoning landed at a 95% rate; an access-controlled retrieval layer (per-namespace ACLs) is what closes the cross-tenant leak. That is a number on the value of namespace enforcement: the gate is not optional polish, it is the difference between a system that leaks every time and one that mostly does not. Interview angle. “How do you enforce auth at retrieval time?” is a near-guaranteed probe; the strong answer names a concrete pattern (pre/post-filter ReBAC, below), not “we have role-based access.”
The canonical pattern, from AuthZed’s SpiceDB walkthrough with Pinecone, models documents and groups as a Zanzibar-style relationship graph (ReBAC) — e.g. definition article { relation viewer: user; permission view = viewer } — and enforces it in the retrieval critical path one of two ways:
01Pre-filter — first call SpiceDB LookupResources to get the set of document IDs this user may view, then pass that set to the vector DB as a filter (an $in on article IDs). The LLM only ever receives permitted context. Best when the corpus is large and the user can see only a small slice.
02Post-filter — let the vector DB retrieve top-k first, then run a CheckPermission request per returned document ID, forwarding only the chunks that pass. Best when hit rates are high and the negative set is small.
python
1# Post-filter ReBAC: retrieve first, then re-check every hit against the authz graph.2def retrieve_authorized(query, user_id, k=20):3 hits = vector_search(query, k=k) # vector DB knows nothing of identity4 allowed = []5 for h in hits:6 ok = spicedb.check( # microsecond-latency relationship check7 resource=f"document:{h['doc_id']}",8 permission="view",9 subject=f"user:{user_id}",10 )11 if ok:12 allowed.append(h) # only permitted chunks reach the prompt13 return allowed1415# Pre-filter variant: ids = spicedb.lookup_resources("document", "view", f"user:{user_id}")16# hits = vector_search(query, k=k, filter={"doc_id": {"$in": ids}})17# Revoke access later? One relationship write in SpiceDB -- the model never sees the doc again.
The mechanism that makes ReBAC correct where metadata ACLs fail: the authorization decision is computed on read, not precomputed at index time. Bake an ACL into vector metadata and it goes stale the instant a group membership changes — a revoked user keeps getting hits until you re-embed. Compute it on read and a revocation is a single relationship write that takes effect immediately, with the check itself running at microsecond latency synchronously between the retriever and prompt assembly. This is the same bar enterprise search vendors hold themselves to: Glean builds a fully permission-aware assistant that “only sources information the user has explicit access to,” enforcing least privilege by inheriting and live-syncing the ACLs of each source system (Confluence, SharePoint, Box) — and its own engineers flag that the hard part is keeping permissions current at scale, since “permission rules change often” and source API rate limits can lag the sync. The senior rule: ReBAC (or an equivalent policy engine) is the default; never store the gate as embedded metadata alone.
ReBAC is one of three ecosystem flavors you will meet, and the architectural pattern is identical across them — identity checked against every chunk before prompt assembly — so do not let the framework choice become the argument. SpiceDB / AuthZed is the Zanzibar-style relationship graph (Google’s model for Drive-scale sharing). OpenFGA is the CNCF Zanzibar-flavored alternative. OPA (Open Policy Agent) evaluates a policy bundle per request and is the right fit when the customer already runs OPA for Kubernetes/infra authz. The selection is about which the customer already operates, not which is “best.” Interview angle. If asked “SpiceDB or OPA?”, the strong answer is “whichever the customer’s platform already speaks — the pattern (per-chunk check, computed on read, in deterministic code) is the same; I would not introduce a second policy engine just for RAG.”
Pick the pattern by your selectivity
Pre vs post is not a coin flip — it is a selectivity decision with real cost implications, and getting it wrong creates the filtered-ANN recall cliff. If a user can see only 0.5% of a 50M-doc corpus, pre-filtering to that ID set is right, but a naive metadata filter that excludes 99.5% of candidates can wreck ANN recall: HNSW walks a proximity graph, and if almost every neighbor is filtered out the search either returns far fewer than k real hits or has to explore enormously more of the graph to find them. The fixes are filter-aware indexes (partition by tenant) or pre-resolving the allowed set and searching only within it. If instead the user can see most of the corpus, post-filter is cheaper: retrieve k, drop the few they cannot see.
code
1CHOOSING pre- vs post-filter authorization23 Situation Pattern Why4 -------------------------------- ----------- ----------------------------------5 User sees small slice (<5%) pre-filter don't waste ANN on forbidden docs;6 of a large corpus resolve allowed IDs, search within7 User sees most of the corpus post-filter retrieve k, drop the few denied;8 cheaper than resolving a huge set9 Filter excludes ~99% of docs pre-filter + naive metadata filter cratters HNSW10 partitioned recall -> partition index by tenant11 index12 Tenancy is hard isolation per-tenant separate namespace/index per tenant;13 (compliance) namespace no cross-tenant vector can be returned1415 Rule: selectivity drives the choice; a filter that excludes almost everything16 needs a partitioned index, not a post-hoc predicate.
Multi-tenancy has its own named spectrum — AWS frames it as Silo / Pool / Bridge and it is a favorite interview probe. Silo gives each tenant a fully separate stack: highest isolation, most expensive. Pool shares the whole end-to-end pipeline: cheapest, but “does not offer performance isolation at the vector-store level” and a noisy neighbor degrades everyone. Bridge sits between (supports up to ~100 tenants) but “does not allow per-tenant end-to-end encryption.” AWS’s newer pattern uses Amazon Verified Permissions evaluating Cedar policies at query time to derive a metadata filter for retrieval — but note the caveat AWS itself states: that gives “filter-level (logical) isolation, not IAM-enforced (infrastructure) isolation.” Interview angle. When asked which tenancy model to pick, the strong answer is “it depends on tenant tier and compliance” with the tradeoff axis (cost vs isolation vs blast radius) — naming Silo/Pool/Bridge and then comparing them, not just cataloging the names.
SSO must reach the agent identity, not just the human
Enterprise auth means SAML/OIDC SSO bound to the customer’s IdP (Okta, Entra ID) with just-in-time user provisioning — not shared tenant accounts that destroy attribution. The subtle senior point is that SSO must extend down to the agent and the connectors, not stop at the human. The assistant, the MCP servers it calls, and the data connectors should all share one identity fabric — not a side-band API key sitting in a .env file. MintMCP ships this as “OAuth & SSO authentication” plus “centralized credential management” precisely so the agent’s actions are attributable to the same identity graph as the user’s, which is what makes the audit trail meaningful.
Why this matters for audit: if the agent acts under a shared service account, every tool call in your log reads as “service-bot did X” and you have lost the chain back to the human who triggered it. With identity propagated end to end, the log says “user:kim, via the assistant, retrieved document:123” — which is the difference between an audit trail and a useless append-only file. The principle of least privilege is the axis this lives on: scope the agent’s token to exactly the tools and data the current user may touch, so an over-broad service account can never become an over-privileged god-account that retrieves across tenants.
The audit schema a regulator actually wants
Auditability comes from a single chokepoint logging every retrieval and generation. The test: if you cannot replay a query and its exact retrieval set six months later, you cannot defend the system. Regulators are explicit about why — EU AI Act Article 12 (enforcement from 2 August 2026, penalties up to €35M or 7% of global revenue) requires operators of high-risk AI to maintain logs sufficient for post-hoc auditing; HIPAA requires audit controls for PHI access; SOC 2 CC7 requires system-operation monitoring. The plain-language ask is always the same: which documents were retrieved, for which query, by whom, and what controls governed that retrieval.
code
1MINIMUM audit event per request (log BEFORE the call, inside the request flow)23 Field Why an auditor / on-call needs it4 ------------------------- ------------------------------------------------5 user_id, session_id attribution -- who asked, in what session6 trace_id, span_id cross-correlate with distributed traces7 query_text (or hash) what was asked (hash if the query itself is PII)8 doc_ids_returned the EXACT retrieval set -- the replayable core9 + chunk_ids, corpus_id localize to the chunk and the index generation10 embedding_model + version detect representation drift across re-embeds11 chunking_params reproduce the index state at the time12 authz_decision (per doc) which docs passed/failed the ACL check, and why13 guardrail_decision PII/injection/jailbreak verdict + action taken14 prompt_hash, response_hash integrity of what went in and came out15 url_category / reliability source tier (green/yellow/red) for the answer1617 Log the doc_id + chunk_id, NOT the raw chunk text -- or the audit log itself18 becomes a PII store you must then re-protect. Sign + hash; retain to the19 strictest regulator in scope (HIPAA 6 yrs, financial services 7+ yrs).
Two senior refinements the field converges on. First, the audit log should be append-only with cryptographic integrity (a tamper-evident chain), so an auditor can verify chain-of-custody rather than trust a mutable file — Pathway’s llm-app has an open issue specifically requesting signed, tamper-proof audit logs across ingest/retrieve/generate for exactly this reason. Second, store the doc_id and chunk_id, not the raw chunk content; if you log full retrieved text you have built a second copy of the sensitive corpus inside your logging system, and now you must apply the same ACLs to the logs. Keep the chunk text in the access-controlled store and reference it by ID.
A junior analyst’s RAG query returns a privileged strategy memo they are not staffed on, yet firewalls, SIEM, and DLP fired no alert. What does this tell you, and what is the fix?
ADLP just needs better regex patterns to catch the memoBPerimeter tools cannot see retrieval-layer leakage; enforce identity on every retrieved document (pre/post-filter ReBAC) before chunks reach the promptCRestrict who can open the chat assistant
Your corpus has 50M docs and each user can typically see under 1% of them. Which authorization pattern fits, and what must you watch for?
APost-filter: retrieve top-k, then drop the documents the user cannot seeBPre-filter to the allowed ID set, using a tenant-partitioned/filter-aware index to avoid the ANN recall cliffCBake the allowed roles into each chunk’s metadata at ingest and rely on that alone
Why is computing the authorization decision on read (ReBAC) preferable to precomputing ACLs into vector metadata?
AIt is faster, because metadata filters are always slower than graph checksBIt avoids embeddings entirely, removing the vector DB from the auth pathCA computed-on-read decision reflects revocations immediately (one relationship write), while baked metadata goes stale until a full re-embed
You are designing the audit log. Which design best satisfies a SOC 2 / EU AI Act auditor without creating a second PII store?
AAppend-only, signed events with user_id, trace_id, the returned doc_ids + chunk_ids, the per-doc authz decision, and model/index versions — referencing chunk text by ID, not copying itBLog the full retrieved chunk text for every request so nothing is lostCLog only aggregate daily counts of queries and answers
An interviewer asks you to choose a multi-tenant isolation model for a RAG product with mixed customer tiers. Strongest response?
AAlways use Pool — sharing the whole pipeline is cheapest and simplestBAlways use Silo — full isolation per tenant is the only safe optionCIt depends on tenant tier and compliance: Silo for regulated/high-tier, Pool for low-tier, Bridge in between — trading cost vs isolation vs blast radius
The authorization portion of an FDE RAG interview is where most candidates are exposed: they say “we have role-based access” and stop. The interviewer wants the retrieval-time enforcement pattern, named, with its tradeoff and its known limitation. Structure the security section as (1) the threat — the retriever is ungoverned; (2) the pattern — pre/post-filter ReBAC or Cedar-at-query-time; (3) the tenancy model — Silo/Pool/Bridge, chosen by tier; (4) the audit schema — what you log and why. Quote a verbatim caveat where you have one (“contextual grounding helps mitigate injection but does not eliminate all vectors”) — it signals production experience.
01“How do you enforce auth at retrieval time?” → pre-filter (LookupResources → $in) or post-filter (CheckPermission per hit) against a ReBAC graph; the LLM only sees permitted chunks.
02“Why not just tag chunks with roles at ingest?” → pre-baked ACLs go stale on every group change and can be missing/wrong; compute on read so revocations are immediate.
03“Pre-filter or post-filter — when?” → selectivity: pre-filter when the user sees a small slice, post-filter when they see most of the corpus.
04“What breaks when a filter excludes 99% of docs?” → the filtered-ANN recall cliff; fix with a tenant-partitioned/filter-aware index or search within the resolved allowed set.
05“Silo vs Pool vs Bridge?” → isolation vs cost vs blast radius; pick by tenant tier and compliance, and name the per-tenant-encryption / noisy-neighbor caveats.
06“Does SSO stop at the human?” → no — propagate identity to the agent and connectors (least privilege), or the audit trail loses the chain back to the user.
07“What goes in the audit log?” → user_id, trace/span, the returned doc_ids+chunk_ids, per-doc authz decision, guardrail verdict, model+index versions — append-only, signed, chunk-by-reference.
08“How long do you retain it?” → to the strictest regulator in scope (HIPAA 6 yrs, financial 7+), with cryptographic integrity for chain-of-custody.
Going deeper. The follow-ups probe whether you have actually shipped this. “A user’s access is revoked mid-session — what happens to in-flight and future retrievals?” (post-filter re-checks on every query, so the next retrieval excludes the doc; a relationship write is the only state change). “How do you red-team a permission leak?” (a friendly user tries to retrieve a doc they do not own, phrased 20 different ways; anything that works is a re-auth bug — put it on the quarterly agenda). “Is the audit log itself a compliance risk?” (yes if it contains raw chunks or query PII — log by ID and hash). “Contextual grounding — is that enough to stop injection?” (no; it limits responses to retrieved context but does not eliminate all injection vectors — quote the AWS caveat). The hidden rubric is whether you treat the retriever as a governed boundary.
Could you implement pre/post-filter ReBAC, choose a tenancy model by tier, and specify an audit schema a regulator would accept?
New to itGetting thereConfident
Takeaways
RBAC on the chat surface is worthless; RBAC on the retrieved document is the contract — the retriever is the ungoverned boundary.
Enforce identity at retrieval with pre-filter ($in on allowed IDs) or post-filter (CheckPermission per hit) against a ReBAC graph.
Compute authorization on read, not as baked metadata — revocations must be immediate, not “after the next re-embed.”
Selectivity picks pre vs post; a filter excluding ~99% needs a partitioned index to avoid the ANN recall cliff.
Propagate SSO identity to the agent and connectors (least privilege), or the audit trail loses the human.
Audit = append-only, signed, per-request: user_id, returned doc/chunk IDs, per-doc authz + guardrail decisions, model/index versions — chunk by reference.
Next: private/VPC and air-gapped deployment — terminating every egress in something the customer can route and audit.