Three connector archetypes — CRM (Salesforce), ticketing (ServiceNow/Zendesk), doc stores (SharePoint/Confluence) — share one playbook: schema map, scoped per-customer OAuth, idempotent writes, verified webhooks for change-data-capture. Plus the integration-pattern choices, the SSO friction that eats weeks, and the per-document permission trap that leaks data.
The systems where the work actually lives
Three connector archetypes cover almost every FDE engagement: CRM (Salesforce, HubSpot), ticketing (Zendesk, Jira, ServiceNow), and doc stores (Notion, Confluence, SharePoint, Google Drive). They share one playbook — schema-map, scoped OAuth/PAT, idempotent writes, webhook inbound for change-data-capture — and they share one truth: the technology is the easy part, the integration is where the work lives. This lesson is the connector layer that wires the agent from lesson 3 into the customer’s real systems, plus the two things that quietly eat weeks: SSO and per-document permissions.
The cross-cutting rules from the research, applied to all three archetypes: (a) every write goes out with an idempotency-key header; (b) every inbound webhook is HMAC-verified; (c) every change-data-capture stream is checkpointed so a replay doesn’t double-apply; (d) the connector’s blast radius is bounded by a per-customer API token. Design the connector schema-agnostic around one canonical row — <source_type, source_id, mime, last_modified, etag, body> — and commit to it in a Postgres control table before you pull any vendor SDK. That single decision is what lets you add the second and third customer without rewriting the connector.
CRM (Salesforce): pick the integration style on purpose
Salesforce’s Data 360 architecture defines five integration styles you choose deliberately, not by default: request-and-reply (synchronous, waits on the response), fire-and-forget (async, non-blocking), batch (Apex batch classes), data virtualization (real-time external data without replication), and Data Cloud (identity resolution across source-of-record systems). The architect docs are explicit that “remote procedures must be idempotent… if idempotency isn’t implemented, repeated invocations can have unintended consequences” — the lesson-2 envelope, restated by the platform itself. The dominant failure modes are documented and predictable: (1) daily API-call limit exceeded (license-tier throttling not modeled as back-pressure), (2) renamed/mismatched field (schema drift), (3) deprecated API version, (4) type drift (text → picklist), (5) expired token / IP-allowlist change. Interview angle. “The customer’s Salesforce keeps throttling your sync” — the strong answer models the API-call ceiling as back-pressure, batches instead of streaming everything real-time, and negotiates limits before writing code.
The Salesforce FDE judgment call from the research: balance real-time and batch. Real-time updates give immediate responsiveness but increase cost and complexity; batching improves resilience and reduces load. The moment volume crosses the daily cap, batch-with-retry-envelopes is the only viable pattern. And the deeper Palantir doctrine applies here: when a Salesforce field has been repurposed three times, write transformation code — do not click through Setup. Configured solutions on a schema that lies silently produce wrong answers, which is the worst failure mode.
Two operational failure modes round out the Salesforce list and both are auth hygiene: an expired token or IP-allowlist change silently breaks a long-lived integration, and a deprecated API version kills a consumer that skipped the Salesforce retirement cycle. The mitigations are unglamorous and load-bearing: rotate OAuth tokens in a way the customer can audit without paging you, pin a specific API version per environment, and keep a centralized integration registry (endpoint, version, owner, schema hash) so when something breaks you know which connector, on which version, owned by whom. Run all custom logic in a sandbox first — but know that the sandbox catches field-level drift, not license-tier rate limits, so mirror the real limits in staging.
Ticketing (ServiceNow/Zendesk): the pattern decision tree
ServiceNow’s integration decision tree codifies the choices: Integration Hub Spokes (low-code third-party connectivity), Zero Copy Connector (read-only cross-instance reporting, no duplication), REST/JSON (preferred over SOAP for almost everything — lighter, easier to debug), Remote Process Sync (RPS) (real-time, task-based, guaranteed execution order, resilient to temporary connectivity loss), and an event-driven flow with a MID server (the bridge when ServiceNow sits inside a private VPC and you can’t route outbound to your control plane). For bidirectional AI-on-ticketing, RPS is the default: when the customer runs ServiceNow as ITSM-of-record and you run a separate CRM-of-record, RPS preserves order so an incident created in ServiceNow, mirrored to your AI workflow, processed, and mirrored back cannot be replayed out of order under flap conditions.
For Zendesk, the webhook framework subscribes to activity across Support/Guide/Gather/Messaging and is strictly versioned — so pin to a specific API version per environment to avoid signature drift. The connector pattern: subscribe to ticket.created, ticket.updated, comment.created, route through one queue that deduplicates by ticket id (the Idempotent Receiver again). The recurring rule across every ticketing seam — create, status update, reassignment — is that all of them must be safe under at-least-once delivery.
Doc stores: untrusted foreign systems with per-document permissions
Notion, Confluence, and SharePoint expose REST + webhook surfaces; the ingest playbook is uniform — pull docs (Markdown/HTML), chunk by heading (~800-token chunks with ~100-token overlap), embed, upsert — and for outbound use a content-hash idempotency key so a doc update doesn’t create a duplicate in the vector index. The Enterprise Integration Patterns frame is the Document Message: it “passes data and lets the receiver decide what to do,” which is exactly your situation at a customer running SharePoint 2016 with custom workflows you’ve never seen — producer and consumer share no command vocabulary.
The trap that leaks customer data: document stores are permission-scoped at the per-document level, the content might be stale, and two SharePoint sites will claim contradictory authoritative policies. So build a normalization layer that surfaces “documents this user is allowed to see,” not “documents we found.” The three discrete layers — indexing, retrieval, write-back — must stay separate; conflating them is how a RAG answer cites a document the asking user was never cleared to read. Interview angle. “Build an internal AI document search for a bank with complex access control” — the strong answer enforces row/document-level permissions at retrieval time, scoped to the user’s identity, and never indexes around the ACL.
SSO & identity: the per-customer OAuth client boundary
SAML/SSO “takes hours on paper, weeks in practice,” and the research names exactly why: the customer’s IdP may not advertise the metadata URL you expect; their network policy filters by source IP and breaks your redirects; their identity store has group claims you assumed were absent (and lacks ones you assumed present); and their audit log demands every login carry a “source system” attribute you must mint. The mitigations: stand up a discovery worker on-site that can dump the IdP metadata.xml within an hour of arrival, and build your service-to-service auth on top of the SSO you ship so a user re-auth through the IdP doesn’t wedge your backend.
The single most important identity rule for AI deployments, stated plainly in the research: request a separate OAuth client per customer, scoped only to the user’s identity and the data the user is permitted to see — never share clients across customers. A shared client is a cross-tenant data leak waiting for one permission misconfiguration. This is the connector-level expression of the lesson-2 rule that the blast radius is bounded by a per-customer token.
Document store integration for any AI deployment is really three layers — indexing, retrieval, and write-back — over a foreign system whose docs might be stale, might conflict, and are permission-scoped per document. Surface what the user may see, not what you found.
The on-site mitigation that turns weeks into hours: stand up a discovery worker the moment you arrive that introspects the customer’s identity and data landscape — dump the IdP metadata.xml, enumerate the group claims actually present, and list the OAuth scopes the customer is willing to grant. You’d rather discover that their SAML metadata URL isn’t where you expected, or that a group claim you assumed is missing, in hour one than on go-live day. The same instinct applies to change-data-capture: checkpoint every CDC stream (store the last processed cursor/etag), so a connector restart or a webhook replay resumes rather than re-applying the whole history.
python
1# Day-one discovery worker: surface the identity + data surface before you build.2def discover(customer):3 md = fetch(customer.idp_metadata_url) # often NOT where docs say it is4 claims = sorted(set(c for u in sample_users(customer) for c in u.group_claims))5 scopes = customer.granted_oauth_scopes # what they'll actually allow6 return {7 "saml_entity_id": md.entity_id,8 "group_claims_present": claims, # assumed-absent ones bite hardest9 "scopes": scopes,10 "cdc_cursor": load_checkpoint(customer) or "FULL_BACKFILL",11 }12# Knowing this in hour one is the difference between a 1-hour and a 3-week SSO setup.
Case studies: connectors at named companies
Salesforce itself now posts a “Forward Deployed Engineering Lead — Data Science & Integration” role to build tailored integrations for “high-profile customer engagements” — the FDE archetype applied to the CRM-of-record. ServiceNow’s hub-and-spoke topology is the recommended antidote to point-to-point sprawl: every new customer attaches to a small set of centralized adapters (a Salesforce adapter, a ServiceNow adapter, a SharePoint adapter), with bespoke customer-specific code living in a per-customer profile behind the adapter — without it, you ship a unique connector per customer and the platform team lags the field forever. BA Insight’s Confluence connector illustrates the production reality that any “document store integration” is really three layers (indexing, retrieval, write-back). The synthesis: the connectors converge on the same envelope; only the vertical object (account, incident, document) changes.
Build a normalization layer that surfaces “documents you are allowed to see,” not “documents we found.” Per-customer OAuth client; per-document permissions at retrieval. — the controls that separate a deployment from a data breach.
Interview prep
FDE system-design rounds lean hard on connectors and access control — Palantir on ontology + messy-source integration + securing analytics, Anthropic on MCP servers for CRM and RAG over enterprise docs with human-in-the-loop. Lead with the shared playbook (schema-map, scoped OAuth, idempotent writes, verified webhooks), then specialize to the archetype and name the permission/throttle trap.
01“Design a bidirectional sync between our app and the customer’s ServiceNow.” → RPS for ordered incident replication, idempotent writes, dedupe by ticket id, MID server if it’s in a private VPC.
02“The customer’s Salesforce keeps throttling you.” → model the daily API-call limit as back-pressure, batch over real-time, negotiate limits up front.
03“Build AI doc search for a bank with complex ACLs.” → enforce per-document permissions at retrieval against the user’s identity; never index around the ACL.
04“How do you connect to a SharePoint store with workflows you’ve never seen?” → Document Message pattern, schema-agnostic canonical row, content-hash idempotency, three discrete layers.
05“How do you avoid cross-tenant data leaks?” → a separate OAuth client per customer scoped to that user’s data; bounded blast radius per token.
06“Real-time or batch CRM sync?” → balance both; real-time for responsiveness, batch for resilience and to stay under the API ceiling.
07“A Salesforce field was repurposed three times — configure or code?” → write transformation code; configured solutions on a lying schema produce silent wrong answers.
08“Why does ‘simple SSO’ take weeks?” → IdP metadata, IP-filtered redirects, unexpected group claims, audit-log source attributes — stand up a discovery worker on day one.
Follow-ups probe the seams: “how do you split the work between your team and the customer’s?” (you own the adapters and mappers; they own canonical ownership, ACL definitions, and IdP config), “what breaks first in prod?” (an API-limit breach or a renamed field — name it and show the back-pressure/observability that catches it), and “how do you keep a doc update from duplicating in the index?” (content-hash idempotency key on the canonical row). The interviewer — often a real FDE — is checking you’ve wired these into a live customer system, so anchor to a specific connector you built and what surprised you.
Building AI doc search for a bank, you must guarantee a user never sees a citation to a document they can’t access. Correct design?
AIndex everything with a service account, then strip restricted results from the UI before displayBEnforce per-document permissions at retrieval time against the asking user’s identity, so restricted docs never enter the candidate setCEncrypt the documents and let the model decrypt only allowed ones
You’re deploying the same assistant to a second customer. What keeps one customer’s data from ever appearing in the other’s answers?
AA shared OAuth app with a customer_id column to separate the dataBRun both customers in the same vector index and filter by tenant at query timeCA separate OAuth client per customer, scoped to that customer’s users and data — bounded blast radius per token
The customer’s Salesforce field “Status” was originally free text, later repurposed twice. The PM suggests configuring a mapping in Setup. Your call?
AWrite explicit transformation code that disambiguates the historical meanings and stamps provenance, rather than configuring over a lying schemaBConfigure the mapping in Setup — it’s faster than writing codeCTrust the most recent meaning and ignore the historical values
A doc-store sync re-imports a page every time it’s edited, creating duplicate vectors in the index. Cleanest fix?
AUse a content-hash (or etag) idempotency key on the canonical row so an edit upserts in place instead of inserting a duplicateBPeriodically run a dedup job over the vector index to remove duplicatesCStop syncing on edits and only do a full re-index nightly
Could you wire the agent into a customer’s CRM, ticketing, and doc store — with scoped OAuth, idempotent writes, verified CDC webhooks, and per-document ACLs — and design it in an interview?
New to itGetting thereConfident
Takeaways
Three archetypes (CRM, ticketing, doc store) share one playbook: schema-map, scoped OAuth, idempotent writes, verified webhook CDC.
Salesforce: pick the integration style on purpose; model the daily API limit as back-pressure; code over Setup on a repurposed field.
Ticketing: RPS for ordered bidirectional sync; dedupe by ticket id; pin the webhook API version.