Idempotency keys at the business-event level, full-jitter backoff that doesn’t synchronize into a thundering herd, HMAC-verified webhooks with a replay window and a dead-letter queue. The Stripe-grade integration envelope that stops your first deployment from double-charging the customer.
Where FDE engagements silently corrupt customer state
This is the lesson that stops your first deployment from double-charging the customer or losing a ticket update. Three patterns form one inseparable architectural unit — idempotency, retries, webhooks — and the moment you call a write across a network boundary you owe all three. Get one wrong and the failure is silent: a retried request creates a duplicate Salesforce account, a synchronized retry storm turns a 30-second blip into an outage, an unverified webhook lets a spoofed event mutate state. Salesforce’s own architect docs are blunt: “remote procedures must be idempotent… if idempotency isn’t implemented, repeated invocations of the same message can have unintended consequences.” Interview angle. FDE system-design rounds will hand you exactly this — “design a resilient integration between our platform and the customer’s CRM” — and grade whether the envelope runs across the boundary, not inconsistently per call.
The senior framing: an integration is not “call the API.” It is a signed, idempotent, retry-wrapped, dead-lettered envelope that you run identically on every write and every inbound event. The three pillars are documented to the letter in the Stripe API reference and the system-design-primer, and they compose — idempotency is what makes retries safe, retries are what make transient failures survivable, and webhooks are how you learn the customer’s system changed without polling it to death.
Stripe’s rule, verbatim and worth memorizing: “for every given idempotency key, Stripe saves the resulting body and status code of the first request, regardless of whether the request succeeded or failed, including 500 errors. Subsequent requests using the same key will return the same stored result.” The concrete constraints: keys are up to 255 characters, you use v4 UUIDs or high-entropy random strings, you never use PII or email addresses as keys, the key applies to POST (GET/DELETE are idempotent by definition), keys are pruned after 24 hours, and — critically — if you reuse a key with different parameters, Stripe returns an error rather than silently applying the new payload. Brandur Leach’s framing is the philosophy: operations that must run “exactly once and no more” have to be idempotent, or “inconsistencies in distributed state caused by failures” become double-charges and lost webhooks.
The mistake that bites FDEs hardest is where you generate the key. Generate it at the business-event level — once per inbound email, per Slack command, per ServiceNow incident — and persist it on the request row. If you derive it from a server-side timestamp or regenerate it per retry, then a queue that redelivers the same logical event three times resolves to three records, which is the exact bug idempotency was supposed to prevent. Interview angle. “Where does the idempotency key come from?” Strong answer: the call site, keyed to a stable business identifier, generated once and stored — not the retry layer, never a timestamp.
python
1import uuid, httpx23# Generate the key ONCE per business event (here: an inbound support email),4# persist it on the row, and reuse it on every retry of that same event.5def push_to_crm(event_row, payload):6 if event_row.idempotency_key is None:7 event_row.idempotency_key = str(uuid.uuid4()) # v4, <=255 chars, no PII8 db.save(event_row) # survive process restart910 return httpx.post(11 "https://crm.example.com/v2/contacts",12 json=payload,13 headers={"Idempotency-Key": event_row.idempotency_key},14 )1516# Retried by the queue 3x for one email -> ONE contact, because the key is stable.17# Reuse the key with a DIFFERENT payload -> the server should reject (param drift),18# not silently apply the new body. Stripe returns an error here, by design.
Retries: full jitter, or you build a thundering herd
The system-design-primer encodes the canonical posture: when a queue saturates and returns HTTP 503, “clients can retry the request at a later time, perhaps with exponential backoff.” But the AWS Builders Library makes the sharper point that most engineers miss: naive (deterministic) exponential backoff is dangerous, because it synchronizes retries across correlated failures — every client that failed at the same instant retries at the same instant, and the retry storm extends the very outage it’s reacting to. The fix is full jitter: sleep = random(0, min(cap, base * 2 ** attempt)). This matters acutely in FDE work because the customer’s downstream (ServiceNow, SAP, Salesforce) fails intermittently, and a synchronized storm on top of that blip is the difference between “incident” and “outage.”
code
1RETRY STRATEGIES UNDER A CORRELATED DOWNSTREAM OUTAGE23 Strategy Behavior under correlated failure Use it?4 ---------------------------- -------------------------------------- --------------------5 constant retry (every 1s) sustains maximum overload never on an external call6 deterministic exp backoff synchronizes retries, extends outage almost never7 decorrelated jitter decouples retry timing good intermediate8 full jitter maximizes decoupling AWS-recommended default9 adaptive + circuit breaker skips retries when failure > SLO best for tenant fairness1011 full jitter: sleep = random(0, min(cap, base * 2 ** attempt))12 Hard rules: never retry 4xx (validation) -- only 5xx / 429 / 408.13 cap attempts (~5); after that, DEAD-LETTER for human review.
Two hard rules ride along. Never retry a 4xx — a 400/422 is a validation error and will fail identically forever; only retry transient classes (5xx, 429, 408). And cap attempts (≈5) with a hard give-up that dead-letters the job, so a permanently-broken downstream doesn’t spin retries until the heat death of the universe. In Python reach for tenacity, in TypeScript p-retry, but configure them — the defaults are rarely full-jitter. Interview angle. “Walk me through your retry policy.” The tell of a senior answer is naming full jitter (not “exponential backoff”), the 4xx-vs-5xx distinction, and the dead-letter terminus.
A free lever many FDEs miss: HTTP verb choice gives you idempotency for nothing. GET, PUT, and DELETE are idempotent by definition — a PUT that sets a record to a target state can be retried infinitely with no extra effect, whereas POST (create) is the one that needs an explicit idempotency key. So when you control the verb, prefer a set-to-statePUT over an increment or append operation; you’ve made the call naturally idempotent rather than bolting safety on. The Enterprise Integration Patterns term for the second category is natural idempotency — the operation’s own semantics make replay safe — versus explicit de-duplication via a key store.
Webhooks: verify, then parse, then dedupe, then act
Inbound webhooks are how you do change-data-capture without polling — but they are an at-least-once channel (Stripe retries up to 3 days with exponential backoff), so duplicates are normal, not exceptional. Stripe signs every event with HMAC-SHA256 in a Stripe-Signature header and the receiver must verify it against the raw request body + the signing secret — “any manipulation to the raw body will cause the verification to fail.” The header includes a timestamp with a default 5-minute tolerance to mitigate replay attacks. The non-negotiable ordering is: (1) verify signature → (2) parse → (3) deduplicate by event id → (4) side-effect. Invert any two and you have a vulnerability or a double-apply.
python
1import hmac, hashlib, time23# (1) VERIFY before anything else -- on the RAW body, with timestamp tolerance.4def verify(raw_body: bytes, sig_header: str, secret: str, tolerance=300):5 ts, sig = parse_header(sig_header) # "t=...,v1=..."6 if abs(time.time() - int(ts)) > tolerance: # 5-min replay window7 raise ValueError("timestamp outside tolerance")8 expected = hmac.new(secret.encode(), f"{ts}.".encode() + raw_body,9 hashlib.sha256).hexdigest()10 if not hmac.compare_digest(expected, sig): # constant-time compare11 raise ValueError("bad signature")1213def handle(raw_body, sig_header):14 verify(raw_body, sig_header, WEBHOOK_SECRET) # 1. verify15 event = json.loads(raw_body) # 2. parse (only after verify)16 if seen(event["id"]): # 3. dedupe on (source, event_id)17 return 200 # at-least-once -> idempotent receiver18 enqueue(event) # 4. side-effect via durable queue19 return 200
For resilience at the customer site, the Stripe.dev “resilient webhook handlers” architecture is the reference: verify at the edge, persist the raw event to a durable store first, dispatch side-effect work through a queue + worker, and route anything that still fails after N attempts into a dead-letter queue with the full payload preserved so you can replay it once the underlying failure (often a downstream Salesforce outage) clears. The Enterprise Integration Patterns name for step 3 is the Idempotent Receiver: a dedup store keyed on (source, event_id), because “most applications provide an at-least-once policy.”
The subtle bug that bites at day 4: Stripe prunes idempotency keys after 24 hours, but a source can replay a webhook up to 3 days later. If your only defense is the 24h key window, a replayed event on day 3 lands as a fresh side effect. The fix is to combine the short response-cache window with a longer-lived dedup store (or a downstream checksum) keyed on the event id, so replays outside the cache window still resolve to a no-op. This is exactly the kind of seam an interviewer probes after you say “I’d dedupe the webhook.”
Back-pressure & circuit breakers: respecting the customer’s ceiling
Full-jitter retries are necessary but not sufficient when the downstream is a hard rate ceiling rather than a transient blip — and at enterprise customers it almost always is. Salesforce enforces a daily API-call limit tied to license tier; blow past it and the whole org’s integrations get throttled, not just yours. The senior pattern is to layer an adaptive limiter + circuit breaker on top of jittered retries: when the downstream failure rate crosses your SLO, the breaker opens and you stop sending (failing fast or shedding to a queue) instead of hammering a saturated dependency. This is also the tenant-fairness mechanism — one runaway customer can’t starve the others.
python
1# Circuit breaker over the customer dependency: stop retrying when it's clearly down.2class Breaker:3 def __init__(self, threshold=5, cooldown=30):4 self.fails = 0; self.threshold = threshold5 self.open_until = 0; self.cooldown = cooldown67 def call(self, fn, *a):8 if time.time() < self.open_until:9 raise Tripped("breaker open -- shed to queue, don't hammer")10 try:11 r = fn(*a); self.fails = 0; return r # success resets12 except Transient:13 self.fails += 114 if self.fails >= self.threshold: # too many failures15 self.open_until = time.time() + self.cooldown # OPEN: stop sending16 raise1718# Pair with: back-pressure (bounded queue), batching under the API ceiling,19# and limits negotiated with the customer BEFORE go-live -- not discovered in prod.
Case studies: the integration envelope at scale
Stripe is the canonical reference precisely because it absorbs the operational complexity (24h response caching, 5-minute tolerance windows, 3-day webhook retries) so its customers don’t have to — the a16z thesis is that this is the same trade FDE-led AI startups make: take on the plumbing so the customer accepts a worse-than-SaaS margin in exchange for reliability. ServiceNow’s Remote Process Sync (RPS) pattern guarantees execution order across the link so an incident mirrored to your AI workflow and back can’t be replayed out of order under flap. Salesforce’s top integration failure modes are dominated by the daily API-call limit (license-tier throttling not modeled in back-pressure) and renamed/deprecated fields — which is why the FDE’s first job is often to negotiate API limits, not write code. The common rule: real-time everything hits the ceiling; batch with retry envelopes is what survives.
Verify the signature first, parse second, deduplicate third, side-effect fourth — do not invert that order. And generate the idempotency key at the business event, not the retry. The whole envelope, every write, every event. — the FDE integration floor.
Interview prep
The FDE system-design round anchors in real integration architecture, not “design Twitter.” Expect “make this integration resilient,” and lead with the envelope. Strong candidates name the specific mechanism (full jitter, HMAC on raw body, business-event key) and the failure mode it prevents; weak ones say “add retries and webhooks” without the safety properties.
01“Make this write safe to retry.” → idempotency key at the business-event level, v4 UUID, stored on the row; server rejects param drift.
02“Design your retry policy.” → full jitter (not plain exponential), retry only 5xx/429/408, cap ~5 attempts, dead-letter the rest.
03“Why is naive exponential backoff bad?” → it synchronizes retries across correlated failures into a thundering herd that extends the outage.
04“How do you secure an inbound webhook?” → HMAC-SHA256 on the raw body + signing secret, reject outside a ~5-min timestamp tolerance.
05“Webhooks are at-least-once — how do you avoid double-applying?” → Idempotent Receiver: dedupe on (source, event_id) before any side effect.
06“Where do you put the side-effect work?” → persist the raw event first, dispatch via a durable queue/worker, DLQ with full payload for replay.
07“The customer’s CRM keeps throttling you.” → model the API-call limit as back-pressure, batch instead of real-time, negotiate limits up-front.
08“GET vs PUT vs POST for idempotency?” → GET/PUT/DELETE are idempotent by definition; POST needs an explicit idempotency key.
Follow-ups probe the seams: “what’s your dedup window, and what happens on day 4 when a source replays?” (combine the 24h key window with a longer-lived dedup store or a downstream checksum), “how do you replay a dead-lettered event safely?” (the idempotency key makes replay a no-op if it already applied), and “what do you do when the customer’s SLA demands fixed retry intervals but you want jitter?” (you own the behavior — full jitter; they own the expectation — don’t sell jitter as an interval promise). These are the questions a real FDE who’s been paged at 3am asks.
Your queue redelivers the same inbound email event 3 times. You want exactly one CRM contact created. Where do you generate the idempotency key?
AInside the retry wrapper, fresh each attemptBFrom a server-side timestamp at send timeCOnce at the business-event level (per email), persisted on the row and reused on every retry
The customer’s ServiceNow has a 20-second outage. 500 of your workers were mid-call and all retry. What prevents the retry from extending the outage?
AFull jitter: sleep = random(0, min(cap, base * 2^attempt)) so retries spread out instead of synchronizingBDeterministic exponential backoff so every worker waits the same growing intervalCRetry immediately and continuously until ServiceNow recovers
An inbound webhook arrives. Which order of operations is correct and safe?
AParse JSON → run business logic → verify signature on the parsed objectBVerify HMAC on the raw body (within timestamp tolerance) → parse → dedupe on event id → side-effect via a durable queueCDedupe by event id → verify signature → parse → act
Your integration gets a 422 Unprocessable Entity from the customer’s API. What should the retry layer do?
ARetry it with full jitter up to 5 timesBDo not retry — 4xx is a permanent validation error; surface it / dead-letter for human fix, and only retry 5xx/429/408CRetry it once immediately, then give up
A client reuses an idempotency key but with a different request body (a code bug). What should a well-designed server do?
ASilently apply the new body — the key was already seenBIgnore the request entirely and return the old cached resultCReturn an error flagging the parameter mismatch, rather than applying the changed payload
Could you design a resilient CRM/ticketing integration — idempotent writes, full-jitter retries, HMAC-verified webhooks, a DLQ — and defend the safety properties in a system-design round?
New to itGetting thereConfident
Takeaways
Idempotency + retries + webhooks are one envelope you run on every write and every event.
Generate the idempotency key at the business-event level (v4 UUID, stored), apply it to every state-changing call, reject param drift.
Use full jitter, not deterministic backoff; retry only 5xx/429/408; cap attempts and dead-letter the rest.
Webhooks are at-least-once: verify HMAC on the raw body → parse → dedupe on (source,event_id) → act.
Persist raw events first, dispatch via a durable queue, DLQ with full payload for safe replay.
Real-time everything hits API ceilings; batch with retry envelopes and negotiate limits up front.
Next: the LLM app/agent layer — function calling, structured outputs, and eval-driven design on top of these integrations.