Lesson 2 of 6 · 48 min

APIs & integration patterns

Idempotency keys at the business-event level, full-jitter backoff that doesn’t synchronize into a thundering herd, HMAC-verified webhooks with a replay window and a dead-letter queue. The Stripe-grade integration envelope that stops your first deployment from double-charging the customer.

Where FDE engagements silently corrupt customer state

This is the lesson that stops your first deployment from double-charging the customer or losing a ticket update. Three patterns form one inseparable architectural unit — idempotency, retries, webhooks — and the moment you call a write across a network boundary you owe all three. Get one wrong and the failure is silent: a retried request creates a duplicate Salesforce account, a synchronized retry storm turns a 30-second blip into an outage, an unverified webhook lets a spoofed event mutate state. Salesforce’s own architect docs are blunt: “remote procedures must be idempotent… if idempotency isn’t implemented, repeated invocations of the same message can have unintended consequences.” Interview angle. FDE system-design rounds will hand you exactly this — “design a resilient integration between our platform and the customer’s CRM” — and grade whether the envelope runs across the boundary, not inconsistently per call.
The senior framing: an integration is not “call the API.” It is a signed, idempotent, retry-wrapped, dead-lettered envelope that you run identically on every write and every inbound event. The three pillars are documented to the letter in the Stripe API reference and the system-design-primer, and they compose — idempotency is what makes retries safe, retries are what make transient failures survivable, and webhooks are how you learn the customer’s system changed without polling it to death.
Designing Idempotent API Endpoints for Payments at StripeArpit Bhayani

Idempotency: the single most important pattern

Stripe’s rule, verbatim and worth memorizing: “for every given idempotency key, Stripe saves the resulting body and status code of the first request, regardless of whether the request succeeded or failed, including 500 errors. Subsequent requests using the same key will return the same stored result.” The concrete constraints: keys are up to 255 characters, you use v4 UUIDs or high-entropy random strings, you never use PII or email addresses as keys, the key applies to POST (GET/DELETE are idempotent by definition), keys are pruned after 24 hours, and — critically — if you reuse a key with different parameters, Stripe returns an error rather than silently applying the new payload. Brandur Leach’s framing is the philosophy: operations that must run “exactly once and no more” have to be idempotent, or “inconsistencies in distributed state caused by failures” become double-charges and lost webhooks.
The mistake that bites FDEs hardest is where you generate the key. Generate it at the business-event level — once per inbound email, per Slack command, per ServiceNow incident — and persist it on the request row. If you derive it from a server-side timestamp or regenerate it per retry, then a queue that redelivers the same logical event three times resolves to three records, which is the exact bug idempotency was supposed to prevent. Interview angle. “Where does the idempotency key come from?” Strong answer: the call site, keyed to a stable business identifier, generated once and stored — not the retry layer, never a timestamp.
python
1import uuid, httpx23# Generate the key ONCE per business event (here: an inbound support email),4# persist it on the row, and reuse it on every retry of that same event.5def push_to_crm(event_row, payload):6    if event_row.idempotency_key is None:7        event_row.idempotency_key = str(uuid.uuid4())   # v4, <=255 chars, no PII8        db.save(event_row)                               # survive process restart910    return httpx.post(11        "https://crm.example.com/v2/contacts",12        json=payload,13        headers={"Idempotency-Key": event_row.idempotency_key},14    )1516# Retried by the queue 3x for one email -> ONE contact, because the key is stable.17# Reuse the key with a DIFFERENT payload -> the server should reject (param drift),18# not silently apply the new body. Stripe returns an error here, by design.

Retries: full jitter, or you build a thundering herd

The system-design-primer encodes the canonical posture: when a queue saturates and returns HTTP 503, “clients can retry the request at a later time, perhaps with exponential backoff.” But the AWS Builders Library makes the sharper point that most engineers miss: naive (deterministic) exponential backoff is dangerous, because it synchronizes retries across correlated failures — every client that failed at the same instant retries at the same instant, and the retry storm extends the very outage it’s reacting to. The fix is full jitter: sleep = random(0, min(cap, base * 2 ** attempt)). This matters acutely in FDE work because the customer’s downstream (ServiceNow, SAP, Salesforce) fails intermittently, and a synchronized storm on top of that blip is the difference between “incident” and “outage.”
code
1RETRY STRATEGIES UNDER A CORRELATED DOWNSTREAM OUTAGE23  Strategy                       Behavior under correlated failure        Use it?4  ----------------------------   --------------------------------------   --------------------5  constant retry (every 1s)      sustains maximum overload                never on an external call6  deterministic exp backoff      synchronizes retries, extends outage     almost never7  decorrelated jitter            decouples retry timing                   good intermediate8  full jitter                    maximizes decoupling                     AWS-recommended default9  adaptive + circuit breaker     skips retries when failure > SLO         best for tenant fairness1011  full jitter:  sleep = random(0, min(cap, base * 2 ** attempt))12  Hard rules:  never retry 4xx (validation) -- only 5xx / 429 / 408.13               cap attempts (~5); after that, DEAD-LETTER for human review.
Two hard rules ride along. Never retry a 4xx — a 400/422 is a validation error and will fail identically forever; only retry transient classes (5xx, 429, 408). And cap attempts (≈5) with a hard give-up that dead-letters the job, so a permanently-broken downstream doesn’t spin retries until the heat death of the universe. In Python reach for tenacity, in TypeScript p-retry, but configure them — the defaults are rarely full-jitter. Interview angle. “Walk me through your retry policy.” The tell of a senior answer is naming full jitter (not “exponential backoff”), the 4xx-vs-5xx distinction, and the dead-letter terminus.
A free lever many FDEs miss: HTTP verb choice gives you idempotency for nothing. GET, PUT, and DELETE are idempotent by definition — a PUT that sets a record to a target state can be retried infinitely with no extra effect, whereas POST (create) is the one that needs an explicit idempotency key. So when you control the verb, prefer a set-to-state PUT over an increment or append operation; you’ve made the call naturally idempotent rather than bolting safety on. The Enterprise Integration Patterns term for the second category is natural idempotency — the operation’s own semantics make replay safe — versus explicit de-duplication via a key store.

Webhooks: verify, then parse, then dedupe, then act

Inbound webhooks are how you do change-data-capture without polling — but they are an at-least-once channel (Stripe retries up to 3 days with exponential backoff), so duplicates are normal, not exceptional. Stripe signs every event with HMAC-SHA256 in a Stripe-Signature header and the receiver must verify it against the raw request body + the signing secret — “any manipulation to the raw body will cause the verification to fail.” The header includes a timestamp with a default 5-minute tolerance to mitigate replay attacks. The non-negotiable ordering is: (1) verify signature → (2) parse → (3) deduplicate by event id → (4) side-effect. Invert any two and you have a vulnerability or a double-apply.
python
1import hmac, hashlib, time23# (1) VERIFY before anything else -- on the RAW body, with timestamp tolerance.4def verify(raw_body: bytes, sig_header: str, secret: str, tolerance=300):5    ts, sig = parse_header(sig_header)            # "t=...,v1=..."6    if abs(time.time() - int(ts)) > tolerance:    # 5-min replay window7        raise ValueError("timestamp outside tolerance")8    expected = hmac.new(secret.encode(), f"{ts}.".encode() + raw_body,9                        hashlib.sha256).hexdigest()10    if not hmac.compare_digest(expected, sig):    # constant-time compare11        raise ValueError("bad signature")1213def handle(raw_body, sig_header):14    verify(raw_body, sig_header, WEBHOOK_SECRET)  # 1. verify15    event = json.loads(raw_body)                  # 2. parse (only after verify)16    if seen(event["id"]):                         # 3. dedupe on (source, event_id)17        return 200                                #    at-least-once -> idempotent receiver18    enqueue(event)                                # 4. side-effect via durable queue19    return 200
For resilience at the customer site, the Stripe.dev “resilient webhook handlers” architecture is the reference: verify at the edge, persist the raw event to a durable store first, dispatch side-effect work through a queue + worker, and route anything that still fails after N attempts into a dead-letter queue with the full payload preserved so you can replay it once the underlying failure (often a downstream Salesforce outage) clears. The Enterprise Integration Patterns name for step 3 is the Idempotent Receiver: a dedup store keyed on (source, event_id), because “most applications provide an at-least-once policy.”
The subtle bug that bites at day 4: Stripe prunes idempotency keys after 24 hours, but a source can replay a webhook up to 3 days later. If your only defense is the 24h key window, a replayed event on day 3 lands as a fresh side effect. The fix is to combine the short response-cache window with a longer-lived dedup store (or a downstream checksum) keyed on the event id, so replays outside the cache window still resolve to a no-op. This is exactly the kind of seam an interviewer probes after you say “I’d dedupe the webhook.”

Back-pressure & circuit breakers: respecting the customer’s ceiling

Full-jitter retries are necessary but not sufficient when the downstream is a hard rate ceiling rather than a transient blip — and at enterprise customers it almost always is. Salesforce enforces a daily API-call limit tied to license tier; blow past it and the whole org’s integrations get throttled, not just yours. The senior pattern is to layer an adaptive limiter + circuit breaker on top of jittered retries: when the downstream failure rate crosses your SLO, the breaker opens and you stop sending (failing fast or shedding to a queue) instead of hammering a saturated dependency. This is also the tenant-fairness mechanism — one runaway customer can’t starve the others.
python
1# Circuit breaker over the customer dependency: stop retrying when it's clearly down.2class Breaker:3    def __init__(self, threshold=5, cooldown=30):4        self.fails = 0; self.threshold = threshold5        self.open_until = 0; self.cooldown = cooldown67    def call(self, fn, *a):8        if time.time() < self.open_until:9            raise Tripped("breaker open -- shed to queue, don't hammer")10        try:11            r = fn(*a); self.fails = 0; return r           # success resets12        except Transient:13            self.fails += 114            if self.fails >= self.threshold:               # too many failures15                self.open_until = time.time() + self.cooldown   # OPEN: stop sending16            raise1718# Pair with: back-pressure (bounded queue), batching under the API ceiling,19# and limits negotiated with the customer BEFORE go-live -- not discovered in prod.

Case studies: the integration envelope at scale

Stripe is the canonical reference precisely because it absorbs the operational complexity (24h response caching, 5-minute tolerance windows, 3-day webhook retries) so its customers don’t have to — the a16z thesis is that this is the same trade FDE-led AI startups make: take on the plumbing so the customer accepts a worse-than-SaaS margin in exchange for reliability. ServiceNow’s Remote Process Sync (RPS) pattern guarantees execution order across the link so an incident mirrored to your AI workflow and back can’t be replayed out of order under flap. Salesforce’s top integration failure modes are dominated by the daily API-call limit (license-tier throttling not modeled in back-pressure) and renamed/deprecated fields — which is why the FDE’s first job is often to negotiate API limits, not write code. The common rule: real-time everything hits the ceiling; batch with retry envelopes is what survives.
Verify the signature first, parse second, deduplicate third, side-effect fourth — do not invert that order. And generate the idempotency key at the business event, not the retry. The whole envelope, every write, every event. — the FDE integration floor.

Interview prep

The FDE system-design round anchors in real integration architecture, not “design Twitter.” Expect “make this integration resilient,” and lead with the envelope. Strong candidates name the specific mechanism (full jitter, HMAC on raw body, business-event key) and the failure mode it prevents; weak ones say “add retries and webhooks” without the safety properties.
  1. 01“Make this write safe to retry.” → idempotency key at the business-event level, v4 UUID, stored on the row; server rejects param drift.
  2. 02“Design your retry policy.” → full jitter (not plain exponential), retry only 5xx/429/408, cap ~5 attempts, dead-letter the rest.
  3. 03“Why is naive exponential backoff bad?” → it synchronizes retries across correlated failures into a thundering herd that extends the outage.
  4. 04“How do you secure an inbound webhook?” → HMAC-SHA256 on the raw body + signing secret, reject outside a ~5-min timestamp tolerance.
  5. 05“Webhooks are at-least-once — how do you avoid double-applying?” → Idempotent Receiver: dedupe on (source, event_id) before any side effect.
  6. 06“Where do you put the side-effect work?” → persist the raw event first, dispatch via a durable queue/worker, DLQ with full payload for replay.
  7. 07“The customer’s CRM keeps throttling you.” → model the API-call limit as back-pressure, batch instead of real-time, negotiate limits up-front.
  8. 08“GET vs PUT vs POST for idempotency?” → GET/PUT/DELETE are idempotent by definition; POST needs an explicit idempotency key.
Follow-ups probe the seams: “what’s your dedup window, and what happens on day 4 when a source replays?” (combine the 24h key window with a longer-lived dedup store or a downstream checksum), “how do you replay a dead-lettered event safely?” (the idempotency key makes replay a no-op if it already applied), and “what do you do when the customer’s SLA demands fixed retry intervals but you want jitter?” (you own the behavior — full jitter; they own the expectation — don’t sell jitter as an interval promise). These are the questions a real FDE who’s been paged at 3am asks.
docsIdempotent requests — the canonical client contract (255-char key, param-drift error)Stripe API ReferencearticleTimeouts, retries and backoff with jitter (why full jitter beats deterministic)Amazon Builders’ LibrarydocsReceive Stripe events in your webhook endpoint (HMAC + raw body + tolerance)Stripereposystem-design-primer — idempotency, message queues, exponential backoffdonnemartinvideoBeyond Webhooks: The Future of Scalable API Event DeliveryNordic APIs

Checkpoint

Your queue redelivers the same inbound email event 3 times. You want exactly one CRM contact created. Where do you generate the idempotency key?

AInside the retry wrapper, fresh each attemptBFrom a server-side timestamp at send timeCOnce at the business-event level (per email), persisted on the row and reused on every retry
Sign up free to answer and see why

Checkpoint

The customer’s ServiceNow has a 20-second outage. 500 of your workers were mid-call and all retry. What prevents the retry from extending the outage?

AFull jitter: sleep = random(0, min(cap, base * 2^attempt)) so retries spread out instead of synchronizingBDeterministic exponential backoff so every worker waits the same growing intervalCRetry immediately and continuously until ServiceNow recovers
Sign up free to answer and see why

Checkpoint

An inbound webhook arrives. Which order of operations is correct and safe?

AParse JSON → run business logic → verify signature on the parsed objectBVerify HMAC on the raw body (within timestamp tolerance) → parse → dedupe on event id → side-effect via a durable queueCDedupe by event id → verify signature → parse → act
Sign up free to answer and see why

Checkpoint

Your integration gets a 422 Unprocessable Entity from the customer’s API. What should the retry layer do?

ARetry it with full jitter up to 5 timesBDo not retry — 4xx is a permanent validation error; surface it / dead-letter for human fix, and only retry 5xx/429/408CRetry it once immediately, then give up
Sign up free to answer and see why

Checkpoint

A client reuses an idempotency key but with a different request body (a code bug). What should a well-designed server do?

ASilently apply the new body — the key was already seenBIgnore the request entirely and return the old cached resultCReturn an error flagging the parameter mismatch, rather than applying the changed payload
Sign up free to answer and see why

Could you design a resilient CRM/ticketing integration — idempotent writes, full-jitter retries, HMAC-verified webhooks, a DLQ — and defend the safety properties in a system-design round?

New to itGetting thereConfident

Takeaways

  • Idempotency + retries + webhooks are one envelope you run on every write and every event.
  • Generate the idempotency key at the business-event level (v4 UUID, stored), apply it to every state-changing call, reject param drift.
  • Use full jitter, not deterministic backoff; retry only 5xx/429/408; cap attempts and dead-letter the rest.
  • Webhooks are at-least-once: verify HMAC on the raw body → parse → dedupe on (source,event_id) → act.
  • Persist raw events first, dispatch via a durable queue, DLQ with full payload for safe replay.
  • Real-time everything hits API ceilings; batch with retry envelopes and negotiate limits up front.

Next: the LLM app/agent layer — function calling, structured outputs, and eval-driven design on top of these integrations.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.