Lesson 3 of 8 · 55 min

Multi-region, multi-tenancy, cost, ops

Active-active vs passive, silo vs pool, noisy neighbors, cardinality traps, unit cost, deploy/rollback — evolved through a B2B analytics SaaS story.

What mid-level answers skip

L4 answers often stop at a single-region happy path. L5/L6 answers add topology under failure, tenant isolation, cost attribution, and operability (deploy, rollback, on-call, cardinality). Netflix’s active-active multi-region story (classic eng blog, still cited) and Slack’s Unified Grid re-architecture for large customers (reported) are production anchors for “global + multi-tenant” thinking. This lesson is pure depth on those axes — we close with a B2B analytics SaaS evolving from single-region to multi-region multi-tenant. FinOps one-liner to memorize: tag every fleet by service and tenant tier; review the top three cost drivers weekly; design scale-down as carefully as scale-up. Multi-region doubles baseline cost before traffic arrives — say that when you propose active-active.
Use this layer when the prompt mentions global users, enterprise customers, noisy neighbors, data residency, or “design for 10×.” Even if the interviewer never says “multi-region,” dropping one crisp sentence about failover and cost is free senior signal — as long as you do not derail the MVP. Always pair topology with an RPO/RTO sentence. Multi-region topologies. Active-passive: one region takes writes; standby replicates async; failover is a deliberate cutover (DNS/traffic shift). Simpler consistency; RTO/RPO tradeoffs. Active-active (read local, write local or pinned): lower latency globally; conflict risk if two regions accept writes to the same entity. Netflix’s public active-active write-ups emphasize multi-region resiliency with conflict-free patterns (LWW/CRDT-class techniques for appropriate data) rather than pretending every write is a global serializable transaction. Write-primary pin: user or tenant pinned to a home region for writes; reads local with async replication — common compromise. Routing: geo-DNS / anycast / global LB. Data residency: EU tenants stay in EU (pool or silo per region). Failover drills matter more than the diagram — say how you detect region death (health checks, synthetic probes) and what is lost (RPO) if replication lag was 30s. Uber’s multi-region Kafka DR writing (reported) is a good mental model for “async pipelines need their own failover story,” not only the OLTP primary.
code
1MULTI-REGION TOPOLOGY TRADEOFF2Strategy              RPO              RTO           Cost      Use when3-------------------   ---------------  ------------  --------  ------------------------4Backup + restore      hours            hours         low       cold / non-critical5Pilot light           minutes          ~30 min       medium    async replica warm-ish6Warm standby          seconds          minutes       med-high  read-heavy OK with lag7Active-active multi   ~0 per region*   ~0            2×+       global mission-critical89* Active-active still has conflict windows for concurrent cross-region writes.10Write-home pin avoids most conflicts at the cost of cross-region write latency for travelers.

Conflict strategies (altitude)

Pick by data type: LWW timestamps (simple, loses concurrent updates), version vectors (detect concurrent edits), CRDTs (converge automatically for counters/sets — expensive to get right), tenant-home primary (avoid conflicts). Interview altitude: name 2 strategies, pick default for the entity (profile vs counter vs document), move on unless collab-editing is the prompt. Spanner/Cockroach-class global SQL exists when you truly need external consistency — say the latency and cost tax explicitly.

Multi-tenancy: silo vs pool vs bridge

Silo: per-tenant DB/cluster/account. Max isolation, noisy-neighbor proof, easier residency and delete-the-tenant, high cost and ops toil (migrations × N). Pool: shared cluster, tenant_id on every row. Efficient, needs ruthless quotas and query isolation. Bridge/hybrid: pool by default, silo for enterprise whales — the production default for serious SaaS. Partition keys almost always include tenant_id first for pool models. Never trust tenant_id from the client body alone — bind from auth context.
python
1# Pool model sketch2# events(tenant_id, event_id, ts, payload)  PRIMARY KEY (tenant_id, ts, event_id)3# Enforce tenant_id from auth context — never trust client body alone.45def query_events(tenant_id, start, end):6	assert tenant_id == auth.tenant_id7	return db.query(8		"SELECT * FROM events WHERE tenant_id=%s AND ts BETWEEN %s AND %s LIMIT %s",9		tenant_id, start, end, PAGE,10	)1112# Noisy neighbor: per-tenant QPS quota + max scan bytes + separate rate limit keys1314TENANCY TIER15               Isolation   Cost     Ops toil   Compliance fit16Pool           low         low      low        SMB / standard17Bridge/hybrid  medium      medium   medium     most SaaS at scale18Silo           high        high     high       enterprise / residency / noisy

Senior signals vs mid-level failures

code
1AREA          MID-LEVEL FAILURE                 SENIOR SIGNAL2------------  --------------------------------  -----------------------------------------3Multi-region  “Use Route 53 geo-DNS”            Write path, conflict rules, failover4                                               direction, blast radius, RPO/RTO5Tenancy       “tenant_id column everywhere”     Isolation tier (pool/bridge/silo) + quotas6Cost          “bigger boxes”                    $/request, tags, scale-to-zero off-peak7Ops           “add monitoring later”            SLOs + error budgets from day 18DR            “we have backups”                 Named RTO/RPO, game days, restore drills9Migrations    “drop column in one deploy”       Expand/contract dual-write, backfill, cut

Noisy neighbor, quotas, billing meters

Noisy neighbor: one tenant’s analytical scan saturates shared CPU/IO. Mitigations: per-tenant RPS and concurrency limits, max query cost, separate priority lanes, sticky pools for whales, kill switches. Quotas need a metering pipeline (usage events → aggregator → enforce + bill). Enforcement points: gateway (cheap) + data plane (accurate). Stripe-like products often combine logical isolation with hard API quotas — the systems lesson is “meter then enforce,” not “hope tenants are nice.”

Cardinality traps in observability

Metrics with unbounded labels (user_id, raw_url, email) explode time-series cardinality → cost spikes and slow dashboards. Prefer bounded labels (route template, tenant tier, region, status class). Logs/traces sample on error paths; exemplars link rare traces to metrics. Senior phrase: “I’d label by tenant tier, not tenant id, on the hot metric.” Google SRE practices (error budgets, toil) are the mental model: SLOs decide what pages, not a 40-dashboard museum.

Cost & FinOps sketch

Name a unit cost: $/1k requests, $/tenant/month, storage $/GB-month. Cost drivers: chatty APIs, over-replicated hot data, multi-region write amplification, full-table analytics on OLTP, uncached media. Levers: reserved/committed use for steady state, spot for batch, right-size (many fleets are 2× over-provisioned), S3 lifecycle, CDN offload, kill zombie resources. Scaling ceiling: “this design is fine to ~20K QPS; beyond that we split read replicas and freeze schema changes during reshard.” Interviewers love one honest ceiling.

Deploy, rollback, on-call, zero-downtime migrations

Deploy: rolling / canary / blue-green; feature flags for risky paths. Rollback: expand-contract migrations (compatible dual-read), never break consumers mid-flight. Never drop a column in one deploy. On-call: what pages (error rate, lag, saturation), blast radius, runbook one-liners. Incident flow: stop bleed (rollback/kill switch) → mitigate (rate limit, fail open/closed by risk) → communicate → root cause → postmortem. You do not need a full SRE interview — you need to show the system is operable.

Worked: B2B analytics SaaS evolution

v1 — single region, pool tenancy. Ingest API → queue → workers → events table partitioned by (tenant_id, day). Dashboard reads pre-aggregated rollups in Redis/OLAP. Quotas per API key. Good for 500 tenants. Capacity check (heuristic): 500 tenants × 100 RPS each peak = 50k RPS at gateway — but p99 dominated by the top 5 tenants; average math lies without a heavy-hitter tier. v2 — enterprise whale + noisy neighbor. One tenant’s export job starves others. Move whale to silo cluster; add query admission control; separate batch vs interactive pools; bill on scanned bytes. This is the bridge model in production: 95% tenants pool, 5% silo. v3 — multi-region. EU residency: EU tenants pinned to EU pool. US default region. Async replicate aggregate stats only (not raw EU events) to global executive dashboards. Failover: active-passive per region pool; RPO = replication lag. Conflict: raw events are append-only (easy); tenant config uses home-region primary. GitLab’s classic DB incident postmortem (historical) is a reminder that replication without restore drills is cosplay.
code
1CLIENT (EU)2  → geo-DNS → EU API3  → auth (tenant home=EU)4  → ingest queue (EU)5  → events OLTP/OLAP (EU only)6  → rollup job → eu_rollups7  → (optional) sanitized aggregates → global BI (no raw PII)89CLIENT (US whale silo)10  → US silo cluster, dedicated quotas, same code, different connection string
Staff-level close: “Here’s how we isolate tenants, where data lives, what failover loses, and what metric tells us a noisy neighbor is back.”

Interview answers — senior layer

  1. 01Postgres multi-region? → read replicas local; writes home-region; Spanner/Cockroach only if global writes truly required.
  2. 02Noisy tenant? → tier model, per-tenant pools, rate limits, circuit breakers, silo whales.
  3. 03Cut AWS bill in half? → commit discounts, spot batch, right-size, lifecycle, CDN, kill zombies.
  4. 04RTO/RPO? → give numbers tied to revenue; e.g. RTO 1h RPO 5m for many B2C apps.
  5. 05Zero-downtime migration? → expand/contract dual-write backfill cutover drop.
  6. 06App down walkthrough? → stop bleed, mitigate, communicate, RCA, postmortem.
  7. 07Active-active vs passive? → cost vs availability; most teams over-buy active-active.
  8. 08Schema across 200 services? → backward compatible only; never one coordinated drop.
  9. 09Silo or pool? → default pool; silo for contract/residency/blast radius.
  10. 10What pages you? → error rate, saturation, consumer lag — controlled cardinality labels.
  11. 11What breaks multi-region writes? → conflicts + lag; pin writes or typed merge.
  12. 12Cost sentence? → unit cost + biggest driver + 10× trigger.
articleNetflix — Active-Active for Multi-Regional ResiliencyNetflix Tech BlogarticleSlack — Unified Grid re-architecture for largest customersSlack EngineeringarticleDiscord — How Discord Stores Trillions of MessagesDiscord EngineeringdocsGoogle SRE book — Eliminating toilGoogle SREdocsStripe — Idempotent requestsStripe

Capacity that forces multi-region choices

code
1# Heuristic: global product, 100M DAU, p99 interactive read 100ms2# Cross-ocean RTT ~100–150ms class → cannot serve EU users from US-only primary on hot path3# DECISION: regional read path (replica or cache) minimum4#5# Writes: 5k write QPS global, 70% US / 20% EU / 10% APAC6# If active-active all regions accept writes to same entities → conflict rate matters7# DECISION default: write-home pin by user/tenant; local reads8#9# DR: RPO 5 min, RTO 30 min product promise10# → async replica + runbook may suffice; active-active is overkill if cost doubles for little RTO gain11#12# Tenancy: 10k SMB + 50 enterprise whales13# top 1% tenants = 60% load → hybrid silo for whales before multi-region fancy
Notice the pattern from L2: geography and heavy-hitter tenants appear in the math before the diagram. Mid candidates draw three regions first. Seniors ask whether the product’s p99 and residency actually require it, then pick the cheapest topology that meets RPO/RTO.

Incident and DR game day (interview altitude)

Say you run a game day quarterly: kill a region’s primary, verify traffic shifts, measure actual RPO from lag metrics, restore a backup to a scratch cluster monthly. Cite postmortems as learning: Slack’s 2-22-22 incident write-up (public) is a cascade lesson; GitLab’s 2017 DB incident is the classic “backups you never restore” cautionary tale. In interview you need the habits, not the full postmortem essay.
code
1RPO/RTO CHEAT SHEET (say numbers)2RPO — how much data you may lose (time)3RTO — how long until service is back45Backup only     RPO hours     RTO hours6Async replica   RPO seconds–min   RTO minutes7Sync standby    RPO ~0 (costly)   RTO minutes8Active-active   RPO~0/region      RTO~0 (conflict tax)910Always answer: what do users experience during failover?
Compliance overlays architecture: data residency, right-to-delete, audit logs, encryption keys per region. Enterprise RFPs often force silo + region pin even when pool would be cheaper. Senior signal: treat compliance as a first-class NFR in requirements, not a bolt-on after the diagram.

Unit economics mini worked example

Suppose ingest costs $0.12 / million events and storage $0.023 / GB-month (illustrative cloud list-class heuristics — not a quote). A tenant at 500M events/month and 2 TB stored is ~$60 ingest + ~$46 storage before query compute. If your price is $99 flat, that tenant is underwater once query compute is added — so you need metering, tiered pricing, or silo with reserved capacity. Systems design interviews increasingly reward one sentence of unit economics next to the boxes.

Checkpoint

A pool-tenant analytics cluster has one customer scanning 80% of IO. Best first isolation move?

ADelete the tenantBEnforce per-tenant scan/QPS quotas + consider moving the whale to a silo poolCTurn off monitoring to save CPU
Sign up free to answer and see why

Checkpoint

EU customers require data residency. Product wants a global exec dashboard. Best approach?

AReplicate all raw EU events to US in real time for simplicityBKeep raw EU events in EU; export only approved aggregates/anonymized rollups to global BICRefuse all dashboards
Sign up free to answer and see why

Checkpoint

Enterprise customer demands their data never leaves the EU and never shares a DB with another tenant. Isolation tier?

AShared DB, shared schema, tenant_id column onlyBDedicated DB per tenant + region pinning to EUCDedicated compute but shared multi-region DB
Sign up free to answer and see why

Checkpoint

You add a Prometheus label user_id on every HTTP metric. What goes wrong?

ANothing — more detail is always betterBTime-series cardinality explodes → cost/slow queries; use bounded labels and traces for per-user debugCTLS will break
Sign up free to answer and see why

Checkpoint

Active-active writes for user profile documents across US/EU — safest default conflict strategy for profiles?

ACRDT for free-text bio fields without app logicBPin profile writes to a home region (or require merge UI); avoid silent LWW on critical fieldsCSynchronous dual-commit on every keystroke worldwide
Sign up free to answer and see why

Senior layer decision card

code
1MULTI-REGION PICKER2Need global read p99 < ~100ms? → local reads (replica/cache) minimum3Need RPO~0 and RTO~0 worldwide? → active-active + conflict rules + 2x cost4Most products? → write-home pin + regional reads + async DR5Residency EU? → pin data plane; export aggregates only67TENANCY PICKER8SMB long tail → pool + tenant_id from auth + quotas9Enterprise whale / noisy / contract → silo or bridge10Delete-tenant hard requirement → silo easier; pool needs careful purge1112DR NUMBERS TO SAY13RPO: max data loss window (e.g. 5 min)14RTO: time to restore service (e.g. 30 min)15Validate with game day + restore drill (not backup checkbox)1617COST UNIT EXAMPLES (heuristic list prices vary — label heuristic)18$/1k API calls, $/tenant-month, $/GB-month storage, $/M events ingest19Biggest multi-region footgun: chatty cross-region sync on request path2021OBSERVABILITY22RED: rate, errors, duration23USE: utilization, saturation, errors24Labels: route, region, tenant_tier, status_class — NOT raw user_id25Traces for high-cardinality debug2627MIGRATION28expand → dual-write/backfill → cut reads → drop29Never coordinated drop across 200 services in one deploy3031INCIDENT ORDER32stop bleed → mitigate → communicate → RCA → postmortem

Can you evolve a single-region pool design into multi-region + hybrid tenancy with quotas, failover RPO, and a unit cost sentence?

New to itGetting thereConfident

Takeaways

  • Multi-region: active-passive vs active-active vs write-home — pick by latency vs conflict vs residency.
  • Tenancy: pool default, silo for whales/contracts; always authz tenant_id.
  • Noisy neighbor = quotas + admission + optional silo.
  • Cardinality kills metrics bills; cost needs a unit and a 10× trigger.
  • Operability: canary, expand-contract, pages with blast radius, named RPO/RTO.

Next: full worked design — URL shortener, the cleanest place to practice estimation → storage/ID decisions.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.