Lesson 3 of 8 · 55 min
Multi-region, multi-tenancy, cost, ops
Active-active vs passive, silo vs pool, noisy neighbors, cardinality traps, unit cost, deploy/rollback — evolved through a B2B analytics SaaS story.
What mid-level answers skip
1MULTI-REGION TOPOLOGY TRADEOFF2Strategy RPO RTO Cost Use when3------------------- --------------- ------------ -------- ------------------------4Backup + restore hours hours low cold / non-critical5Pilot light minutes ~30 min medium async replica warm-ish6Warm standby seconds minutes med-high read-heavy OK with lag7Active-active multi ~0 per region* ~0 2×+ global mission-critical89* Active-active still has conflict windows for concurrent cross-region writes.10Write-home pin avoids most conflicts at the cost of cross-region write latency for travelers.Common mistake
“Active-active means strong consistency everywhere.”
Conflict strategies (altitude)
Multi-tenancy: silo vs pool vs bridge
tenant_id on every row. Efficient, needs ruthless quotas and query isolation. Bridge/hybrid: pool by default, silo for enterprise whales — the production default for serious SaaS. Partition keys almost always include tenant_id first for pool models. Never trust tenant_id from the client body alone — bind from auth context.1# Pool model sketch2# events(tenant_id, event_id, ts, payload) PRIMARY KEY (tenant_id, ts, event_id)3# Enforce tenant_id from auth context — never trust client body alone.45def query_events(tenant_id, start, end):6 assert tenant_id == auth.tenant_id7 return db.query(8 "SELECT * FROM events WHERE tenant_id=%s AND ts BETWEEN %s AND %s LIMIT %s",9 tenant_id, start, end, PAGE,10 )1112# Noisy neighbor: per-tenant QPS quota + max scan bytes + separate rate limit keys1314TENANCY TIER15 Isolation Cost Ops toil Compliance fit16Pool low low low SMB / standard17Bridge/hybrid medium medium medium most SaaS at scale18Silo high high high enterprise / residency / noisyKey idea
Key idea
Senior signals vs mid-level failures
1AREA MID-LEVEL FAILURE SENIOR SIGNAL2------------ -------------------------------- -----------------------------------------3Multi-region “Use Route 53 geo-DNS” Write path, conflict rules, failover4 direction, blast radius, RPO/RTO5Tenancy “tenant_id column everywhere” Isolation tier (pool/bridge/silo) + quotas6Cost “bigger boxes” $/request, tags, scale-to-zero off-peak7Ops “add monitoring later” SLOs + error budgets from day 18DR “we have backups” Named RTO/RPO, game days, restore drills9Migrations “drop column in one deploy” Expand/contract dual-write, backfill, cutNoisy neighbor, quotas, billing meters
Cardinality traps in observability
user_id, raw_url, email) explode time-series cardinality → cost spikes and slow dashboards. Prefer bounded labels (route template, tenant tier, region, status class). Logs/traces sample on error paths; exemplars link rare traces to metrics. Senior phrase: “I’d label by tenant tier, not tenant id, on the hot metric.” Google SRE practices (error budgets, toil) are the mental model: SLOs decide what pages, not a 40-dashboard museum.Common mistake
“More metrics always means better production readiness.”
Cost & FinOps sketch
Deploy, rollback, on-call, zero-downtime migrations
Worked: B2B analytics SaaS evolution
1CLIENT (EU)2 → geo-DNS → EU API3 → auth (tenant home=EU)4 → ingest queue (EU)5 → events OLTP/OLAP (EU only)6 → rollup job → eu_rollups7 → (optional) sanitized aggregates → global BI (no raw PII)89CLIENT (US whale silo)10 → US silo cluster, dedicated quotas, same code, different connection stringStaff-level close: “Here’s how we isolate tenants, where data lives, what failover loses, and what metric tells us a noisy neighbor is back.”
Interview answers — senior layer
- 01Postgres multi-region? → read replicas local; writes home-region; Spanner/Cockroach only if global writes truly required.
- 02Noisy tenant? → tier model, per-tenant pools, rate limits, circuit breakers, silo whales.
- 03Cut AWS bill in half? → commit discounts, spot batch, right-size, lifecycle, CDN, kill zombies.
- 04RTO/RPO? → give numbers tied to revenue; e.g. RTO 1h RPO 5m for many B2C apps.
- 05Zero-downtime migration? → expand/contract dual-write backfill cutover drop.
- 06App down walkthrough? → stop bleed, mitigate, communicate, RCA, postmortem.
- 07Active-active vs passive? → cost vs availability; most teams over-buy active-active.
- 08Schema across 200 services? → backward compatible only; never one coordinated drop.
- 09Silo or pool? → default pool; silo for contract/residency/blast radius.
- 10What pages you? → error rate, saturation, consumer lag — controlled cardinality labels.
- 11What breaks multi-region writes? → conflicts + lag; pin writes or typed merge.
- 12Cost sentence? → unit cost + biggest driver + 10× trigger.
Capacity that forces multi-region choices
1# Heuristic: global product, 100M DAU, p99 interactive read 100ms2# Cross-ocean RTT ~100–150ms class → cannot serve EU users from US-only primary on hot path3# DECISION: regional read path (replica or cache) minimum4#5# Writes: 5k write QPS global, 70% US / 20% EU / 10% APAC6# If active-active all regions accept writes to same entities → conflict rate matters7# DECISION default: write-home pin by user/tenant; local reads8#9# DR: RPO 5 min, RTO 30 min product promise10# → async replica + runbook may suffice; active-active is overkill if cost doubles for little RTO gain11#12# Tenancy: 10k SMB + 50 enterprise whales13# top 1% tenants = 60% load → hybrid silo for whales before multi-region fancyIncident and DR game day (interview altitude)
1RPO/RTO CHEAT SHEET (say numbers)2RPO — how much data you may lose (time)3RTO — how long until service is back45Backup only RPO hours RTO hours6Async replica RPO seconds–min RTO minutes7Sync standby RPO ~0 (costly) RTO minutes8Active-active RPO~0/region RTO~0 (conflict tax)910Always answer: what do users experience during failover?Unit economics mini worked example
Checkpoint
A pool-tenant analytics cluster has one customer scanning 80% of IO. Best first isolation move?
Checkpoint
EU customers require data residency. Product wants a global exec dashboard. Best approach?
Checkpoint
Enterprise customer demands their data never leaves the EU and never shares a DB with another tenant. Isolation tier?
Checkpoint
You add a Prometheus label user_id on every HTTP metric. What goes wrong?
Checkpoint
Active-active writes for user profile documents across US/EU — safest default conflict strategy for profiles?
Senior layer decision card
1MULTI-REGION PICKER2Need global read p99 < ~100ms? → local reads (replica/cache) minimum3Need RPO~0 and RTO~0 worldwide? → active-active + conflict rules + 2x cost4Most products? → write-home pin + regional reads + async DR5Residency EU? → pin data plane; export aggregates only67TENANCY PICKER8SMB long tail → pool + tenant_id from auth + quotas9Enterprise whale / noisy / contract → silo or bridge10Delete-tenant hard requirement → silo easier; pool needs careful purge1112DR NUMBERS TO SAY13RPO: max data loss window (e.g. 5 min)14RTO: time to restore service (e.g. 30 min)15Validate with game day + restore drill (not backup checkbox)1617COST UNIT EXAMPLES (heuristic list prices vary — label heuristic)18$/1k API calls, $/tenant-month, $/GB-month storage, $/M events ingest19Biggest multi-region footgun: chatty cross-region sync on request path2021OBSERVABILITY22RED: rate, errors, duration23USE: utilization, saturation, errors24Labels: route, region, tenant_tier, status_class — NOT raw user_id25Traces for high-cardinality debug2627MIGRATION28expand → dual-write/backfill → cut reads → drop29Never coordinated drop across 200 services in one deploy3031INCIDENT ORDER32stop bleed → mitigate → communicate → RCA → postmortemCan you evolve a single-region pool design into multi-region + hybrid tenancy with quotas, failover RPO, and a unit cost sentence?
Takeaways
- Multi-region: active-passive vs active-active vs write-home — pick by latency vs conflict vs residency.
- Tenancy: pool default, silo for whales/contracts; always authz tenant_id.
- Noisy neighbor = quotas + admission + optional silo.
- Cardinality kills metrics bills; cost needs a unit and a 10× trigger.
- Operability: canary, expand-contract, pages with blast radius, named RPO/RTO.
Next: full worked design — URL shortener, the cleanest place to practice estimation → storage/ID decisions.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.