Lesson 3 of 8 · 49 min

Enrichment waterfalls: coverage, cost, and freshness

Design enrichment waterfalls as policy: ordered providers, confidence and validation gates, cache/TTL, clobber rules, cost caps, and bounce feedback — optimizing coverage without burning budget or domains.

Single-vendor hope → waterfall policy

One data vendor never covers your ICP cleanly. GTM engineers design enrichment waterfalls: ordered providers, stop conditions, confidence thresholds, caches, and cost caps — the same idea as a multi-CDN failover, but for emails, titles, and firmographics. This lesson is the production craft of coverage vs cost vs freshness.
Enrichment turns a weak identity (domain + name) into send-ready fields. Required field sets differ by motion: cold email needs valid work email + persona; phone motions need direct dials; ABM ads may only need domain + industry. Define the minimum viable enrichment (MVE) per motion before shopping APIs.
Waterfalls are not “call every API every time.” That pattern burns money and can worsen data by overwriting a high-confidence email with a low-confidence guess. Policy: ordered attempts, accept when confidence ≥ threshold, else continue; never clobber better data without rules.

Waterfall anatomy

text
1EMAIL WATERFALL (simplified policy)23  inputs: first, last, domain, company_id4  providers: [clearbit, apollo, hunter, rocketreach, findymail]5  accept if confidence >= 0.85 OR (validator=valid and conf>=0.70)67  for p in providers:8      if cache_hit(key, ttl=30d): return cache9      resp = p.find_email(inputs)10      meter(p, cost=resp.cost)11      if accept(resp): write(value, source=p, conf, ts); break12  else:13      mark enrich_failed reason=no_email1415  post: validate_email(value) unless provider already validated16  never send on catch-all unless policy.allow_catch_all and band==A
  1. 01Identity waterfall — resolve person/company IDs across systems.
  2. 02Firmographic waterfall — industry, size, revenue, geo.
  3. 03Technographic waterfall — stack detection providers.
  4. 04Email/phone waterfall — highest risk; couple with validation.
  5. 05Insights waterfall — news, jobs, funding (often LLM+search).
  6. 06Verification hop — never skip for cold email at scale.

Coverage, cost, freshness — pick two under budget

You always trade the three. 98% email coverage on a weird ICP may cost 5× and still include catch-alls. Interviews want you to set a target coverage at max $/contact and design the waterfall to hit it — then measure actuals.
text
1ENRICHMENT SCOREBOARD23  metric                         definition4  ---------------------------    ----------------------------------5  coverage_mve                   % eligible with all MVE fields6  cost_per_success               $ / contact reaching MVE7  cost_per_attempt               $ including failures8  median_freshness_days          now - observed_at for email9  validator_valid_rate           valid / validated10  provider_yield[p]              accepts from p / calls to p11  clobber_rate                   overwrites of higher conf data1213  Targets example: coverage_mve ≥ 75%, cost_per_success ≤ $0.35,14  median email freshness ≤ 45d for A-band sends.
Optimize provider order by yield per dollar on your ICP slice, not vendor marketing decks. Re-learn order quarterly — providers drift. Keep a holdout to detect quality regressions when you reorder.

Caching, idempotency, and clobber rules

Cache enrichment results with TTLs by field class: industry long, email medium, intent short. Idempotent keys: (person_key, field, provider) or (person_key, field) with best-value semantics. Clobber rules prevent a late cheap provider from overwriting a validated email.
  1. 01Never overwrite validated email with unvalidated guess.
  2. 02Allow overwrite if new conf higher and validator still valid.
  3. 03Force refresh on bounce, role-change signal, or TTL expiry before re-sequence.
  4. 04Person key — stable ID (CRM id, LinkedIn URL hash, email). Domain+name is weak.
  5. 05Company key — domain primary; legal name secondary; handle acquisitions.

Validation placement

Where you validate matters. Validate too early and you pay on contacts that fail ICP. Validate too late and you burn sequence steps. Common pattern: cheap firmographic filter → email find → validate → only then enroll. Catch-all domains need an explicit policy: skip, send only A-band, or use secondary channel.
text
1VALIDATION OUTCOMES → ACTIONS23  valid          → allow sequence4  invalid        → suppress email channel; try phone/LinkedIn or fail5  catch_all      → policy gate (band, domain reputation, volume)6  unknown        → recheck later or secondary validator; no mass send7  role_account   → suppress for personalized outbound (info@, sales@)89  On hard bounce: mark invalid, stop sequence, quarantine domain if spike.

Clay waterfalls in production

Clay’s waterfall columns are the popular implementation of this pattern — great for iteration. Productionize by: locking schemas, metering credits per base, writing provenance columns, and exporting to CRM. Do not leave the only copy of emails inside an unversioned workbook.

PII, compliance, and multi-client isolation

Enrichment moves personal data. Know region constraints, vendor DPAs, and retention. For agencies (capstone), isolate client waterfalls and keys — no shared caches across clients that could leak contacts. Log access. Delete on request paths must reach cache + SEP + CRM.
  1. 01Q: How do you choose waterfall order? Benchmark on a labeled sample from your ICP: yield, valid rate, cost, latency. Sort by valid-accept per dollar with constraints on max latency. Re-run when vendors change packages.
  2. 02Q: What if coverage is 60% vs 80% target? Do not silently lower confidence thresholds to “hit the number.” Expand providers, improve identity inputs (LinkedIn URL), or narrow ICP. Fake coverage becomes bounce rate.
  3. 03Q: Batch vs real-time enrichment? Batch for list loads and recalibration; real-time for inbound or high-value triggers. Same policy engine both paths — only the scheduler changes.
Interview sketch: MVE fields → providers → accept rules → validation → cache/TTL → cost cap → metrics → bounce feedback loop. Mention one vendor outage plan (skip step, degrade coverage, alert).

Credit metering and noisy neighbors

Treat enrichment credits like a cloud bill: per-tenant budgets, anomaly alerts, and circuit breakers when a bad list explodes spend overnight. Shared agency pools without meters create noisy-neighbor incidents where Client A’s scrape burns Client B’s month.
  1. 01Budget cap — hard stop enrich when tenant hits daily/monthly $.
  2. 02Per-provider caps — avoid one API burning the whole envelope.
  3. 03Dry-run mode — estimate cost on sample before full list load.
  4. 04Anomaly alert — 3× baseline spend/hour pages the owner.
text
1COST GUARD (pseudo)23  before call(provider, tenant):4    if tenant.spend_today + provider.est > tenant.daily_cap: raise BudgetExceeded5    if provider.error_rate_5m > 0.3: skip_provider (degrade)6  after call:7    meter(actual_cost); emit enrich_attempt event
Enrichment is a supply chain. Cheap parts that fail inspection are more expensive than good parts that cost more up front.
docsClay — waterfalls and enrichmentClaydocsHunter — email finder & verifier conceptsHunterdocsNeverBounce validation overviewNeverBouncedocsApollo data & enrichmentApollo

Checkpoint

Waterfall finds an email at conf 0.6, then a later provider returns different email at conf 0.9 unvalidated. Policy?

AAlways keep the first email to preserve idempotency foreverBPrefer higher conf only after validation (or provider-grade proof); never clobber validated with unvalidatedCAverage the confidences and pick randomly between emails
Sign up free to answer and see why

Checkpoint

Finance caps enrichment at $0.20/contact but leadership wants 90% MVE coverage. Current best is 72% at $0.20. Senior move?

ASilently raise conf accept threshold... wait, lower it to 0.4 to hit 90%BPresent the efficient frontier: coverage vs $; propose ICP tighten, better inputs, or budget trade — do not fake coverageCBuy the most expensive vendor only and delete the waterfall
Sign up free to answer and see why

Checkpoint

Hard bounces spike on a list enriched 90 days ago. What system fix is incomplete if you only pause the campaign?

ANothing else — pausing is sufficient and lists age well for a yearBFeed bounces to invalid suppressions, force re-validate/re-waterfall before re-entry, and shorten TTLs for that sourceCSwitch ESP and keep the same addresses immediately
Sign up free to answer and see why

Checkpoint

Where should validation sit for a cost-sensitive cold email motion?

AValidate every raw lead before ICP scoring to maximize cleanlinessBAfter email find and ICP eligibility; before sequence enroll; special-case catch-all by policyCSkip validation if the finder UI shows a green check from 2022
Sign up free to answer and see why

Checkpoint

Agency runs two clients in one Clay workspace with shared credit pool and cache. Primary risk?

ASlightly harder billing — otherwise fine at any scaleBPII/cross-client leakage and entangled metering; isolate workspaces/keys/caches per clientCClay credits are not PII so isolation never matters
Sign up free to answer and see why

Can you design a waterfall with accept rules, validation placement, cost caps, and bounce feedback?

New to itGetting thereConfident

Takeaways

  • Waterfalls optimize valid coverage per dollar with provenance on every field.
  • Clobber rules, TTLs, and bounce loops keep data from rotting into domain damage.
  • Clay is a strong builder; CRM/warehouse + policy hold production truth.
  • Next: sequence engines, branching, and human-in-the-loop (gos-sequences).

Next lesson: sequence graphs that stop on reply, branch on signals, and respect capacity.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.