Lesson 2 of 8 · 48 min

ICP definition and scoring services that don't rot

Build ICP as a versioned scoring service: structured features, bands under capacity constraints, reason codes, human overrides with TTL, and anti-rot calibration — not a static slide.

Static ICP slide → living scoring service

ICPs die as PDF slides. Six months later SDRs still spray “Series B+ SaaS in the US” while the winning segment drifted to fintech ops teams on Snowflake. GTM interviews grade whether you treat ICP as a versioned scoring service with features, weights, thresholds, reason codes, and decay — not a brainstorm sticky note. This lesson builds that service on top of L1’s state machine.
Scoring answers one question: is this account/contact eligible for scarce outbound capacity? Capacity is limited by domain reputation, enrichment budget, and AE calendars. Scoring is prioritization under constraint. If everything scores 90, you built a participation trophy, not a model.
Separate ICP definition (who we serve and non-goals) from scoring implementation (features → score → band). Definition is product/GTM strategy. Implementation is engineering: feature store-ish tables in Clay or warehouse, batch + streaming updates, audit logs when humans override.

ICP definition — include, exclude, evidence

Write ICP as testable predicates. Soft vibes (“innovative teams”) cannot be scored. Prefer firmographics, technographics, behavioral intent, and fit-to-motion (PLG vs enterprise). Always list non-goals with the same precision as goals.
text
1ICP CARD (versioned document + machine fields)23  icp_version: 2026-07-14  motion: outbound_mm_fintech_ops5  includes:6    - industry in {fintech, banking_software}7    - employees 50..5008    - geo in {US, CA, UK}9    - tech_any {salesforce, netsuite, snowflake}10  excludes:11    - industry in {crypto_exchange, payday_lending}  # risk/brand12    - employees < 20                                 # support cost13    - existing customer or opp open                  # suppression14  positive_signals: hiring_ops, tech_job_posts, funding_24m15  evidence: 12 closed-won, 4 churn notes, win-theme tags1617  Rule: no score deploy without icp_version bump + changelog.
  1. 01Firmographic — industry, size, geo, ownership, funding stage.
  2. 02Technographic — stack fit / displacement triggers (CRM, data warehouse, SEP).
  3. 03Motion fit — can they buy how we sell (PLG seat vs procurement)?
  4. 04Intent / timing — jobs, news, product usage (if PLG), web intent.
  5. 05Relationship — champion present, competitor install, prior conversations.
  6. 06Suppression — customer, competitor, do-not-contact, litigation, student.

Feature design for scores

Features should be observable, refreshable, and costed. A feature that needs a $2 enrichment to compute cannot sit on every raw lead — put it behind a cheap prefilter. Production metaphor: feature computation is a waterfall cousin (L3).
text
1SCORING FEATURES (example)23  feature                  type        refresh     cost4  -----------------------  ----------  ----------  --------5  industry_naics           categorical  90d         low6  employee_band            ordinal      30d         low7  geo_country              categorical  90d         low8  has_salesforce           boolean      60d         med9  hiring_ops_roles_30d     boolean      7d          med10  funding_last_24m         boolean      30d         low11  intent_topic_fit         0..1         7d          high12  open_opp_or_customer     boolean      1d          low (CRM)1314  score = weighted sum (or rules tree) → band15  bands: A (auto-sequence), B (review), C (nurture/ads), Z (suppress)
Start with a transparent rules model. ML ranking can come later when you have labeled outcomes (meetings, opps, revenue) and a feedback loop. Interviews distrust black boxes without training labels or calibration plots — and GTM rarely has clean labels on day one.

Thresholds, bands, and capacity

Thresholds should be set by capacity, not by vibes. If you can ethically sequence 2,000 new contacts/week, the A-band cutoff is wherever the top 2,000 fall after suppressions — then you validate quality with spot checks and outcome rates.
  1. 01A-band — auto-enroll to standard sequence; highest fit.
  2. 02B-band — human research or personalized POD; high value, incomplete data.
  3. 03C-band — holdout, ads, or product-led only; not worth cold domain risk.
  4. 04Z-band — hard suppress with reason; never “just one more experiment.”
  5. 05Override — AE/SDR can promote with logged reason + expiry (not forever).

Rot, decay, and recalibration

ICP rot is real: markets shift, your product ships upmarket, a competitor saturates a segment. Build time decay on intent features, stale feature flags, and a quarterly calibration ritual that compares score bands to win rates.
text
1ANTI-ROT CHECKLIST23  [ ] icp_version + owner + review date on the score record4  [ ] feature freshness timestamps; block A-band if critical feature stale5  [ ] intent signals decay (e.g. half-life 14–30 days)6  [ ] monthly band→outcome report: meetings, opps, ACV, cycle time7  [ ] shadow mode for new weights before cutover8  [ ] suppressions synced from CRM nightly (customers, opps, unsub)9  [ ] sample 20 A-band accounts for human QA weekly1011  Red flags: A-band win rate ≈ C-band; override rate >15%;12  >20% of A-band missing industry or employee_band.
Production metaphor: scoring is a feature flag + config service for who enters the factory. Rot is config drift. Shadow mode is canary deploy for weights. You would not push a payment config without canary — do not push ICP vNext to 100% of domains either.

Clay as scoring workspace

Clay tables are excellent for composing features from APIs and CRM — treat them as a transformation layer, not the long-term system of record for scores. Write scores and reason codes back to CRM or a warehouse table with icp_version. Other tools (sequences, ads) consume the band, not a fragile Clay view URL.

Interview ways — scoring design

  1. 01Q: How do you stop SDRs from working bad accounts? Suppress Z-band in the sequence engine, not only in a Notion doc. Route B-band to review queues. Publish band→outcome dashboards so culture follows the data. Policy without enforcement is a slide.
  2. 02Q: Rules vs ML for ICP? Rules first for auditability and cold-start. ML when you have enough labeled outcomes and stable features; keep a rules fallback. Hybrid: ML ranks inside A/B-eligible population defined by hard rules.
  3. 03Q: Account score vs contact score? Account fit gates budget; contact persona + seniority gates messaging and routing. Multi-thread only inside eligible accounts. Never sequence a perfect persona at a Z-band account.
When prompted live, sketch: ICP card → feature table → formula/weights → bands → suppressions → write-back → calibration metrics. Mention one failure mode (stale technographics, geo mis-tag, student emails) and how the system detects it.

Suppression as part of scoring (not a side spreadsheet)

Eligibility is score and suppressions. Customers, open opps, competitors, litigation holds, unsubs, and hard bounces must gate enroll regardless of a high fit score. Sync suppressions on a short TTL from CRM and ESP — overnight is often too slow for unsub compliance.
  1. 01Hard suppress — never enroll (unsub, competitor, legal, customer if policy says so).
  2. 02Soft suppress — cooldown TTL after breakup sequence or “not now.”
  3. 03Channel suppress — email dead but LinkedIn allowed (explicit policy).
  4. 04Persona suppress — wrong seniority; keep account eligible for other contacts.
Interview closer: “Scoring without suppressions is a recommendation engine; scoring with enforced suppressions is a production eligibility service.” Say both halves.
ICP is a product decision encoded as data. If only founders can explain who you sell to, your scoring service has a single point of failure.
articleClay blog — finding and scoring ICP accountsClaydocsClearbit / Breeze style firmographic enrichment conceptsClearbitdocsHubSpot target account scoring patternsHubSpotdocsSalesforce Einstein lead scoring (conceptual companion)Salesforce

Checkpoint

Your A-band includes any company with “AI” in the About page via an LLM scrape. Win rate is flat vs random. What broke?

ANothing — brand mention of AI is a gold-standard firmographicBFeature quality: ungrounded LLM keyword fit replaced testable predicates; recalibrate with structured industry/size/tech and measure band liftCIncrease the weight of the AI mention feature until meetings rise
Sign up free to answer and see why

Checkpoint

Sales wants permanent override to force any account into A-band. Your design?

AAllow unlimited overrides — sales knows best and data is always wrongBAllow override with reason code, owner, TTL (e.g. 30–60d), and dashboard of override vs model win ratesCBan all overrides to protect model purity
Sign up free to answer and see why

Checkpoint

Intent “hiring ops managers” is still scoring full points 5 months later. Fix?

ALeave it — hiring pages often stay up and still indicate budgetBApply time decay / expiry on intent features and re-enrich on a short TTL before A-band enrollmentCDelete all intent features because they go stale
Sign up free to answer and see why

Checkpoint

Interviewer asks where scores should live long-term. Best answer?

AOnly inside the sequence tool’s tags — closest to send timeBMaterialize band, score, reason codes, icp_version, computed_at to CRM/warehouse; Clay/transforms are buildersCIn a Slack canvas updated weekly by RevOps
Sign up free to answer and see why

Checkpoint

You have 800 closed-won with tags and 40k cold accounts. First scoring approach?

ATrain a deep model on day one — you have “enough” enterprise dataBShip versioned rules from win analysis + suppressions; track band lift; consider ML ranker only after clean outcomes and stable featuresCScore everyone equally and let sequences A/B the market
Sign up free to answer and see why

Can you design an ICP scoring service with features, bands, overrides, and anti-rot controls?

New to itGetting thereConfident

Takeaways

  • ICP is versioned policy; scoring is a service with features, bands, and reason codes.
  • Capacity sets thresholds; overrides need TTL and measurement.
  • Decay + calibration fight rot; Clay builds, CRM/warehouse remembers.
  • Next: enrichment waterfalls — coverage, cost, freshness (gos-enrichment).

Next lesson: compose enrichment waterfalls that hit coverage targets without melting the budget.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.