
Sales Data Hygiene
- 55 installs
- 96 repo stars
- Updated August 1, 2026
- sales-skills/sales
Sales Data Hygiene is an agent skill that audits CRM completeness, accuracy, duplication, and decay before you fix or enrich records.
About
Sales Data Hygiene is an agent skill for solo founders and indie builders running pipeline in a CRM who need a repeatable quality framework before scaling outreach or trusting dashboards. It structures discovery around completeness percentages on critical fields, accuracy validation on a sample (deliverability, direct dials, title currency against LinkedIn), duplication at contact and account levels including cross-object overlap, and decay awareness with industry-oriented staleness benchmarks cited in the guide. The skill also carries a community learnings log the agent reads at the start of each run and appends to as new platform-specific quirks appear—with an optional path to share back via sales-request-skill when patterns mature. Use it when lists feel “full but wrong,” when reply rates drop without a messaging change, or when you are preparing a migration or integration and cannot afford garbage-in. It measures before it prescribes fixes so effort targets the highest-leverage broken fields rather than blanket enrichment spend.
- Four-pillar audit: completeness, accuracy, duplication rate, and decay rate
- Documented completeness targets (e.g. 95%+ email on contacts, 100% company on accounts)
- Accuracy checks via sampled email verification, phone checks, and LinkedIn title currency
- Calls out cross-object duplication (leads vs contacts) and account spelling variants
- Maintains a living learnings file read each invocation for accumulated CRM gotchas
Sales Data Hygiene by the numbers
- 55 all-time installs (skills.sh)
- +2 installs in the week ending Aug 5, 2026 (Skillselion tracking)
- Ranked #552 of 853 Sales & Marketing skills by installs in the Skillselion catalog
- Security screen: MEDIUM risk (skills.sh audit)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/sales-skills/sales --skill sales-data-hygieneAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 55 |
|---|---|
| repo stars | ★ 96 |
| Security audit | 2 / 3 scanners passed |
| Last updated | August 1, 2026 |
| Repository | sales-skills/sales ↗ |
What it does
Audit and improve CRM record quality—completeness, accuracy, duplicates, and decay—before outbound or pipeline reporting misleads you.
Who is it for?
Best when you're doing B2B sales with a CRM and need structured audits before campaigns, handoffs, or reporting.
Skip if: Pure product analytics funnels with no sales objects, or teams wanting a one-click auto-clean without human sampling and verification discipline.
When should I use this skill?
Starting CRM cleanup, post-import validation, or when pipeline metrics disagree with reality.
What you get
You get a prioritized picture of what is broken, field-level targets, and hygiene actions grounded in measured completeness and sample-based accuracy checks.
- Completeness and accuracy scorecard
- Duplication and decay assessment
- Appended hygiene learnings log entries
By the numbers
- Completeness targets include 95%+ email on contacts and 100% company on accounts
- Framework cites ~30% annual B2B data decay, ~20% sales contact job churn, ~15% direct-dial invalidation
Files
CRM Data Hygiene & Quality
Help the user clean, deduplicate, normalize, and maintain CRM data quality. This skill is tool-agnostic but includes platform-specific guidance for ZoomInfo OperationsOS, Salesforce native tools, HubSpot Operations Hub, Clay, LeanData, RingLead, Openprise, and DemandTools.
Step 1 — Gather context
If references/learnings.md exists, read it first for accumulated knowledge.
Ask the user:
1. What's the main data problem?
- A) Duplicate contacts, leads, or accounts
- B) Stale/outdated records (job changes, company changes)
- C) Missing fields (no phone, no email, incomplete company data)
- D) Inconsistent data (job titles, industries, company names formatted differently)
- E) Compliance issues (opt-outs, GDPR, stale consent)
- F) General data audit — don't know what's wrong yet
- G) Setting up ongoing data hygiene automation
- H) Other — describe it
2. What CRM are you using?
- A) Salesforce
- B) HubSpot
- C) Microsoft Dynamics
- D) Pipedrive
- E) Other CRM
- F) Custom/in-house system
3. How many records are affected?
- A) Under 1,000 (small cleanup)
- B) 1,000-10,000 (moderate)
- C) 10,000-100,000 (large)
- D) 100,000+ (enterprise-scale)
- E) Not sure — need to audit first
4. What tools do you have for data operations?
- A) ZoomInfo OperationsOS
- B) HubSpot Operations Hub
- C) Salesforce native (duplicate management, data.com)
- D) Clay
- E) LeanData / RingLead / Openprise
- F) DemandTools (Validity)
- G) None — using manual processes
- H) Other — describe it
Step 2 — Strategy and approach
Read `references/platform-guide.md` for detailed audit frameworks, deduplication strategies, normalization tables, enrichment automation, and platform-specific guidance.
You no longer need the platform guide details — focus on the user's specific situation.
Step 3 — Actionable guidance
Quick wins (do these first)
1. Remove obvious duplicates — exact email match dedup is safe and fast 2. Fix formatting — standardize phone numbers, capitalize names, normalize countries 3. Fill critical gaps — bulk enrich records missing email or phone 4. Remove dead records — hard bounces, invalid emails, disconnected phones
Ongoing hygiene program
1. Prevent duplicates at entry — enable duplicate rules on record creation 2. Enrich on create — auto-enrich new records within minutes of creation 3. Monthly dedup sweep — run fuzzy match dedup monthly, review and merge 4. Quarterly refresh — re-enrich all active records every 90 days 5. Annual purge — remove records with no activity in 12+ months (archive, don't delete)
Metrics to track
- Duplicate rate — % of records with duplicates (target: <2%)
- Field completeness — % of critical fields filled (target: 95%+)
- Bounce rate — email bounce rate on outbound (target: <3%)
- Data age — median days since last enrichment (target: <90)
- Merge rate — duplicates merged per month (should trend down over time)
Gotchas
1. Merge before you enrich — enriching duplicate records wastes credits. Dedup first, then enrich the surviving records.
2. Test dedup rules on a sample first — fuzzy matching can produce false positives (merging records that shouldn't be merged). Always review a sample of 50-100 merge candidates before running bulk operations.
3. Preserve lead source on merge — the most common post-merge complaint is losing original lead source attribution. Configure merge rules to keep the oldest record's lead source.
4. Don't delete — archive — instead of deleting stale records, move them to an archive status. Deleted records lose history; archived records can be reactivated if the contact returns.
5. GDPR and compliance — data hygiene must respect opt-out and consent records. Never re-enrich a contact who has opted out. Check compliance status before any bulk enrichment operation.
- Self-improving: If you discover something not covered here, append it to
references/learnings.mdwith today's date.
Before recommending a specific platform skill
This skill covers a strategy domain across many platforms. Before pointing the user to any specific platform skill (any /sales-{platform} listed in ## Related skills, e.g., /sales-mailshake, /sales-klaviyo, /sales-apollo), read that platform skill's actual SKILL.md first. The 1-line description in ## Related skills is enough to identify a candidate — it's not enough to commit to it or to write a prompt that invokes it well.
How to read it:
- If
~/.claude/skills/{skill-name}/SKILL.mdexists locally,Readit. - For
sales-*skills,WebFetchdirectly from this repo:https://raw.githubusercontent.com/sales-skills/sales/main/skills/{skill-name}/SKILL.md— e.g., forsales-mailshake:https://raw.githubusercontent.com/sales-skills/sales/main/skills/sales-mailshake/SKILL.md. - For non-
sales-*skills (third-party), look up{org}/{repo}in~/.claude/skills/sales-do/references/skill-sources.mdif installed and fetch the sameskills/{skill-name}/SKILL.mdpath under that repo.
After reading, ground your recommendation in something concrete from the SKILL.md (its scope, a sub-flow, its argument-hint shape, or a "Do NOT use for..." negative trigger). Align any generated invocation with the platform skill's argument-hint. If the platform skill turns out not to fit the user's situation, swap to another or handle the question here directly rather than recommending a poor fit.
Related skills
/sales-hubspot— HubSpot platform help (Data Hub data sync, data quality automation, deduplication)/sales-attio— Attio platform help (AI-native CRM with custom objects)/sales-blueconic— BlueConic CDP — profile unification, identity resolution, audience activation/sales-tealium— Tealium CDP — Real-Time CDP, identity resolution, 1300+ connectors/sales-cdp— CDP comparison and selection strategy across platforms/sales-clay— Clay platform help/sales-zoominfo— ZoomInfo platform help (for OperationsOS-specific setup)/sales-clearbit— Clearbit platform help (enrichment, reveal, prospector)/sales-enrich— enrichment strategy across all providers/sales-lead-routing— lead assignment and territory rules (often paired with dedup)/sales-lead-score— lead scoring models (depend on clean data)/sales-integration— connecting data tools to CRM/sales-prospect-list— building prospect lists (data quality at the source)/sales-do— Not sure which skill to use? The router matches any sales objective to the right skill. Install:npx skills add sales-skills/sales --skill sales-do
Examples
Example 1: CRM data audit
User says: "Our Salesforce has 50,000 contacts and I suspect a lot of them are duplicates or outdated. Where do I start?" Skill does: Walks through the data quality audit framework — measure completeness, accuracy, duplication rate, and decay. Recommends starting with exact-match email dedup (safest), then running a field completeness report, then sampling 100 records against LinkedIn to estimate accuracy. Result: User has a data quality scorecard and prioritized cleanup plan.
Example 2: Setting up ongoing hygiene
User says: "We keep getting duplicates in HubSpot and our data goes stale within months. How do we automate this?" Skill does: Recommends HubSpot Operations Hub for dedup + ZoomInfo or Clay for enrichment. Sets up duplicate prevention rules on creation, auto-enrichment for new records, and a quarterly re-enrichment schedule. Result: User has an automated hygiene program that prevents duplicates and keeps data fresh.
Example 3: Pre-campaign data cleanup
User says: "We're about to launch a big outbound campaign to 10,000 contacts. How do I make sure the data is clean first?" Skill does: Recommends a pre-campaign checklist: dedup the list, verify emails with a dedicated verification tool, re-enrich records older than 90 days, remove contacts at companies that no longer fit ICP, and check opt-out/DNC status. Result: User launches campaign with verified, deduplicated, compliant data — lower bounce rate, higher deliverability.
Troubleshooting
Dedup merging wrong records
Symptom: Fuzzy match dedup merged two different people who happen to have similar names at the same company Cause: Match rules too loose — matching on name + company without additional criteria Solution: Tighten match rules: require email OR phone match in addition to name + company. Always run in "review" mode before "auto-merge" mode. Add title or department as a tiebreaker.
Enrichment not filling expected fields
Symptom: Auto-enrichment runs but many records still have empty phone or email fields Cause: Single enrichment provider doesn't have coverage for all contacts. Coverage varies by geography, seniority, and industry. Solution: Implement waterfall enrichment — try Provider A, if no result try Provider B, then Provider C. Use /sales-enrich for waterfall setup. Common waterfall: ZoomInfo → Apollo → Lusha.
Data quality metrics not improving
Symptom: Running monthly dedup and enrichment but duplicate rate and completeness aren't improving Cause: New duplicates are being created faster than they're being merged. Root cause is usually web forms, imports, or integrations creating records without duplicate checks. Solution: Fix the source — enable duplicate prevention rules on all record creation paths (web forms, API imports, manual creation, integration syncs). Prevention is more effective than cleanup.
CRM Data Hygiene & Quality Learnings
Accumulated tips, gotchas, and corrections discovered during use. Claude reads this at the start of each invocation and appends new learnings as they're discovered. Once significant learnings have accumulated, use /sales-request-skill to share them back to the community. Shared and declined entries are marked so they won't be re-prompted.
<!-- Add entries below in format: YYYY-MM-DD: Learning description -->
CRM Data Hygiene Platform Guide
Data Quality Audit Framework
Before fixing data, measure what's broken:
1. Completeness — what % of records have all critical fields filled?
- Email: target 95%+ for contacts
- Phone: target 70%+ for key personas
- Company: target 100% for accounts
- Title/department: target 90%+ for contacts
2. Accuracy — what % of filled fields are actually correct?
- Email deliverability: verify a sample with an email verification tool
- Phone connectivity: check a sample of direct dials
- Job title currency: compare against LinkedIn for a sample
3. Duplication rate — what % of records are duplicates?
- Contact-level: same person, multiple records
- Account-level: same company, different spellings
- Cross-object: leads that are also contacts
4. Decay rate — how fast does your data go stale?
- Industry average: 30% of B2B data decays annually
- Sales contacts: ~20% change jobs each year
- Direct dials: ~15% become invalid annually
- Emails: ~22% bounce rate after 12 months without refresh
5. Consistency — are the same things called the same thing?
- Job titles: "VP Sales" vs "Vice President of Sales" vs "VP, Sales"
- Industries: "SaaS" vs "Software" vs "Technology"
- Company names: "IBM" vs "International Business Machines" vs "IBM Corp"
Deduplication Strategy
| Approach | When to use | Risk level |
|---|---|---|
| Exact match | Email, phone, domain — safest | Low |
| Fuzzy match | Names, company names, addresses | Medium — review matches before merging |
| Rule-based | Combine multiple fields (name + company + title) | Medium |
| ML-based | Large datasets with complex patterns | Low (if trained well) — but expensive |
Merge rules (which record wins):
- Most recently updated record keeps modifiable fields
- Most complete record keeps enrichment data
- Oldest record keeps the original owner/source
- Always preserve: original lead source, first touch date, opt-in status
Data Normalization
| Field | Common problems | Solution |
|---|---|---|
| Job title | Abbreviations, variations, custom titles | Map to standard taxonomy (C-Level, VP, Director, Manager, IC) |
| Industry | Free-text, overlapping categories | Map to SIC/NAICS or your internal taxonomy |
| Company name | Abbreviations, legal suffixes, DBA names | Normalize to official name, store variants as aliases |
| Phone | Mixed formats, extensions, country codes | E.164 format (+1XXXXXXXXXX) |
| Address | Inconsistent formatting, abbreviations | USPS standardization or Google Maps API |
| Country | Mix of codes and names | ISO 3166-1 alpha-2 codes |
Enrichment Automation
Set up ongoing enrichment to prevent decay:
1. Trigger-based — enrich when a record is created or updated 2. Scheduled — monthly/quarterly batch enrichment of all records 3. Decay-based — re-enrich records older than X days 4. Event-based — re-enrich when a contact's company has a news event (funding, acquisition)
Platform-Specific Guidance
In ZoomInfo (OperationsOS)
ZoomInfo OperationsOS is purpose-built for CRM data management at scale.
Deduplication:
- OperationsOS identifies duplicates across contacts, leads, and accounts using fuzzy matching
- Configurable match rules: email, name+company, phone, domain
- Bulk merge with configurable "winning record" rules
- Cross-object dedup: find leads that already exist as contacts
Data orchestration:
- Build automated workflows: new record -> match existing -> enrich -> normalize -> route
- Configure matching rules to prevent duplicates before they're created
- Set up enrichment triggers on record creation, field change, or scheduled interval
- Normalization rules for job titles, industries, company names
Data decay management:
- Auto-detect job changes and company changes
- Flag stale records based on last-enriched date
- Configure re-enrichment schedules (monthly recommended for active pipeline)
- Track data quality metrics over time
Setup: ZoomInfo admin -> OperationsOS -> Data Orchestration -> Create Workflow.
In Salesforce (Native)
Duplicate Management:
- Setup -> Duplicate Management -> Duplicate Rules
- Standard rules: match on email, name+company, phone
- Custom matching rules for complex scenarios
- Block or alert on duplicate creation
Data.com Clean (if licensed):
- Batch cleaning of contacts and accounts
- Auto-enrichment on record creation
- Scheduled cleanups
Limitations: Native dedup is basic — no fuzzy matching, no cross-object dedup, no automated merge. For enterprise-scale, pair with DemandTools or ZoomInfo OperationsOS.
In HubSpot (Operations Hub)
Deduplication:
- Operations Hub includes AI-powered duplicate detection
- Suggests merge candidates with confidence scores
- Bulk merge with "primary record" selection
- Available on Operations Hub Professional+
Data quality automation:
- Programmable automation (Operations Hub Professional): custom code actions for normalization
- Data quality command center: monitor property completeness, formatting issues
- Automated formatting: capitalize names, standardize phone numbers, clean URLs
Limitations: HubSpot dedup is contact/company only — no custom object dedup. Formatting automation requires Operations Hub Professional ($800/mo+).
In Clay
- CRM enrichment & refresh: Import contacts/accounts from Salesforce, HubSpot, or Dynamics 365 into Clay tables. Run waterfall enrichment to fill missing fields and refresh stale data. Push updated records back via bidirectional sync.
- Automated data maintenance: Set up scheduled imports to regularly pull CRM records into Clay, re-enrich, and sync back. Keeps contact data fresh without manual effort.
- Duplicate detection: Use enrichment data (LinkedIn URLs, company domains, verified emails) to identify and flag duplicate records before syncing back to CRM.
- Data standardization: Use Sculptor workflows to normalize job titles, company names, industry classifications, and other fields. Apply consistent formatting before pushing to CRM.
- Plan gate: CRM sync requires Growth plan ($446-495/mo). Free/Launch users can export enriched data as CSV for manual CRM import.
- Best for: RevOps teams wanting to automate CRM data enrichment and cleanup on a recurring basis.
In LeanData / RingLead
LeanData:
- Lead-to-account matching (which leads belong to which accounts)
- Lead routing based on matching results
- Deduplication with merge automation
- Salesforce-native (runs inside SFDC)
RingLead (now ZoomInfo — acquired):
- Duplicate prevention on record creation
- Bulk dedup with configurable match rules
- Data normalization and standardization
- Works with Salesforce, HubSpot, Marketo
In DemandTools (Validity)
- Most powerful Salesforce dedup tool
- Scenario-based: build complex match/merge rules
- Mass operations: update, delete, deduplicate at scale
- Import management: clean data before it enters CRM
- Single Table Dedupe, Table-to-Table Dedupe (cross-object)
In Clearbit
Clearbit (now Breeze Intelligence in HubSpot) focuses on enrichment-driven data hygiene — standardizing and filling CRM records with firmographic and demographic data.
Enrichment-based cleanup:
- Enrich existing CRM records with standardized firmographic/demographic data
- Bulk enrichment via API or Breeze Intelligence in HubSpot
- Continuous data refresh: subscribe to enrichment updates when data changes (
subscribe: trueparameter) — records stay current without manual re-enrichment
Normalization:
- Standardized industry codes (NAICS, GICS, SIC) — normalizes messy free-text industry fields
- Normalized role and seniority classifications — standardizes job titles across records into consistent categories
- Corporate hierarchy: parent company and ultimate parent domain fields help deduplicate subsidiary records
Data quality signals:
emailProviderflag identifies personal email addresses (gmail, yahoo) vs business emails — useful for filtering low-quality records- Tech stack data helps identify outdated technology fields in CRM
Best for: Filling missing fields, standardizing industries and titles, flagging personal emails, and keeping records fresh with continuous enrichment. Pair with a dedup tool (ZoomInfo, DemandTools) for full hygiene coverage.
In Attio
Attio's flexible data model makes hygiene both easier (custom attributes, no rigid schema) and harder (more objects to maintain).
Built-in hygiene features:
- Auto-enrichment fills contact and company fields from email/domain data on record creation
- AI Research Agent (Pro plan) enriches records from public web sources
- Merge duplicates via UI — select records and merge with field-level control
- Record references enforce relationships between objects
Custom data model considerations:
- Custom objects multiply hygiene surface area — each object needs its own dedup and completeness rules
- No built-in formula fields — computed hygiene scores (completeness %, last activity) require automations or external tools
- Attribute types enforce data formats (email, phone, domain) — use these over free-text to prevent format drift
Automation-based hygiene:
- Set up automations to flag records missing critical fields (e.g., no email, no company)
- Use "record updated" triggers to normalize data on entry (e.g., standardize country names)
- Schedule periodic enrichment re-runs via API for stale records
Best for: Teams under 50 using Attio as primary CRM. For enterprise-scale hygiene with advanced dedup rules, ZoomInfo OperationsOS or DemandTools are more powerful.
In Treasure Data (CDP)
Treasure Data approaches data hygiene from a CDP perspective — unifying messy data from 400+ sources into clean, deduplicated customer profiles.
Identity resolution as dedup:
- Parent table defines the unified profile schema. Child tables (website events, CRM, POS, email) feed into it.
- Identity resolution rules match records across sources: deterministic (email, phone, customer ID) and probabilistic (device IDs, cookies).
- Source priority order resolves conflicts — if CRM says "John" and email says "Jonathan", the higher-priority source wins.
- Test unification in QA sandbox before production — bad identity rules create false merges that are hard to undo.
Data quality at ingestion:
- Schema-flexible ingestion means bad data gets in easily. Define validation rules in Treasure Workflows to catch issues before they reach the parent table.
- Postback API is case-sensitive — column names must match the target table schema exactly (including casing). Mismatched casing silently drops data.
- Monitor profile count growth — overly loose identity rules create too many profiles (inflating "No Compute" pricing costs).
Ongoing maintenance:
- Schedule re-unification jobs to pick up new identity signals as more data arrives.
- Use Treasure Workflows (DAGs) to orchestrate: ingest → validate → unify → segment. Don't manage individual job schedules.
- Audit source priority order quarterly — data source quality changes over time.
Best for: Enterprise B2C companies with 400+ data sources that need centralized identity resolution across channels. Requires SQL knowledge for advanced transformations. Implementation takes 8-12 weeks with professional services ($30K-$100K+).
In Tealium (CDP)
Tealium approaches data hygiene through real-time identity resolution in AudienceStream CDP, unifying profiles from 1,300+ sources.
Identity resolution as dedup:
- Visitor Switching merges profiles when a shared identifier (email, customer ID, phone) matches across devices/channels.
- Profile merge rules are configurable — deterministic matching (email, login ID) is safer than probabilistic (device fingerprints, cookies).
- Test merge rules in QA environment before production — bad rules create false merges that blend distinct customers.
Data quality at ingestion:
- EventStream connectors can validate and transform data before it reaches AudienceStream profiles.
- The Collect HTTP API is case-sensitive for all field names — mismatched casing silently drops data or creates duplicate attributes.
- Event-based pricing means dirty data (duplicate events, bot traffic) inflates costs. Filter at the source.
Ongoing maintenance:
- Monitor profile count growth — overly loose identity rules create too many profiles (inflating event-based pricing).
- Use AudienceStream enrichments to flag data quality issues (e.g., badge for "missing email", metric for "days since last event").
- Connector error rates surface data quality problems downstream — if a connector consistently fails, the source data may be malformed.
Best for: Enterprise organizations with 1,300+ integration touchpoints needing centralized real-time identity resolution. Marketer-friendly AudienceStream UI for non-technical teams. Implementation takes 4-12 weeks with professional services.
In Cognism (CRM Enrichment)
Cognism approaches data hygiene through automated CRM enrichment — refreshing stale records with verified contact and company data.
Automated record refresh:
- Connect Salesforce (Professional+) or HubSpot (2-way sync) to Cognism for ongoing enrichment.
- Director-level data is refreshed every 30 days — keeps job titles, companies, and phone numbers current.
- Job change detection (Elevate plan) catches contacts who've moved companies — prevents outreach to people who've left.
Duplicate prevention:
- Import existing CRM contacts into Cognism to enable exclusion filters — prevents creating duplicate records when prospecting.
- HubSpot integration imports Contacts, Companies, and Tasks for matching.
- Set up exclusion lists for customers, competitors, and opted-out contacts.
Data decay management:
- B2B data decays at ~30% annually. Cognism's scheduled enrichment catches changes faster than manual audits.
- Diamond Data phone verification (Elevate plan) adds human-verified mobile numbers — reduces the "wrong number" problem that plagues CRM phone fields.
- Email addresses may be pattern-generated — run periodic validation through ZeroBounce or SafetyMails to catch invalid emails before they bounce.
Fair-use cap consideration:
- ~2,000 records/user/month under "unlimited" plans. For CRMs with 50K+ contacts, enrichment must be batched across multiple users or months.
- Prioritize enrichment on active pipeline contacts and recently engaged leads — don't waste fair-use allocation on cold, inactive records.
Best for: Mid-market to enterprise teams with EMEA-focused CRMs that need ongoing phone number and contact verification. Cognism's Diamond Data is uniquely valuable for keeping EMEA phone fields accurate. For US-heavy CRMs, ZoomInfo OperationsOS may be stronger.
In People.ai (Backstory)
People.ai approaches data hygiene by eliminating the root cause of stale CRM data: manual rep entry. Instead of cleaning bad data after the fact, it prevents bad data from entering.
Automatic activity capture:
- Captures every email, call, meeting, and chat message and auto-associates with CRM contacts, accounts, and opportunities.
- No rep action required — activities are logged server-side from Gmail/Outlook/Zoom/Teams/Slack.
- Eliminates the "garbage in" problem: CRM data reflects actual activity, not stale rep estimates.
- Analyzes 2 years of historical data on day one — fills in activity gaps retroactively.
Contact creation and association:
- Auto-creates contacts in CRM from email/meeting participants when they don't already exist.
- Associates activities with the correct opportunity based on participant and timing signals.
- Multi-CRM support (Salesforce, Dynamics, Oracle) — maintains data quality across CRM instances.
Pipeline data accuracy:
- Deal intelligence signals (engagement scoring, single-threading detection) surface data quality issues at the deal level.
- Stakeholder mapping identifies which contacts are missing from opportunities — a form of data completeness checking.
Limitations:
- People.ai captures activity metadata, not call content. It won't populate methodology fields (MEDDPICC/BANT) from calls — for that, use Gong, Sybill, or Scratchpad.
- Contact matching depends on email addresses: if CRM contacts have outdated emails, activities won't associate correctly. Run an email verification pass first.
- Enterprise-only pricing. No free tier or self-serve. Budget $50-100+/user/month.
- Activity processing can take 24-48 hours for call data — not real-time.
Best for: Enterprise teams where CRM data decay is primarily caused by reps not logging activities. People.ai solves the input problem — once activities flow automatically, downstream data hygiene issues (stale deals, missing contacts, incomplete opportunity records) decrease dramatically. For CRM field enrichment (job titles, phone numbers, company data), pair with ZoomInfo or Cognism.
Related skills
How it compares
Audit framework and operational CRM hygiene—not a single-vendor enrichment API skill.
FAQ
Who is sales-data-hygiene for?
SaaS founders and small sales motions who own their CRM data and need agent-guided quality measurement before scaling touches.
When should I use sales-data-hygiene?
During Grow lifecycle work before major outbound, after importing lists, when win rates slip, or quarterly to re-baseline duplication and decay.
Is sales-data-hygiene safe to install?
It may suggest reading or updating CRM records and external verification samples; review the Security Audits panel on this Prism page and avoid pasting production credentials into chat.