How-to guide

How to clean CRM data at scale

Dirty CRM data is the single most expensive problem revenue teams quietly tolerate. Reps waste hours chasing stale phone numbers, forecasts rot because duplicate accounts split pipeline, and marketing targets ghosts. This guide walks you through a repeatable program: profile the mess, dedupe the obvious collisions, enrich the thin records, archive the dead weight, standardize what remains, and lock the whole thing behind guardrails so you never have to run a six-week emergency cleanup again.

Before you start

What you need.

Time: 2 hours to plan, 2 weeks to execute the first full pass

  • Admin access to your CRM so you can edit records, run merges, and change validation rules
  • Export permissions on accounts, contacts, leads, and opportunities for baseline profiling
  • A documented list of required fields today, even if nobody follows it, so you know what you\'re measuring against
  • Access to a data provider or enrichment source your team already pays for (Clearbit, ZoomInfo, Apollo, or equivalent)
  • Buy-in from sales leadership that a two-week cleanup is worth two weeks of lower-priority prospecting
  • A sandbox environment or at least a dry-run habit; never run bulk merge or bulk delete directly against production first
Clean CRM data at scale across accounts, contacts, and deals

Step by step.

  1. 1

    Profile the data before you touch anything

    You cannot clean what you have not measured. Before the first merge, export every account, contact, lead, and opportunity to CSV and run basic profiling queries. For each object, measure: percentage of records with the required fields populated, duplicate rate by exact-match email and by fuzzy-match company name, percentage with no activity in the last ninety days, and percentage missing a related record (contacts with no account, deals with no primary contact, accounts with zero contacts). Write these numbers down. They are your baseline. Without them, you will finish the cleanup and have no way to prove it worked, which means leadership will quietly stop funding the work and the problem will come back in six months. Profile across the full dataset, not a sample. Fuzzy duplicates cluster in the long tail and sampling hides them.

    • Export accounts, contacts, leads, and opportunities to CSV in full, not sampled
    • Measure required-field completeness per object as a percentage, not a count
    • Measure exact-email duplicate rate and fuzzy-company-name duplicate rate separately
    • Measure the percentage of each object with zero activity in the last ninety days
    • Measure orphan rates: contacts without accounts, deals without primary contacts, accounts without contacts
    • Save the baseline numbers in a shared doc and date them; you will compare against these after each phase
    Tip: If your profiling shows duplicate rates above fifteen percent on accounts, do not start with enrichment. Dedupe first; enriching duplicates just makes two polished copies of the same thing.
  2. 2

    Define what a duplicate means before you write a merge rule

    The word duplicate hides more disagreement than any other term in CRM work. Before you build matching rules, decide and document what constitutes a duplicate in each object. For accounts, the usual rule is same legal entity, which means same domain for software companies, same legal name plus state for services firms, and same parent company plus billing country for multinationals. For contacts, the rule is usually same person, which means matching work email, or matching personal email plus phone, or matching first name plus last name plus current employer. For leads, you have an extra decision: do leads with the same email as an existing contact count as duplicates, or do they get merged into the contact. Pick one answer and stick to it. Write the rules down in plain English before translating them into match criteria. Rules that live only in your head produce merges nobody can explain six months later.

    • Write the account duplicate rule in one sentence, then enumerate the exceptions
    • Write the contact duplicate rule the same way; name whether personal email counts
    • Decide the lead-to-contact collision policy: merge into contact, or hold as separate lead until converted
    • Document the rules in your admin wiki and reference them in the merge job name
    Tip: Resist the urge to use company name alone as an account match key. "Acme" matches 400 real companies. Domain plus country is almost always safer.
  3. 3

    Dedupe accounts first, then contacts, then leads

    Order matters. Dedupe accounts first because every contact and deal points at one, and merging accounts automatically repoints the children. If you dedupe contacts before accounts, you will merge two contacts that each point at a different account, then merge the accounts, and the second merge will produce new conflicts you have to resolve by hand. Run the account merge as a two-phase job. Phase one: exact-domain matches. These are safe, high-confidence, and typically account for sixty to eighty percent of the duplicate volume. Automate the merge, pick the record with the most recent activity as the winner, and preserve all notes, activities, and related records from the loser. Phase two: fuzzy matches on company name, same billing country, created more than thirty days apart. Review these in a merge queue with a human in the loop. Never bulk-merge fuzzy matches without review; the false-positive rate is high enough to corrupt real customer relationships. After accounts, dedupe contacts using the rules you wrote in step two, then do leads last.

    • Phase one accounts: exact-domain merge, automated, winner is most-recent-activity record
    • Phase two accounts: fuzzy-name merge, human-reviewed in a queue, batched at fifty per sitting
    • Dedupe contacts after accounts settle; email-exact is safe to automate, personal-email-plus-phone needs review
    • Dedupe leads last; if your policy is to merge leads into existing contacts, run that sweep now
    • Log every merge in a cleanup audit table so you can roll back a specific merge if a rep flags it
    Tip: When you merge, the record with the most recent activity usually wins, not the oldest record. Oldest records tend to have the most stale data; newest records have the most context.
  4. 4

    Enrich the records you kept, not the ones you lost

    Now that the record count is smaller and cleaner, run enrichment. The economics here matter. Most enrichment providers charge per record, so enriching before dedupe wastes thirty percent or more of your spend on records that will be merged away. For accounts, enrich firmographics: industry, employee count, revenue band, HQ country, and tech stack if that signal matters to your motion. For contacts, enrich title normalization, seniority, department, and verified work email. Pick a single source of truth per field; if ZoomInfo and Clearbit disagree on employee count, decide which one wins before the sync runs, do not let the last write win. Set enrichment to run on create and then on a quarterly refresh for stable fields and monthly for volatile fields like job titles. Do not enrich every field every day; it costs money, hits rate limits, and makes change tracking impossible.

    • List the firmographic and demographic fields that actually drive segmentation or routing; enrich only those
    • Pick a single provider as source of truth per field; do not allow two providers to write the same field
    • Set enrichment cadence: on create, quarterly refresh for stable fields, monthly for job titles and employee counts
    • Flag records where enrichment returned low confidence; those go to a review queue, not into production
    • Reconcile billing monthly to catch provider sync loops that burn credits on the same record daily
  5. 5

    Archive the dead weight on a documented policy

    Not every record deserves to live forever. The hardest conversation in CRM cleanup is agreeing to archive records that reps emotionally refuse to delete. Build a written archival policy and get leadership signoff before you archive a single record. A reasonable baseline: accounts with no activity in twenty-four months, no open opportunities, and no active contacts move to an archived state. Contacts with bounced email, no activity in eighteen months, and no related open deal move to archived. Leads older than twelve months with no response move to a nurture pool outside the primary lead object or get archived outright. Archive does not mean delete. Keep the records readable, exclude them from default list views and reports, and preserve them for compliance and historical reporting. Deletion is only for records you are legally required to remove, such as GDPR subject access requests or hard-bounce suppressions. Everything else is archive, not delete.

    • Write the archival rules per object, including every exception leadership wants to carve out
    • Set a status field (active, archived, suppressed) rather than deleting; archive is reversible, delete is not
    • Exclude archived records from default views, routing, and reports; keep them visible on direct lookup
    • Review archival volumes monthly; if archive rates spike, something upstream changed and you need to find it
    • Keep a six-month unarchive window during which a rep can reactivate with a documented reason
    Tip: Reps will fight archival. Pre-empt the fight by promising unarchive in one click and showing them that archived records still appear in global search. The emotional objection is loss of control, not loss of data.
  6. 6

    Standardize the fields that drive reporting

    Standardization is the step most teams skip, and it is the reason most CRM dashboards lie. Every picklist, every country code, every industry label, every stage name needs one spelling and one value. Pick a controlled vocabulary and enforce it. Industry should match a published taxonomy like NAICS or SIC rather than free text; otherwise you will have Software, Software and Technology, SaaS, Software as a Service, and Technology Services all competing in the same report. Country codes should match ISO 3166, not United States, USA, US, and America fighting for the same field. Job titles need a seniority mapping (IC, Manager, Director, VP, C-level) that strips the eight-hundred unique strings reps type into a title field down to a reportable set. Standardization happens in three places: the picklist definition, a nightly normalization job that cleans the long tail of free-text drift, and a validation rule that blocks new drift at creation. If you skip any of the three, drift wins over a quarter.

    • Convert industry, country, state, and seniority to picklists tied to published taxonomies
    • Build a nightly job that normalizes legacy free-text values to the controlled vocabulary
    • Add validation rules that block new records from using off-list values
    • Set a quarterly review to add or retire picklist values based on actual sales patterns, not committee opinion
    • Separate display label from stored value so you can rename labels without breaking historical reports
  7. 7

    Fix the orphans and the broken relationships

    Clean records that point at nothing are almost as useless as duplicates. Walk every orphan class your profiling flagged in step one and resolve them. Contacts without an account: match to an account by email domain where possible; if no match, create a stub account named after the domain and flag it for enrichment. Opportunities without a primary contact: this one is a forecasting hazard because you cannot route communications or run close-plan reports; assign the deal creator as the temporary primary contact and surface the gap in a weekly rep view. Accounts without contacts: if the account has an open deal, flag it urgently; if not, archive it under the step-five policy. Opportunities with closed dates in the past but still-open status: these are the single most common forecast lie; close them with a loss reason of time-expired or push the close date with a required justification note. Fix orphans before you declare the cleanup done. They are the records most likely to break downstream reports once everything else looks tidy.

    • Match orphan contacts to accounts by email domain; stub-create accounts for unmatched contacts
    • Assign a temporary primary contact to every opportunity missing one and surface the gap for the rep
    • Flag accounts without contacts; open-deal accounts are urgent, no-deal accounts follow the archival policy
    • Close or justify every opportunity whose close date is in the past; this is the single most common forecast corruption
    • Rebuild the parent-child relationships on any account merged in step three where a child was accidentally detached
  8. 8

    Install the guardrails so the mess does not come back

    A one-time cleanup is a waste of two weeks if you do not install the guardrails that keep the data clean. Four guardrails matter. First, dedupe at creation: your CRM should check for existing matches when a rep or import tries to create a new account or contact and either block the create or offer to merge. Second, required-field enforcement at stage transitions: a deal cannot move from stage two to stage three without the fields that stage three reports require. Third, decay flags: a daily job that flags any record that has drifted from the controlled vocabulary, any record that lost a required field, and any record that lost its parent relationship. Fourth, a weekly data health review: a thirty-minute meeting with revenue operations and a sales manager where you review the decay dashboard, approve the merges, and sign off on the archivals. If any of the four guardrails is missing, drift will return within a quarter. If all four are in place, you can run the full cleanup pass once a year instead of once a quarter and spend the saved time on actual revenue work.

    • Enable dedupe-at-creation on accounts, contacts, and leads with the matching rules from step two
    • Attach required-field validation to every stage transition in your pipeline
    • Build a decay dashboard that counts drift, missing fields, and orphan creation per week
    • Schedule a thirty-minute weekly data health review with revops and a sales manager; make it recurring
    • Measure the baseline numbers from step one on a monthly recurring cadence; publish the trend line to leadership
    Tip: Guardrails that only flag are weaker than guardrails that block. If leadership will not approve blocking rules, start with daily email reports to managers listing their reps\' violations. Social accountability closes the gap faster than you expect.
Avoid

Common mistakes.

  • Running enrichment before deduplication. You pay the enrichment vendor for both halves of every duplicate pair and still have to merge them afterward, which wastes thirty to forty percent of your enrichment budget on records that will not survive the week.
  • Using company name alone as the account match key. The word Acme matches four hundred real companies in most datasets. Match on domain plus billing country for software firms, and legal name plus state for services firms.
  • Bulk-merging fuzzy matches without a human review step. Fuzzy matching has a false-positive rate in the low single digits, which sounds small until you have merged two real customers into a single corrupted record and a rep notices three weeks later.
  • Treating archive and delete as synonyms. Deletion loses history, breaks referential integrity in downstream warehouses, and triggers compliance incidents when a GDPR request later asks what you held. Archive is reversible and safe. Delete is not.
  • Skipping the guardrails step because the cleanup is done. The cleanup is never done. Without dedupe-at-creation, required-field enforcement, a decay dashboard, and a weekly review, the dataset will regress to its starting state within two quarters, guaranteed.
  • Letting two enrichment providers write the same field. Last-write-wins produces fields that silently flip between ZoomInfo and Clearbit values day to day, which corrupts time-series reporting. Pick one source per field and lock the other out.
FAQ

Frequently asked questions.

How long does a full CRM data cleanup actually take?

The first pass takes two to three weeks of focused work for a dataset under one hundred thousand records, and four to six weeks for anything larger. Profiling runs in a day. Dedupe is the longest phase because fuzzy matches need human review. Enrichment runs on vendor timelines. Standardization and guardrails together add about a week. After the first pass, if you install the guardrails properly, subsequent passes are a half-day monthly health check rather than a multi-week project.

Should I clean the data in place or move to a new CRM?

Clean the data in place first, even if you are planning a migration. Moving dirty data into a new CRM gives you a new CRM full of dirty data, plus a migration bill. The one exception is when the current CRM lacks the tools to run bulk merge, dedupe-at-creation, or required-field validation at stage transitions. In that case, the migration itself becomes the forcing function for cleanup, and you build the guardrails as part of the move.

How often should I rerun the full cleanup?

If your guardrails are installed and working, once a year. The weekly data health review catches the small drift, the monthly baseline measurement catches the medium drift, and the annual pass catches the structural drift (new product lines, new segments, new picklist values that reflect real business change). Teams that run full cleanups quarterly almost always have missing or weak guardrails; the cleanup is a symptom, not a solution.

What is the right duplicate rate to target?

Under three percent on accounts, under five percent on contacts, under seven percent on leads. Zero is not a realistic target because new records always arrive before dedupe-at-creation can match them, and because mergers, acquisitions, and name changes create legitimate near-duplicates that take time to resolve. If your rates are above those thresholds a month after cleanup, your dedupe-at-creation rules are either disabled or too narrow.

Who owns CRM data quality: sales, marketing, or revenue operations?

Revenue operations owns the system, the rules, and the dashboards. Sales managers own the behavior and the enforcement. Marketing owns the inbound hygiene (form validation, enrichment at capture, suppression lists). None of the three can own the whole thing alone. The weekly data health review from step eight is where the three functions meet and agree on priorities, which is why skipping that meeting is the fastest way to let the dataset decay.

What should I do with lead records that have never responded?

Move them to a documented nurture pool with a lower-cost touch cadence (newsletter, educational content) for twelve months. If still no engagement at the twelve-month mark, archive under the step-five policy. Do not keep eighteen-month-old no-response leads in the primary routing pool; they distort coverage math, waste rep time on prospecting that will not land, and inflate database costs with records that produce no revenue.

See it in Strkr

Related product surfaces.

Strkr CRM All features

Run cleanups on a cadence, not as a crisis

Strkr gives you dedupe-at-creation, required-field validation at every stage transition, a decay dashboard, and a weekly data health view out of the box. Clean once, guardrail forever, forecast on data you trust. Strkr AI surfaces the records most likely to drift before a rep ever notices, and Strkr Messaging keeps outreach pointed at the verified contact, not the stale one.

Try it free. Bring your team next week.

No sales call, no migration consultant, no four-month implementation. Enter your card, get 14 days of the full Pro tier, cancel any time before day 14 with zero charge. Spin up a workspace, import your CSV, and have something useful before lunch.