Answers

What is data dedupe?

Every CRM accumulates duplicates. Dedupe is the detection and triage layer that catches them. Merge is the resolution step that fuses two rows into one. The two are related but not interchangeable.

Short answer

Data dedupe is the process of finding and merging duplicate records (contacts, accounts, and leads) in a CRM so each real-world person or company exists exactly once. The job runs match strategies like exact email, domain plus fuzzy name, phone normalization, and Levenshtein distance across every incoming record and every existing row. Dedupe prevents double-touch outreach, broken routing, and messy reporting, and is a core RevOps hygiene task separate from the merge action that actually resolves the duplicate.

Key points

What matters most.

The six ideas every RevOps team should know before standing up a dedupe program: what counts as a duplicate, which match strategies catch which cases, and where dedupe stops and merge begins.

Definition

Finding the same person or company twice.

Data dedupe is the process of identifying records that describe the same real-world entity, even when the field values do not match exactly. Two contacts with the same email are an obvious duplicate. Two accounts named Acme Corp and Acme Corporation at the same domain are a less obvious one. Dedupe surfaces both.

Match strategies

Exact, normalized, and fuzzy.

A dedupe engine runs layered strategies. Exact email catches the clean duplicates. Normalized phone (strip punctuation, country code) catches the same person captured twice from different forms. Domain plus fuzzy name on company catches Acme Corp vs Acme Corporation. Levenshtein distance on a last-name plus email-local-part catches typos. Each layer owns a different failure mode.

Dedupe is not merge

Detection, triage, and resolution are three steps.

Dedupe detects candidate duplicate groups and flags them. A review step triages which pairs are true duplicates and which are false positives (two Johns at the same company). Merge is the resolution action that fuses a confirmed pair into a single row, picking surviving values and preserving history. Mature programs run the three steps as distinct stages, not one blob.

Why it matters

Double-touch, broken routing, messy reports.

Duplicates break the business before anyone sees them. Two reps call the same prospect in the same week. Routing sends a lead to the wrong owner because the real account record is a different row. Pipeline reports double-count the deal because it exists on both the parent account and the subsidiary. Dedupe is the hygiene layer that keeps the pipeline math honest.

When it runs

Inline on write, batch overnight.

Inline dedupe fires at the moment a record is created or imported, checking against the base before the row saves. Batch dedupe runs on a schedule across the whole CRM, catching the duplicates that slipped past inline because they came in through a different channel. Most mature teams run both, with inline owning new-record prevention and batch owning historical cleanup.

Scope

Contacts, accounts, and leads.

Dedupe runs on every object that represents a real-world entity. Contacts dedupe against other contacts and sometimes against leads. Accounts dedupe against other accounts, usually at the domain level. Leads dedupe against both leads and contacts to catch the case where a known buyer fills out a form from a new browser. The match rules differ per object, but the workflow is the same.

The match strategies

What catches which kind of duplicate.

A dedupe engine is only as good as its match rules. The strongest programs layer multiple strategies, because each catches a different failure mode and each has a different false-positive risk. Running a single strategy either misses obvious duplicates or merges distinct people into one row. The point of layering is to raise recall without blowing up precision.

Exact email

The baseline match.

An exact email match on the primary email field is the cleanest duplicate signal in B2B. Two records with the same email address are almost always the same person, with narrow exceptions for shared mailboxes. Exact email runs first, resolves the obvious cases cheaply, and sets the baseline recall every other strategy builds on.

Phone normalization

Same number, different formats.

Phone numbers arrive from forms, imports, and integrations in a dozen formats. Normalization strips punctuation, prefixes a country code when it is missing, and compares the resulting E.164 string. Normalized phone catches the duplicate that bypassed email dedupe because the buyer used a work email the first time and a personal one the second.

Domain plus fuzzy name

Acme Corp vs Acme Corporation.

Account dedupe keys on web domain first, because the domain is a strong identity signal that resists reformatting. Fuzzy name matching sits on top: Acme Corp, Acme Corporation, and Acme Inc at the same domain collapse to one candidate group. Domain plus fuzzy name is the workhorse strategy for account hygiene across the whole base.

Levenshtein distance

Edit distance for typos.

Levenshtein distance counts the number of single-character edits (insert, delete, substitute) that turn one string into another. A dedupe rule that flags last-name plus email-local-part pairs within a Levenshtein distance of two catches typos like jsmith vs jsmth and Jonson vs Johnson. The distance threshold is the knob that trades recall for precision.

Phonetic matching

Soundex for names that sound alike.

Soundex and Metaphone encode names by how they sound, not how they are spelled. Smith and Smyth share a code. Catherine and Katherine share a code. Phonetic matching catches the duplicates where the data entry clerk guessed at the spelling, and it layers cleanly with Levenshtein when both the spelling and the sound are close.

Composite keys

First + last + domain, or similar.

No single field matches every real duplicate. Composite keys combine multiple fields into a single match rule: first name plus last name plus email domain, or company name plus postal code plus main phone. Composite rules catch the cases where every individual field is only partially populated, but the combination still uniquely identifies the entity.

When the job runs

Inline, batch, and the review queue.

Dedupe is a workflow before it is an algorithm. The same match rules can block a save, flag a candidate, or merge silently, and each mode trades friction for safety differently. The programs that work write the rules down: which strategies auto-merge, which flag for review, and which simply warn the user at save time. The programs that fail let a batch job merge aggressively on day one and spend the next quarter unwinding the damage.

Inline on create

Catch duplicates before they land.

Inline dedupe runs at the moment a new record is created, imported, or captured from a form. The engine checks the incoming row against the base in real time and either blocks the save, routes it to a review queue, or merges automatically on a high-confidence match. Inline is the cheapest prevention layer because it stops the duplicate from existing in the first place.

Batch sweep

Nightly scan across the whole base.

Batch dedupe runs on a schedule across every row in the object, surfacing candidate groups the inline job missed. Batch is where historical duplicates (imported CSVs, legacy data, cross-system syncs) finally get flagged. The output is a review queue, not an automatic merge, because batch matches tend to be lower-confidence than inline matches on fresh data.

Review queue

A human confirms the borderline cases.

Not every candidate pair is a true duplicate. Two Johns at the same company, a parent account captured alongside its subsidiary, a legitimate second contact at a shared desk phone. The review queue is where a RevOps admin triages borderline matches, confirms the real duplicates, dismisses the false positives, and keeps the dedupe rules learning from the exceptions.

Confidence thresholds

Auto-merge, flag, or warn.

Every match has a confidence score. The rule set decides which band triggers which action. High-confidence matches (exact email, same domain plus exact company name) can auto-merge safely. Medium-confidence matches flag for review. Low-confidence matches warn the user at save time without blocking. Setting the thresholds is the single most important dedupe decision.

Cross-object dedupe

Leads against contacts, contacts against leads.

A known buyer fills out a new form and lands as a lead, even though they already exist as a contact on a closed-won account. Cross-object dedupe runs lead-to-contact matching before the lead ever reaches the SDR queue, so the sales team does not waste a touch pretending the buyer is net new. The same logic runs on contact-to-lead for inbound captures.

Audit trail

Every match decision, logged.

Every merge, every dismissed candidate, and every auto-merge write lands in an audit log. The log records which records were involved, which rule fired, what confidence score triggered the action, and which admin approved it. Audit coverage is how a team recovers from a bad batch merge, and how a compliance review proves the dedupe program is not silently losing records.

What to look for in a program

Rule governance, merge safety, and reporting.

Picking a match algorithm is the loudest decision but it is not the hardest one. The programs that work treat dedupe as a governed workflow with explicit rule ownership, merge-safety guardrails, cost controls on the review queue, and a quarterly review against ground truth. The programs that fail buy a dedupe tool, point it at the CRM, let it auto-merge aggressively, and discover six months later that half the merges fused unrelated people and the sales team has stopped trusting the data.

Rule ownership

Who writes the match rules.

Match rules are business logic, not IT configuration. The RevOps team owns the rule set because the trade-off between recall and precision is a sales-math decision, not a database decision. The rules live in a documented place, change on a reviewed cadence, and get tested against a sampled ground-truth set before they ship to production.

Merge safety

Pick surviving values, preserve history.

A merge should never silently lose data. The merging workflow picks a surviving value for each field (newest, most-complete, human-entered), preserves all activity history from both rows under the surviving record, and keeps a reversible audit of which record was the loser. A merge the team cannot undo is a merge the team will stop trusting.

False-positive handling

The engine learns from dismissals.

Every dismissed candidate pair is a training signal. A healthy program feeds the dismissed pairs back into the rule tuning: if a specific rule produces a high dismissal rate, the threshold is wrong and the rule needs tightening. Without that feedback loop, the review queue grows faster than the admins can clear it, and the whole program loses credibility.

Scope control

Not every record needs the same strictness.

Dedupe rules can differ by object, segment, or source. Lead dedupe might run tight to avoid wasting SDR time. Account dedupe might run loose on inbound and tight on strategic enterprise rows. Partner-sourced records might skip auto-merge entirely because the data quality is unknown. Scoping the strictness per segment keeps the program fit for the business.

Reporting

How many duplicates, how fast resolved.

A dedupe program needs an executive dashboard. Duplicate rate per object (contacts, accounts, leads). Review-queue age, so stale backlog does not hide the problem. Merge count per week, split by auto-merge and human-approved. Cost per resolved duplicate when a vendor is in the loop. The numbers are the only way to prove the hygiene work is paying for itself.

Quarterly review

Do the rules still fit the data.

Rules drift against the data. A new acquisition brings in records with a different naming convention. A new form captures a field the match rules do not consider. A new integration syncs in a slightly different phone format. A quarterly review samples the current base, measures precision and recall of each rule, and retunes the thresholds that have aged out of fit.

Dedupe inline, review in queue, merge with a full audit trail.

Strkr runs layered match strategies (exact email, phone normalization, domain plus fuzzy name, Levenshtein distance) across contacts, accounts, and leads, with confidence-tiered auto-merge, a review queue for borderline cases, and a reversible audit log on every merge. Pricing is published. The feature pages show exactly what ships today.

People also ask

Related questions.

What is the difference between data dedupe and record merge?

Data dedupe is the detection and triage step: it finds candidate duplicate groups using match strategies like exact email, phone normalization, and fuzzy name matching, and surfaces them for review or action. Record merge is the resolution step that fuses two confirmed duplicates into a single row, picking surviving field values and preserving activity history. Dedupe is the hunt, merge is the fix. A healthy program runs them as distinct stages.

What is a fuzzy match in CRM dedupe?

A fuzzy match flags two records that are not identical but are similar enough to likely represent the same entity. Fuzzy techniques include Levenshtein distance (edit distance on strings), Soundex and Metaphone (phonetic encoding for names), and token-based similarity on company names. Fuzzy matching catches typos, formatting differences, and name variants (Acme Corp vs Acme Corporation) that exact-match rules miss, with a confidence score the workflow uses to decide auto-merge versus review.

What is Levenshtein distance?

Levenshtein distance is a numeric measure of the difference between two strings, counting the smallest number of single-character edits (insert, delete, or substitute) needed to turn one into the other. In dedupe, a Levenshtein threshold of one or two edits catches typos (jsmith vs jsmth, Jonson vs Johnson) on names and email local parts. The threshold is the knob that trades recall (catch more) for precision (fewer false positives).

How often should a CRM run dedupe?

Inline dedupe should run on every record creation and every import, because preventing a duplicate is cheaper than resolving one later. Batch dedupe should sweep the whole base at least weekly against lightweight rules and monthly against the full composite rule set. Review queues should be worked daily so backlog does not hide the real duplicate rate. New integrations and bulk imports should kick off a one-time dedupe pass before the data lands in production reporting.

Should duplicates auto-merge or require review?

It depends on the confidence score. High-confidence matches (exact email on the same person, same domain plus exact company name) can auto-merge safely because the false-positive rate is near zero. Medium-confidence matches should flag for human review because the trade-off between a missed merge and a wrong merge needs a judgment call. Low-confidence matches should warn the user at save time without blocking. Setting the right thresholds is the most important dedupe decision.

What objects should a B2B CRM dedupe?

Contacts, accounts, and leads at minimum. Contacts dedupe against other contacts and against leads to catch the known buyer who re-enters as a new lead. Accounts dedupe against other accounts, usually on web domain first and fuzzy company name second. Leads dedupe against both leads and contacts. Beyond the core three, teams dedupe deals (the same opportunity captured twice in different stages) and even activity records when integrations sync noisy inbound data.

What causes CRM duplicates in the first place?

Four big sources. Forms: the same buyer fills out a webinar form, a demo form, and a content gate on three different days with slight variations. Imports: a CSV from a vendor lands next to records that already exist. Integrations: a marketing automation tool syncs leads that match existing contacts under a different unique key. Manual entry: a rep creates a new contact because they could not find the existing one, usually because of a typo in the search. Dedupe addresses all four.

How does dedupe differ from master data management?

Dedupe is a specific hygiene task: finding and merging duplicate rows inside a single system. Master data management (MDM) is the broader discipline of maintaining a single authoritative record for each real-world entity across multiple systems, often with a golden record stored in a dedicated MDM tool that other systems sync against. Dedupe is a prerequisite for MDM, because the golden record cannot exist until each system has resolved its own internal duplicates first.

Try it free. Bring your team next week.

No sales call, no migration consultant, no four-month implementation. Enter your card, get 14 days of the full Pro tier, cancel any time before day 14 with zero charge. Spin up a workspace, import your CSV, and have something useful before lunch.