Answers

What is duplicate management?

A one-time dedupe run clears the backlog. Duplicate management is what keeps the backlog from growing back. The policy, the match rules, the approval queue, and the lineage log all live under the same owner.

Short answer

Duplicate management is the ongoing policy framework that prevents and resolves duplicate records inside a CRM. It combines prevent-on-save rules that block or warn at create, nightly match jobs that find duplicates after the fact, merge approval workflows that route resolutions through a reviewer, and lineage tracking that records which record absorbed which. Duplicate management is a standing RevOps responsibility, not a one-time dedupe project, and it owns both the rules and the audit trail.

Key points

What matters most.

The six ideas every revenue team should understand before standing up a duplicate management program: the rules that catch duplicates, the jobs that find them, and the audit trail that explains every merge.

Definition

A standing policy, not a cleanup project.

Duplicate management is the ongoing framework that governs how a CRM prevents, detects, and resolves duplicate records. It owns the match rules, the prevent-on-save behavior, the nightly match job, the merge approval queue, and the lineage log. A one-time dedupe clears the base once. Duplicate management keeps it clean forever.

Prevent on save

Catch the duplicate before it lands.

Prevent-on-save rules fire at the moment of create. The form or import pipeline runs a match query, and when a candidate scores above the threshold the system blocks the save, warns the user, or routes to the existing record. Prevention is the cheapest fix because the duplicate never joins the base in the first place.

Match jobs

Nightly sweep against the whole base.

Match jobs run on a schedule, usually nightly, and compare every record against every other record using a configured match rule. Candidates land in a review queue grouped by likely match. The job catches the duplicates that slipped past prevention, including the slow drift of records that only converge after an enrichment update.

Merge approval

A reviewer confirms before records collapse.

Merge approval workflows route each proposed merge through a reviewer with the authority to confirm, decline, or defer. The reviewer picks which record survives, which fields win in a conflict, and whether to escalate for a second opinion. Approval keeps a bad merge from quietly destroying history that only one record had.

Lineage

Every merge leaves a traceable record.

Lineage tracking records which record absorbed which, when the merge happened, who approved it, which match rule proposed it, and which fields moved. The log makes merges reversible in principle, debuggable in practice, and auditable for compliance. Without lineage, a merged record is just a quiet deletion nobody can prove.

Owner

A named RevOps responsibility.

Duplicate management is a standing RevOps function with a named owner. The owner tunes the match rules, reviews the queue cadence, publishes the policy, and reports the duplicate rate as an operating metric. Teams that leave duplicate management unassigned end up with the same dedupe project, every eighteen months, run by whoever has the time.

The four control points

Prevent, detect, resolve, record.

A mature duplicate management program runs on four control points, each with its own tooling, its own owner, and its own metric. Prevention catches duplicates at the entry point. Detection sweeps the base for what prevention missed. Resolution routes the proposed merges through a review workflow. Recording captures the lineage so every change is explainable. Teams that build only one or two of the four points end up with gaps that quietly refill the backlog.

Prevent on save

Block or warn before the record commits.

A prevent-on-save rule runs at create time against every form, import, and API write. The rule queries the base for near-matches on the fields that matter (company domain, normalized email, phone), scores the candidates, and either blocks the save, warns the user with a confirm step, or silently routes the write to the existing record. Prevention is cheaper than detection by an order of magnitude.

Nightly match

Sweep the base for what prevention missed.

A scheduled match job compares every record against every other record using fuzzy matching on normalized keys. The job catches duplicates that enrichment created later, duplicates from legacy imports, and duplicates where the match fields were empty at create time and only populated after. The output is a queue of candidate pairs, grouped and ranked by match confidence.

Review queue

A human reviews before merge.

The match job never merges silently. Every candidate pair lands in a review queue where a reviewer confirms the match, picks the surviving record, resolves field-level conflicts, and approves the merge. High-confidence pairs can be routed to a fast-path reviewer. Low-confidence pairs can be deferred or declined. The queue is where policy becomes action.

Lineage log

Every merge, with source and winner.

After each merge, a lineage entry writes to the surviving record: which record was absorbed, which match rule proposed it, which reviewer approved it, which fields moved from the absorbed record, and when the merge happened. The log is how a sales rep understands why a familiar contact looks different today, and how a compliance review proves data integrity.

Rule tuning

Match thresholds evolve with the base.

The match rules that worked at ten thousand records will produce too many false positives at a hundred thousand. Duplicate management includes a cadence for reviewing match performance: precision, recall, reviewer agreement rate, and queue depth. The RevOps owner tunes the thresholds and the key field weights every quarter against the actual queue outcomes.

Policy doc

What counts as a duplicate, written down.

The policy document names which fields define identity for an account (company legal name plus domain), for a contact (normalized email plus company), and for a lead (email plus phone). It names who can approve a merge, who can override a prevention block, and how a false merge gets reversed. Written policy is what makes the program durable across RevOps turnover.

How the rules actually fire

Match rules, thresholds, and the save path.

The policy document reads clean. The implementation is where programs succeed or fail. Match rules compare normalized values across fuzzy string, phonetic, and token-set metrics. Thresholds decide when a candidate is high enough confidence to block, high enough to warn, or low enough to let through for the nightly job. The save path has to run these checks in a few hundred milliseconds without breaking the form submit experience, which rules out approaches that scan the whole base per write.

Normalized keys

Match the shape, not the string.

A match rule compares normalized values, not raw strings. Emails lowercase and strip aliases. Phones strip formatting and country codes. Company domains strip www and tld variants. Names apply phonetic encoding so Catherine and Kathryn hit the same bucket. Normalization is what turns a probabilistic match problem into a deterministic one for the ninety percent of cases.

Confidence score

Each candidate carries a number.

Every candidate pair gets a confidence score from the match rule, usually zero to one. The score combines the per-field similarity (string distance, phonetic match, token overlap) weighted by how identifying each field is. Email weighs higher than title. Company domain weighs higher than company name. The score is what the threshold compares against.

Block vs warn

Two tiers, two user experiences.

A high-confidence candidate (above the block threshold) refuses the save and tells the user which existing record matched. A medium-confidence candidate (above the warn threshold but below block) shows a confirm step with the candidates listed. A low-confidence candidate lets the save through and routes the pair to the nightly match job. Two tiers protect data without blocking the business.

Entry points

Every save path runs the same rule.

The duplicate rule has to fire on every save path: web form submits, mobile app writes, bulk imports, Zapier-style integrations, API writes from marketing platforms, and manual user entry. A rule that only runs on the web form leaks duplicates from every other path. Centralizing the match in the write pipeline, not the UI, is how the rule stays enforced.

Performance

Match in milliseconds, not seconds.

A prevent-on-save check has a hard latency budget. The match query needs to return in a few hundred milliseconds or the form feels broken. The database schema needs match-optimized indexes on normalized key columns. The match rule needs to run as a bounded candidate lookup, not a full table scan. Performance is a design concern, not a tuning afterthought.

Override

A supervisor can let the save through.

No match rule is perfect. The user occasionally knows the near-match is a different person at the same company, or the same company at a different division. The policy names who can override a block and the audit log records every override with the reason given. Overrides without logging are how teams lose trust in the rule they stood up.

The resolution workflow

From queued pair to merged record.

The nightly match job produces a queue of candidate pairs. The resolution workflow is how those pairs become either merges, dismissals, or deferrals. The workflow has to be fast enough that reviewers actually work through it, structured enough that two reviewers would reach the same answer, and audit-grade enough that a compliance review can see the full trail years later. Duplicate management lives or dies on this workflow, because a queue nobody works is just a report.

Grouping

Pairs collapse into clusters.

The review queue groups candidate pairs into clusters when the pairs share a record. A single company with four duplicates becomes one cluster of four, not six overlapping pairs. The reviewer picks a winner once per cluster and the system merges the remaining three into it. Clustering is what makes the queue tractable at scale.

Field-level winners

The reviewer picks which value survives.

When two records disagree on a field, the reviewer picks which value survives on the merged record. The UI shows the two values side by side with their source and last-modified timestamp. The reviewer picks per field, or accepts a default rule (most-recent-non-blank). The merge preserves field-level provenance so the surviving record remembers where each value came from.

Related records

Children follow the parent.

A merged company brings its contacts, deals, activities, tasks, cases, and notes with it to the surviving record. The merge workflow previews the counts before the reviewer confirms. The surviving record ends up with the union of related records from both sides, with any duplicates among the related sets flagged for a second-pass review.

Decline

Not every candidate is a duplicate.

Reviewers decline candidate pairs that look similar but represent genuinely distinct entities: two contacts at the same email alias, two companies with the same name in different industries, two leads from the same person at different roles. Declined pairs record the reason and the match rule learns to deprioritize that signature. Decline reasons are training data for rule tuning.

Defer

The reviewer needs more context.

A deferred pair stays in the queue but moves off the primary review surface until a specific condition fires: enrichment refreshes the record, a human writes a note, a new activity hits either side. Defer is how reviewers avoid guessing when the signal is thin, without losing the match. Deferred pairs resurface automatically rather than falling into permanent limbo.

Reversal

Every merge is undoable in principle.

A bad merge happens. The lineage log makes the merge technically reversible: the absorbed record can be reconstructed from the stored snapshot plus the field moves, and the related records can be split back along their original parent. Reversal is rare but the capability is non-negotiable, because the alternative is irreversible silent data loss on a reviewer mistake.

Prevent duplicates on save, resolve them through approval.

Strkr ships prevent-on-save duplicate rules, nightly match jobs, a review queue, merge approval workflows, and lineage tracking against the same contact, company, lead, and deal records that drive routing, scoring, and forecasting. Pricing is published. The feature pages show exactly what ships today.

People also ask

Related questions.

What is the difference between duplicate management and data deduplication?

Data deduplication is the one-time act of finding and merging duplicates in a CRM, often run as a project after an import or an acquisition. Duplicate management is the ongoing policy framework that prevents duplicates from entering, detects duplicates that slipped past prevention, routes resolutions through an approval workflow, and records lineage for every merge. Dedupe is an event. Duplicate management is the standing capability that makes the event unnecessary.

What is a prevent-on-save duplicate rule?

A prevent-on-save rule fires at the moment a new record is created. The rule queries the existing base for near-matches on identifying fields (email, company domain, phone), scores the candidates, and either blocks the save, warns the user with a confirm step, or routes the write to the existing record. Prevention is the cheapest form of duplicate management because the duplicate never enters the base, which means no match job has to find it later.

What is a merge approval workflow?

A merge approval workflow routes every proposed merge through a reviewer with authority to confirm, decline, or defer. The reviewer picks which record survives, resolves field-level conflicts, and confirms the merge of related records. Approval prevents silent destructive merges, keeps the audit trail honest, and gives the data owner a decision point before history collapses onto a single record. The workflow sits between the match job output and the actual merge.

What is lineage tracking in duplicate management?

Lineage tracking is the audit record of every merge: which record was absorbed, which record survived, which match rule proposed the pair, which reviewer approved it, which fields moved, and when the merge happened. The lineage log makes merges debuggable (why does this record look different today), reversible in principle (restore the absorbed record from the snapshot), and compliance-grade for data integrity reviews.

How often should duplicate match jobs run?

Prevent-on-save runs on every write in real time. The full-base match job should run at least nightly against new and recently modified records, and weekly against the whole base. Match-rule tuning should happen quarterly, by reviewing queue outcomes and reviewer agreement rates. Enrichment-triggered match should fire whenever an enrichment update populates a previously blank identifying field, because that is the moment two records most often converge.

Who owns duplicate management in a company?

Duplicate management is a standing RevOps responsibility with a named owner, typically a RevOps analyst or a sales operations manager. The owner publishes the policy, tunes the match rules, reviews the queue cadence, trains new reviewers, and reports the duplicate rate as an operating metric. Companies that leave duplicate management unassigned end up with the same dedupe project every eighteen months, run by whoever happens to have the time.

What fields define a duplicate in a CRM?

The identifying fields depend on the object. Accounts usually match on normalized company domain plus legal name. Contacts match on normalized email plus company. Leads match on email plus phone. Deals rarely match on identity (they are events) but can match on account plus close date plus amount band. The policy document names the identity fields per object, because an unnamed match rule drifts from reviewer to reviewer.

Can duplicate management be fully automated?

Prevention and detection can run fully automated. Resolution should not. The highest-confidence candidate pairs (an exact email match on identical normalized domain) can auto-merge safely, but the long tail of near-matches needs a human to pick which record survives and which values win. Fully automated merges without approval are the fastest way to lose history that only one of the two records had, and the slowest to detect because the surviving record still looks populated.

Try it free. Bring your team next week.

No sales call, no migration consultant, no four-month implementation. Enter your card, get 14 days of the full Pro tier, cancel any time before day 14 with zero charge. Spin up a workspace, import your CSV, and have something useful before lunch.