Back to the journal
RSS feed

CRM data synchronisation: why duplicates keep coming back

Stop duplicate contacts returning after a CRM clean-up. Set matching rules, field ownership and safe retries before connecting more tools or adding AI.

Areza Digital

Two customer record cards connect through a matching gate to one consistent record with a check mark

You merge two customer records on Friday. On Monday, the duplicate is back. The website has created another contact, an old import has restored the previous email address, and two salespeople are following up with the same person.

CRM data synchronisation needs rules for recognising an existing customer and deciding which changes to keep. Cleaning a database once cannot prevent the next integration from creating the same problem.

Before adding another connector or an AI assistant, answer four questions: which records belong together, which system owns each field, what happens when a request runs twice, and what should happen after a record is merged or removed?

How one customer becomes three records

Consider this illustrative example, not a client case study.

A buyer asks for a quote through your website. The CRM creates contact C-1042. Your order system later creates customer O-778. A salesperson also imports a spreadsheet containing the buyer’s previous email address.

Three records are not automatically three customers. Keeping one record in each system is normal. The problem starts when the connection between them is missing: every new message or import becomes another opportunity to create a contact.

If the original enquiry never reaches anyone, start with our guide to website enquiry handovers. Here, the enquiry arrives. The failure is deciding where it belongs.

1. Match the customer before copying the fields

A field mapping says that a website’s “email” belongs in the CRM’s email field. A matching rule decides whether the incoming information belongs to an existing person. You need both.

HubSpot’s data-sync documentation distinguishes these steps: matching runs separately from field mappings, and paired records are subsequently linked through their record IDs. Source: HubSpot record matching.

For a new CRM integration, use a stable identifier already shared between systems where one exists. Otherwise, define the evidence needed for an initial match, then preserve the resulting link. In our example, that means retaining the relationship between C-1042 and O-778 rather than rediscovering it from the buyer’s name every time.

Email can help with that first match, but it is not a permanent identity. People change addresses. Several people can use one purchasing inbox. A person’s employer and a company’s domain can change too.

Send ambiguous matches for review. Two people called Alex working at the same company are not enough evidence for an automatic merge.

Also check what your CRM does when it finds a match. Salesforce separates matching rules, which identify possible duplicates, from duplicate rules, which govern how they are handled. Detection and prevention are separate decisions. Source: Salesforce matching rules.

2. Give each field a clear owner

“The CRM is our source of truth” sounds settled until an order changes, a form arrives with a blank phone number, or a salesperson corrects the company name.

Choose the authoritative source for each piece of information. The following is an example policy for a sales team, not a universal configuration.

Swipe or scroll to compare.

Information Authoritative source Rule for incoming changes
Sales owner and deal stage CRM Forms and imports cannot reset them
Payment status Order or accounting system CRM displays the status without rewriting it
Contact phone number Reviewed contact record A blank form field does not erase a known number
Delivery address Individual order A one-off destination does not replace the account address
Linked record IDs Integration’s record mapping Email changes do not break the established link

Make the difference between missing, blank and intentionally cleared explicit. A form that does not ask for a phone number has supplied no update. A person deliberately removing an incorrect number has.

Avoid applying “latest update wins” to every field. A delayed spreadsheet import can arrive after a correction while still containing older information. Arrival time does not make its contents more accurate.

3. Make repeated requests harmless

Imagine the CRM creates a contact successfully, but the connection drops before the website receives confirmation. The website retries. If the integration treats every attempt as a new enquiry, one submission can create two records.

Give each submitted enquiry a stable request identifier that survives retries. Keep a record of the completed action, so repeating that same request returns the existing result. Developers call this idempotency: the same operation can be repeated without creating another copy.

The repeated-request check must also hold when two attempts arrive together. A lookup followed by “create if missing” can still produce duplicates if both attempts check before either finishes. Use the destination’s supported idempotency mechanism or an atomic uniqueness guarantee where available, and test the actual connector’s behaviour.

Do not deduplicate every enquiry from the same customer into one event. A buyer asking for a second quote has made a new request, even when the contact stays the same.

Two-way sync needs another boundary: a change copied from system A into system B should not bounce back as a fresh change indefinitely. Track where an update came from and whether the destination already has the intended value.

4. Plan what happens after a merge

Merging contacts in the CRM does not automatically tell every connected tool which record survived.

If an integration still points to the retired contact, its next update can fail or recreate an unwanted record. A reviewed merge should update the stored links and preserve the relevant conversation, order and activity history.

Treat removal separately. Disconnecting a record, archiving it and deleting it are different actions. Decide how each connected system should respond; do not let a routine import silently recreate something that was deliberately removed.

Before bulk cleanup, keep a recoverable export and check the CRM’s merge behaviour. Start with a small reviewed group so you can inspect the linked history and downstream records.

Five tests before enabling the full sync

Use test records and check the destination, not just a green connector status.

  1. Submit one enquiry twice. One enquiry and one contact should remain, with a visible record of the retry.
  2. Submit another enquiry from the same buyer. It should attach to the existing contact without losing the new request.
  3. Change the buyer’s email address. Previously linked systems should continue updating the same person.
  4. Send conflicting updates. Test a blank phone field and an older import arriving after a correction; the agreed ownership rules should hold.
  5. Merge a duplicate and replay an old update. The update should reach the surviving record or a review queue, without recreating the duplicate.

Check unresolved matches, repeated failures and overwritten corrections during a limited rollout. A connector reporting “success” only tells you that its operation completed; inspect whether the right customer’s information changed.

Where AI helps with CRM data quality

An AI assistant can suggest that two differently written company names refer to the same organisation, summarise conflicting notes, or prepare a review queue. Those suggestions still need evidence and a clear approval step where a wrong merge would matter.

An existing record ID does not need a language model to match it. A repeated request does not need one either. Use fixed rules for those decisions and reserve AI for genuinely ambiguous information.

Areza’s connected systems work brings websites, CRMs and internal tools together around these decisions. If duplicate contacts keep returning, tell us which systems create or update your customer records. That is a more useful starting point than another database clean-up.