To merge duplicate leads without losing useful data, first decide which records represent the same person or company. Then choose the record to keep, fill its missing fields, and resolve conflicting values before applying the merge.
Deleting every repeated email row can discard a phone number, note, or CRM identifier that exists only in another row. Datablist separates duplicate detection from merge decisions, so you can inspect a group before changing it.
For leads spread across multiple databases, start with exports that include each source record ID. This guide uses a file-based workflow: clean the combined lead list in Datablist, then prepare updates for the source systems.
In this step-by-step guide, you will learn:
- How to find duplicate leads automatically
- How to dedupe leads automatically
- How to manually merge remaining duplicate leads
- How to update your CRM with your cleaned lead list
- How to export duplicate groups in an Excel file for external processing
Notes: This guide is about Lead Deduplication. But the process is similar for any list of records: Contacts, Companies, Products, etc. you want to dedupe.
Find duplicate leads
To begin, import your Leads Database into Datablist.
With Datablist, data is organized in collections. A collection stores a list of records sharing the same data model. You must import your leads using external files. Datablist supports CSV and Excel files. Click Import CSV/Excel, then select the file with your lead list.
Click the + to create a new collection. Give it a name (and an icon 🚀). Or click Start with a CSV/Excel file from the home screen.
Then, move to the Properties screen. This step lists the columns found when parsing the CSV file. Datablist checks each column to detect which data type should be used. For example, email addresses and urls are automatically detected.
Manually select data type when needed. Disable import if you have CSV columns that must not be imported.
The next import step displays a preview with the content of your file. Click Import {x} items to launch the import process.
If your leads are spread across several files, import them into a single collection and map equivalent columns to the same properties. Add a Source System column before combining files, and keep the source contact IDs. An ID from CRM A is not interchangeable with an ID from CRM B.
I would start by matching people on an appropriate identifier, such as a personal email or profile URL. Do not match on the source ID alone when each database uses its own ID scheme. For a company list, use company identifiers rather than assuming every person at the same domain is one lead.
Once the leads are loaded, open Clean > Merge duplicates.
Select how your leads should be compared to start the dedupe process. Two modes are available:
- All Properties compares complete lead rows.
- Selected Properties compares the fields that identify one lead.
Notes - In Datablist, the term "Property" is a synonym with Field or Column in other systems.
For lead deduplication, select Selected Properties.
Select the fields that identify the entity you want to deduplicate. A shared address such as sales@example.com can belong to several contacts. A name alone can also group different people.
For email comparisons, choose the processor deliberately. Email deduplication explains when normalized addresses can be useful and when aliases or shared inboxes need review. Start with a narrow comparison, inspect the groups, then broaden it only when the resulting matches represent the same entity.
Configure the algorithm and processor for each selected property, then click Run duplicates check.
Important
- The analysis is a read-only process. No data modification will be done until the next phase and the leads merge.
- Datablist compares text using a case-insensitive algorithm. If two values are similar but one has uppercases, they will be listed as duplicate leads.
Automatically dedupe leads
Select Merge and preserve data. Datablist fills empty fields from complementary leads, then lists the conflicting properties that need a rule.
For each conflict, combine distinct text values, keep the selected lead's value, or use a field survivorship rule. Configure Record to keep to preserve the right CRM or source record.
The preview shows values that will be filled, combined, kept, or discarded. Review it before processing Ready groups. Needs review groups remain unchanged.
Example: keep the contact ID and combine useful notes
These illustrative rows represent one person. The CRM contact has the identifier you want to retain; the event export contains complementary information.
| Source | Contact ID | Phone | Notes | Status | |
|---|---|---|---|---|---|
| CRM | C-104 | alex@example.com | Empty | Requested a demo | Qualified |
| Event export | E-218 | alex@example.com | +1 202 555 0147 | Met at an event | New |
Choose the CRM record as the record to keep. Fill its empty phone field, combine the distinct notes if you want both, and retain Qualified after checking the status conflict. The intended result is:
| Contact ID | Phone | Notes | Status | |
|---|---|---|---|---|
| C-104 | alex@example.com | +1 202 555 0147 | Requested a demo; Met at an event | Qualified |
Keep a copy of both original rows and their source IDs for reconciliation. Do not combine contact IDs into a single value and import that value into a CRM ID field.
Choose which record survives before resolving its fields
Record to keep selects the Datablist item that survives. Most Complete chooses the item with the most populated fields; that is not necessarily your preferred CRM contact. Additional record-selection rules depend on your plan. If no automatic rule reliably selects the intended contact, review that group manually.
For each conflicting property, decide what the field means:
- Notes or alternate contact details: combine distinct text values when retaining them is useful.
- CRM IDs, owners, or account references: keep one valid value after checking which record should survive. Dropping a conflict without this decision can break the mapping back to the source.
- Status: select the value that reflects your business rule. A more complete record does not automatically have the correct sales status.
- Activity dates: compare the business date field itself when choosing the latest activity. The Most recently updated field rule refers to the Datablist item's modification date, not the date of the activity recorded in a column.
Merge and preserve data fills complementary fields and applies the configured conflict rules. Remove duplicates keeps the selected record unchanged and removes the others. Use the second mode only when the discarded rows contain nothing you need.
Check the preview for values that will be discarded. Process Ready groups after reviewing them; leave Needs review groups for a manual decision. Retaining useful information sometimes means leaving a group unmerged until you know which value is correct.
Manually merge remaining duplicate leads
Use Datablist Merging Assistant to merge manually your remaining duplicate leads.
Open the remaining duplicate groups to inspect unresolved records.
On the left of each duplicate lead group, a Manual Merging assistant button opens the Merging Assistant.
It opens a merging tool. On the right, Datablist selects the record with the most data as "Primary item". And on the left, the remaining duplicate leads are called "Secondary Items".
When possible, property values from secondary items are auto-selected to be merged into the primary items. If several values conflict, you will have to make a decision and select which value to keep.
If the resulting "Primary item" suits you, click the Merge button to confirm the merge process. All the secondary leads will be deleted to keep only one combined lead record.
You can also edit or delete your duplicate leads directly from this listing.
Update your CRM with your updated leads list
Manage multiple values in a single cell
Datablist combines values into a single cell. You can end up with a listing with several values merged with a delimiter.
For example, a combined phone field might contain +1 202 555 0147;+1 202 555 0182. If your destination requires one number per field, split it into Phone 1 and Phone 2 before importing it. Preserve the original combined field until you have checked the result.
To manage this transformation, you can:
- Use Datablist Split Property feature to create several properties from multi values data
- Or run a script code directly in Datablist to process this splitting.
- Or export your leads listing into an Excel file and post-process it with Excel or Google Sheets.
If you need one exported row per phone number or email instead of several columns, use the CSV Rows Splitter on the merged file.
How to use "Split Property" to split multi values data into several properties
Datablist has a built-in tool to split the text from a property into new properties. This is a perfect tool to deal with combined results from the deduplication algorithm.
Open the tool by clicking on Split Property in the Edit menu.
Select the property with the multi values. And choose the same delimiter that you used when combining.
The last setting defines how many parts will be created. Choose enough output fields for the largest number of values you intend to retain. Check rows with more values than your chosen output count before processing.
Before processing your data, Datablist shows a preview of the result. Check that the split data correspond to the expected results. Then click on Split Property to process all your data.
After the processing, your initial property is unchanged and new properties are created to store the split values. Rename them to match your CRM import columns.
Split values on delimiter using a JavaScript script on Datablist
For more complex splitting or if you need extra manipulation, Datablist has a powerful tool to run JavaScript code on your data. This tool can be used to split your text into several properties.
First, create extra properties to store your multiple values if they are not created. Create a Phone 2, Phone 3 properties or Email 2, Email 3 that will store a single value after the split.
Then, click on Run Javascript in the Edit menu to open the script editor.
Adapt the script below to fit your properties:
function runOnItem(item){
if(!item.phone) return null;
var parts = item.phone.split(';');
if(parts.length===1) return null;
return {
phone1: parts[0],
phone2: parts[1]
}
}
Note: Process each combined property separately. If you have a property with phone numbers, and another one with email addresses. First process the phone number with a script, then run a second one for your email addresses.
Here is an example of code that split the content of a property with the key phone1. The split is done on a semicolon. And the resulting phone numbers are stored in 2 properties: phone1 and extraphone.
Please contact us if you have questions about how to write the script.
Reconcile the cleaned file with the source databases
A merge in Datablist changes the imported collection. It does not automatically update or delete the corresponding records in your external databases. Prepare that reconciliation separately.
Export the surviving source ID alongside the values you intend to update. For each removed source record, decide whether the destination requires an update, a merge operation, or a deletion. Follow that system's import and merge rules, especially for linked activities and accounts.
Paid plans can download the merge changes list, which includes previous and destination values for tracking changes. Use it alongside your original export. Test the destination update on a small group before applying the full file.
Export duplicate groups into an Excel or CSV file
At any time in your deduplication process, you can export the remaining duplicates. Datablist exports data in Excel or CSV files.
Export the duplicates when you want to clean them manually using Excel or outsource the task to an external provider.
FAQ
What is lead duplication?
Lead duplication means that more than one record represents the same person or company. Lead deduplication finds these records and decides whether to keep, remove, or merge them. Multiple people at the same company are not necessarily duplicates.
How do I handle deduplication in a lead list?
Choose an identifier that fits your entity, run a duplicate check, and inspect the groups before applying changes. Then choose the record to keep and set a rule for each conflicting field. Shared inboxes and similar names need more context than a simple equality match.
How do I merge duplicate leads across multiple databases?
Export the relevant fields and source IDs from each database, add source labels, and map equivalent fields into one Datablist collection. Match on person or company identifiers, resolve conflicts, and export the surviving IDs with the cleaned values. Reconcile those results with each source database separately.
Can all lead properties be combined?
Text-based values can be combined with a delimiter. Fields such as a status, numeric value, date, or CRM ID often need one selected value instead. Combining every conflict can create an output that the destination cannot import correctly.
How do I retain multiple phone numbers?
Combine the distinct numbers in a text field, then split that field into the output columns your CRM expects. Check the delimiter and the required number of columns. See managing multiple values above.
Do I need to resolve every group at once?
No. You can process reviewed groups and leave uncertain matches unchanged. Start with clear matches, then examine ambiguous groups. The preview and remaining groups help you keep uncertain decisions separate.
How long does deduplication take?
The time depends on the list size, comparison settings, browser resources, and conflict review. Detection and the human work of resolving conflicting leads are separate steps. Keep an original export before applying changes; do not rely on a promised number of seconds.
Does this workflow connect directly to my CRM?
This guide uses CSV or Excel exports. Importing those files into Datablist does not establish a live CRM connection. Check your destination's requirements for updating IDs, deleting duplicates, and retaining relationships before importing the cleaned results.



















