Duplicate Records Across Files

Compare two CSV datasets for exact keys and separately labeled similarity candidates.

Files stay on your device No sign-up Free to use
How this works

The tool runs in this browser. Your file or text is not uploaded to UseFreeTools. Check this tool's limits for anything it may save on your device.

Privacy details

Process files controls

Drop your CSV files here

or choose them from your device

Exactly two UTF-8 CSV files, up to 2 MiB each, 10,000 combined data rows and 50 columns per file. Each file needs distinct column names.

    Comma-separated column names present in both files. Names containing commas can be quoted, as in CSV.

    One key column only. One Unicode code point insertion, deletion or substitution after the chosen normalization, not a full grapheme comparison. Keys up to 80 code points; 120,000 signatures, 20,000 candidate checks and 5,000 total pairs. No automatic merging.

    How to use Duplicate Records Across Files

    1. Choose two supported UTF-8 CSV files; use Up/Down to set file A then file B.
    2. Set matching columns and any bounded similarity review options.
    3. Inspect exact pairs separately from similarity candidates and download the review.

    Example: Duplicate Records Across Files

    Two CSV tables with key values 001, 002 and 002, 003 respectively.

    You add
    Two CSV files: a.csv contains id,name with 001,Ana and 002,Lee; b.csv contains 002,Sam and 003,Mia. Key columns: id Trim surrounding whitespace in keys: on Ignore letter case in keys: off Also list near candidates for manual review: off
    You get
    Found 1 cross-file row pairs using id. Sample rows: 3 | 2 | Exact selected-key match | 002 | 002.

    Options

    Key columns
    Choose the same named columns in both files. Matching uses these values after the selected trimming and case rules; it does not compare every cell in a row.
    Near candidates
    Enable this only for one key column. A one-code-point edit can find a possible spelling difference, but it can also connect different records. Review the listed row pair before taking any action.

    Supported inputs and limits

    Two UTF-8 CSV files up to 2 MiB each, 10,000 total rows and 50 columns. Similarity comparisons have a hard candidate cap and can omit unmatched blocks; they are review candidates rather than identity proof. Exports protect risky spreadsheet formula prefixes.

    Where your input is processed

    This tool processes your input in this browser. Your text and files are not uploaded to UseFreeTools. Check this tool's limits for anything it may save on your device.

    A pair is evidence for review, not a merge instruction

    Row numbers count the header as row 1. Two matching keys can have different names, addresses or other data, and repeated keys can produce several row pairs. Open both original rows and compare the fields that matter to your task. The exported match list leaves the source files unchanged; its formula guard also means a spreadsheet cell may display an added apostrophe.

    Questions about Duplicate Records Across Files

    Will matching records be merged?

    No. The output identifies pairs for your review.

    Is a near match a confirmed duplicate?

    No. Similar text can describe different records; inspect the selected fields and original data.

    Why are leading zeros preserved?

    CSV key fields are compared as text so identifiers such as 001 keep their entered form.

    Project manager: Tony Hines · Content updated 4 October 2026 · Report a problem