near-dupes
Find near-duplicate texts, table rows and images with one call, then drop the copies and keep the best one.
Install
pip install near-dupes
Image support needs Pillow: pip install "near-dupes[images]".
Quickstart
import near_dupes
texts = ["The quick brown fox jumps over the lazy dog.", "the quick brown fox jumps over the lazy dog", "Something else entirely."]
result = near_dupes.find_duplicates(texts) # or a DataFrame, or a list of image paths
print(result.summary()) # 1 duplicate group: [0, 1]
print(result.dedupe(texts)) # keeps the longest copy plus the unrelated text
Rows of a DataFrame work the same way: near_dupes.dedupe(df, key=["name", "city"]) compares rows on
those columns and returns the frame without the copies.
What it does
- Text (
list[str]): character n-gram shingles, MinHash (128 permutations, fixed seed) and LSH banding propose candidate pairs; every candidate is then scored with the exact Jaccard similarity of its shingle sets, so the scores you get back are exact. With 2000 or fewer distinct texts there is no MinHash or banding at all: every pair is scored exactly (only when the texts are very long is each pair first screened by a MinHash estimate with a wide safety margin). Case, Unicode form (NFKC) and whitespace are normalized first. - Records (
pandas.DataFrame, a.csv/.parquetpath, or a list of dicts): the key columns of each row (default: all columns) are normalized for case, whitespace and punctuation, joined into one string and run through the text pipeline. Exact duplicates are always found; NaN cells count as empty. - Images (paths or
PIL.Imageobjects): 64-bit difference hash (dHash); similarity is1 - hamming / 64. thresholdis the minimum similarity in(0, 1];1.0means exact duplicates only.- Duplicate groups are the connected components of the pairs found. One representative per group is kept: the longest text, the first row, or the first image.
- Blank items (empty strings, rows whose key cells are all missing) are never treated as duplicates of each other.
- Result indices are 0-based positions in the input, so they line up with
listpositions anddf.iloc. - Everything is deterministic: the MinHash permutations come from a fixed
random_state.
API
near_dupes.find_duplicates(items, *, kind="auto", threshold=0.85, key=None, n_gram=3) -> DuplicateResult
near_dupes.dedupe(items, **kw) # find_duplicates(items, **kw).dedupe(items) in one line
items is a list of strings, a pandas.DataFrame (or a .csv/.parquet path, or a list of dicts), or a
list of image paths / PIL.Image objects. kind is detected from the input unless given ("text",
"records", "images"). key names the DataFrame column(s) to compare rows on; n_gram is the character
shingle length.
import pandas as pd
df = pd.DataFrame({"name": ["Acme Corp", "ACME Corp.", "Globex"], "city": ["New York", "new york", "Springfield"]})
near_dupes.find_duplicates(df, key=["name", "city"]).groups # [[0, 1]]
near_dupes.dedupe(df, key=["name", "city"]) # rows 0 and 2, original index kept
near_dupes.dedupe(["a.jpg", "a_copy.jpg", "b.jpg"]) # image paths; needs near-dupes[images]
DuplicateFinder(kind="auto", threshold=0.85, key=None, n_gram=3, num_perm=128, random_state=0, normalize=True, all_pairs_max=2000)
is the class underneath, with .find(items) and .dedupe(items), for the extra knobs.
text_similarity(a, b, n_gram=3) is the exact shingle Jaccard between two strings.
DuplicateResult
.groups-list[list[int]], each group of near-duplicate indices (size >= 2), ascending.pairs-list[(i, j, score)], the edges found (i < j); exact duplicates score1.0.representatives- the index kept for each group;.keep_indices- all singletons plus one representative per group, ascending;.drop_indices- whatdeduperemoves.n_duplicates,.n_groups,.n_items,.kind,.threshold.summary()- human-readable text;.to_dict()- JSON-safe dict.dedupe(items)-itemswith duplicates removed, same type as the input (a DataFrame keeps its original index labels)
CLI
near-dupes INPUT [INPUT ...] [--kind auto|text|records|images] [--threshold 0.85]
[--key COL [COL ...]] [--n-gram 3] [--json] [--output PATH]
near-dupes contacts.csv --key name emailprints the summary for a table.near-dupes tickets.txttreats each line as one text.near-dupes photos/ornear-dupes a.jpg b.jpgcompares images.--jsonprintsto_dict()as JSON;--output PATHwrites the deduplicated items (.csv/.parquetfor tables, one item per line otherwise).
License
MIT
Metadata
Release files for near-dupes 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| near_dupes-0.1.0.tar.gz | 27.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| near_dupes-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 50.7 kB
Release files / near_dupes-0.1.0.tar.gz
| Download URL | near_dupes-0.1.0.tar.gz |
|---|---|
| Size | 27.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1211561df7b1932e3fbeffc1243a0d01f8ccff23d0cffe8b756f7235713486ed
|
|
BLAKE2b-256 checksum How to use checksums |
d9047ae337c7299ee2e02cc2e3cae067f814c1f78a31fa0c3046e3e23da92c5f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.10.11
|
Release files / near_dupes-0.1.0-py3-none-any.whl
| Download URL | near_dupes-0.1.0-py3-none-any.whl |
|---|---|
| Size | 22.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6b5f1bb6d568978bdf15f598edc642ecb11421a96b34289c3b57df18a5b193dc
|
|
BLAKE2b-256 checksum How to use checksums |
116cb14fa468087f46281d1b9198da2fa8e13a23d39605de0cc0c7f7de0e1435
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.10.11
|