Skip to main content

A robust fuzzy string matching library that combines syntactic and semantic approaches.

Project description

Hybrid Fuzzy Matcher

A robust fuzzy string matching library that combines syntactic (character-based) and semantic (meaning-based) approaches to provide more accurate and context-aware matching.

This package is ideal for tasks like data cleaning, record linkage, and duplicate detection where simple fuzzy matching isn't enough. It leverages thefuzz for syntactic analysis and sentence-transformers for semantic similarity.

Key Features

  • Hybrid Scoring: Combines fuzz.WRatio and fuzz.token_sort_ratio with cosine similarity from sentence embeddings.
  • Configurable: Easily tune weights and thresholds for syntactic and semantic scores to fit your specific data.
  • Preprocessing: Includes text normalization, punctuation removal, and a customizable abbreviation/synonym map.
  • Efficient Blocking: Uses a blocking strategy to reduce the search space and improve performance on larger datasets.
  • Easy to Use: A simple, intuitive API for adding data, finding matches, and detecting duplicates.

Installation

You can install the package via pip:

pip install hybrid-fuzzy-matcher

How to Use

Here is a complete example of how to use the HybridFuzzyMatcher.

1. Initialize the Matcher

First, create an instance of the HybridFuzzyMatcher. You can optionally provide a custom abbreviation map and adjust the weights and thresholds.

from hybrid_fuzzy_matcher import HybridFuzzyMatcher

# Custom abbreviation map (optional)
custom_abbr_map = {
    "dr.": "doctor",
    "st.": "street",
    "co.": "company",
    "inc.": "incorporated",
    "ny": "new york",
    "usa": "united states of america",
}

# Initialize the matcher
matcher = HybridFuzzyMatcher(
    syntactic_weight=0.4,
    semantic_weight=0.6,
    syntactic_threshold=70,
    semantic_threshold=0.6,
    combined_threshold=0.75,
    abbreviation_map=custom_abbr_map
)

2. Add Data

Add the list of strings you want to match against. The matcher will automatically preprocess the text and generate the necessary embeddings.

data_corpus = [
    "Apple iPhone 13 Pro Max, 256GB, Sierra Blue",
    "iPhone 13 Pro Max 256 GB, Blue, Apple Brand",
    "Samsung Galaxy S22 Ultra 512GB Phantom Black",
    "Apple iPhone 12 Mini, 64GB, Red",
    "Dr. John Smith, PhD",
    "Doctor John Smith",
    "New York City Department of Parks and Recreation",
    "NYC Dept. of Parks & Rec",
]

# Add data to the matcher
matcher.add_data(data_corpus)

3. Find Matches for a Query

Use the find_matches method to find the most similar strings in the corpus for a given query.

query = "iPhone 13 Pro Max, 256GB, Blue"
matches = matcher.find_matches(query, top_n=3)

print(f"Query: '{query}'")
for match in matches:
    print(f"  Match: '{match['original_text']}'")
    print(f"    Scores: Syntactic={match['syntactic_score']:.2f}, "
          f"Semantic={match['semantic_score']:.4f}, Combined={match['combined_score']:.4f}")

# Query: 'iPhone 13 Pro Max, 256GB, Blue'
#   Match: 'Apple iPhone 13 Pro Max, 256GB, Sierra Blue'
#     Scores: Syntactic=95.00, Semantic=0.9801, Combined=0.9681
#   Match: 'iPhone 13 Pro Max 256 GB, Blue, Apple Brand'
#     Scores: Syntactic=95.00, Semantic=0.9734, Combined=0.9640

4. Find Duplicates in the Corpus

Use the find_duplicates method to identify highly similar pairs within the entire corpus.

duplicates = matcher.find_duplicates(min_combined_score=0.8)

print("\n--- Finding Duplicate Pairs ---")
for pair in duplicates:
    print(f"\nDuplicate Pair (Score: {pair['combined_score']:.4f}):")
    print(f"  Text 1: '{pair['text1']}'")
    print(f"  Text 2: '{pair['text2']}'")

# --- Finding Duplicate Pairs ---
#
# Duplicate Pair (Score: 0.9933):
#   Text 1: 'Dr. John Smith, PhD'
#   Text 2: 'Doctor John Smith'

How It Works

The matching process follows these steps:

  1. Preprocessing: Text is lowercased, punctuation is removed, and custom abbreviations are expanded.
  2. Blocking: To avoid comparing every string to every other string, candidate pairs are generated based on simple keys (like the first few letters of words). This dramatically speeds up the process.
  3. Syntactic Scoring: Candidates are scored using thefuzz's WRatio and token_sort_ratio. This catches character-level similarities and typos.
  4. Semantic Scoring: The pre-trained sentence-transformers model (all-MiniLM-L6-v2 by default) converts strings into vector embeddings. The cosine similarity between these embeddings measures how close they are in meaning.
  5. Score Combination: The final score is a weighted average of the syntactic and semantic scores. This hybrid score provides a more holistic measure of similarity.

This approach ensures that the matcher can identify similarities even when the wording is different but the meaning is the same.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hybrid_fuzzy_matcher-0.1.0.tar.gz (9.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

hybrid_fuzzy_matcher-0.1.0-py3-none-any.whl (8.5 kB view details)

Uploaded Python 3

File details

Details for the file hybrid_fuzzy_matcher-0.1.0.tar.gz.

File metadata

  • Download URL: hybrid_fuzzy_matcher-0.1.0.tar.gz
  • Upload date:
  • Size: 9.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.5

File hashes

Hashes for hybrid_fuzzy_matcher-0.1.0.tar.gz
Algorithm Hash digest
SHA256 e46544d5e8aa2c2af205163395a4314e23abbc9c76add3dd31ff6dda41fd79bd
MD5 9931fb32a6a5971051fb69da0114e735
BLAKE2b-256 99fe7e4f5f371caf7d40052df902b66e7b3b836c876e764d1312c591bbd2defb

See more details on using hashes here.

File details

Details for the file hybrid_fuzzy_matcher-0.1.0-py3-none-any.whl.

File metadata

File hashes

Hashes for hybrid_fuzzy_matcher-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 8a20ae2336c47cafe912bfd5713a41d0f2f91451a881c7847d430f5b55b81e13
MD5 dbb66d3482aaf16343a3924cd52e0a11
BLAKE2b-256 dff966a6b6e9c7e0eb0930bf4cd195f1ce53c81f4788edc7a5e09b7edb002fe4

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page