Skip to main content

String Grouper

pypi license lastcommit codecov PyPI Downloads

Click to see image

The image displayed above is a visualization of the graph-structure of one of the groups of strings found by string_grouper. Each circle (node) represents a string, and each connecting arc (edge) represents a match between a pair of strings with a similarity score above a given threshold score (here 0.8).

The centroid of the group, as determined by string_grouper (see tutorials/group_representatives.md for an explanation), is the largest node, also with the most edges originating from it. A thick line in the image denotes a strong similarity between the nodes at its ends, while a faint thin line denotes weak similarity.

The power of string_grouper is discernible from this image: in large datasets, string_grouper is often able to resolve indirect associations between strings even when, say, due to memory-resource-limitations, direct matches between those strings cannot be computed using conventional methods with a lower threshold similarity score.

———

This image was designed using the graph-visualization software Gephi 0.9.2 with data generated by string_grouper operating on the sec__edgar_company_info.csv sample data file.


string_grouper is a library that makes finding groups of similar strings within a single, or multiple, lists of strings easy — and fast. string_grouper uses tf-idf to calculate cosine similarities within a single list or between two lists of strings. The full process is described in the blog Super Fast String Matching in Python.

Installing

pip install string-grouper

Speed

string_grouper leverages the blazingly fast sp_matmul_rs (originally based on: sparse_dot_topn) to calculate cosine similarities.

s = datetime.datetime.now()
matches = match_strings(names["name"], number_of_processes=15)
e = datetime.datetime.now()

diff = e - s
print(diff)

Results in:

00:17.80 On an m5 pro, where len(names) = 663 000

in other words, the library is able to perform fuzzy matching of 663 000 names in less than 18 seconds on a 2026 consumer CPU using 15 cores.

Simple Match

import pandas as pd
from string_grouper import match_strings

company_names = 'sec__edgar_company_info.csv'
companies = pd.read_csv(company_names)
# Create all matches:
matches = match_strings(companies['Company Name'])
# Look at only the non-exact matches:
matches[matches['left_Company Name'] != matches['right_Company Name']].head()
left_index left_Company Name similarity right_Company Name right_index
15 14 0210, LLC 0.870291 90210 LLC 4211
167 165 1 800 MUTUALS ADVISOR SERIES 0.931615 1 800 MUTUALS ADVISORS SERIES 166
168 166 1 800 MUTUALS ADVISORS SERIES 0.931615 1 800 MUTUALS ADVISOR SERIES 165
172 168 1 800 RADIATOR FRANCHISE INC 1 1-800-RADIATOR FRANCHISE INC. 201
178 173 1 FINANCIAL MARKETPLACE SECURITIES LLC /BD 0.949364 1 FINANCIAL MARKETPLACE SECURITIES, LLC 174

Group Similar Strings and Find most Common

companies[["group-id", "name_deduped"]] = group_similar_strings(companies['Company Name'])
companies.groupby('name_deduped')['Line Number'].count().sort_values(ascending=False).head(10)
name_deduped Line Number
ADVISORS DISCIPLINED TRUST 1747
NUVEEN TAX EXEMPT UNIT TRUST SERIES 1 916
GUGGENHEIM DEFINED PORTFOLIOS, SERIES 1200 652
U S TECHNOLOGIES INC 632
CAPITAL MANAGEMENT LLC 628
CLAYMORE SECURITIES DEFINED PORTFOLIOS, SERIES 200 611
E ACQUISITION CORP 561
CAPITAL PARTNERS LP 561
FIRST TRUST COMBINED SERIES 1 560
PRINCIPAL LIFE INCOME FUNDINGS TRUST 20 544

Documentation

The documentation can be found here

Backends

The library was originally developed using the sparse_dot_topn library, but has since been rewritten to use the sp_matmul_rs library, which is a Rust implementation of the sparse matrix multiplication algorithm optimized with Claude Fable. To run the library with the original sparse_dot_topn backend, set use_sp_matmul_rs to False.

Metadata

Release files for string-grouper 0.8.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for string-grouper 0.8.0
File Size Uploaded
string_grouper-0.8.0.tar.gz 2.4 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for string-grouper 0.8.0
File Interpreter ABI Platform
string_grouper-0.8.0-py3-none-any.whl Python 3 none any Details

Total release size: 2.5 MB

Release files / string_grouper-0.8.0.tar.gz

Download URL string_grouper-0.8.0.tar.gz
Size 2.4 MB
Tags Source
SHA-256 checksum
How to use checksums
cfd147f7c13ef0f181defc41b026ce468045a4c21f9aac7d20ffbd5b96d1d4e8
BLAKE2b-256 checksum
How to use checksums
696a59d613d3561bf1c97fbb938311ba9f49fd54cd6cb8288c46266d1d9823e3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 26, 2026.

Transparency log

Release files / string_grouper-0.8.0-py3-none-any.whl

Download URL string_grouper-0.8.0-py3-none-any.whl
Size 31.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2607b28840c70d17b4a8f0f38e1ad050afab9c4d45cfef043a5f729e12192680
BLAKE2b-256 checksum
How to use checksums
7eee13e29384d04dfa5d006dbf25c1e7a3d9afbfc98cbbc15c5df1536b38d467
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 26, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.8.0 This release

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.6.1

1 release file

0.6.0

3 release files

0.5.0

1 release file

0.4.0

1 release file

0.3.2

1 release file

0.2.2

1 release file

0.1.2

1 release file

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page