String Grouper
Click to see image
The image displayed above is a visualization of the graph-structure of one of the groups of strings found by string_grouper. Each circle (node) represents a string, and each connecting arc (edge) represents a match between a pair of strings with a similarity score above a given threshold score (here 0.8).
The centroid of the group, as determined by string_grouper (see tutorials/group_representatives.md for an explanation), is the largest node, also with the most edges originating from it. A thick line in the image denotes a strong similarity between the nodes at its ends, while a faint thin line denotes weak similarity.
The power of string_grouper is discernible from this image: in large datasets, string_grouper is often able to resolve indirect associations between strings even when, say, due to memory-resource-limitations, direct matches between those strings cannot be computed using conventional methods with a lower threshold similarity score.
This image was designed using the graph-visualization software Gephi 0.9.2 with data generated by string_grouper operating on the sec__edgar_company_info.csv sample data file.
string_grouper is a library that makes finding groups of similar strings within a single, or multiple, lists of
strings easy — and fast. string_grouper uses tf-idf to calculate cosine similarities
within a single list or between two lists of strings. The full process is described in the blog Super Fast String Matching in Python.
Installing
pip install string-grouper
Speed
string_grouper leverages the blazingly fast sp_matmul_rs (originally based on: sparse_dot_topn)
to calculate cosine similarities.
s = datetime.datetime.now()
matches = match_strings(names["name"], number_of_processes=15)
e = datetime.datetime.now()
diff = e - s
print(diff)
Results in:
00:17.80 On an m5 pro, where len(names) = 663 000
in other words, the library is able to perform fuzzy matching of 663 000 names in less than 18 seconds on a 2026 consumer CPU using 15 cores.
Simple Match
import pandas as pd
from string_grouper import match_strings
company_names = 'sec__edgar_company_info.csv'
companies = pd.read_csv(company_names)
# Create all matches:
matches = match_strings(companies['Company Name'])
# Look at only the non-exact matches:
matches[matches['left_Company Name'] != matches['right_Company Name']].head()
| left_index | left_Company Name | similarity | right_Company Name | right_index | |
|---|---|---|---|---|---|
| 15 | 14 | 0210, LLC | 0.870291 | 90210 LLC | 4211 |
| 167 | 165 | 1 800 MUTUALS ADVISOR SERIES | 0.931615 | 1 800 MUTUALS ADVISORS SERIES | 166 |
| 168 | 166 | 1 800 MUTUALS ADVISORS SERIES | 0.931615 | 1 800 MUTUALS ADVISOR SERIES | 165 |
| 172 | 168 | 1 800 RADIATOR FRANCHISE INC | 1 | 1-800-RADIATOR FRANCHISE INC. | 201 |
| 178 | 173 | 1 FINANCIAL MARKETPLACE SECURITIES LLC /BD | 0.949364 | 1 FINANCIAL MARKETPLACE SECURITIES, LLC | 174 |
Group Similar Strings and Find most Common
companies[["group-id", "name_deduped"]] = group_similar_strings(companies['Company Name'])
companies.groupby('name_deduped')['Line Number'].count().sort_values(ascending=False).head(10)
| name_deduped | Line Number |
|---|---|
| ADVISORS DISCIPLINED TRUST | 1747 |
| NUVEEN TAX EXEMPT UNIT TRUST SERIES 1 | 916 |
| GUGGENHEIM DEFINED PORTFOLIOS, SERIES 1200 | 652 |
| U S TECHNOLOGIES INC | 632 |
| CAPITAL MANAGEMENT LLC | 628 |
| CLAYMORE SECURITIES DEFINED PORTFOLIOS, SERIES 200 | 611 |
| E ACQUISITION CORP | 561 |
| CAPITAL PARTNERS LP | 561 |
| FIRST TRUST COMBINED SERIES 1 | 560 |
| PRINCIPAL LIFE INCOME FUNDINGS TRUST 20 | 544 |
Documentation
The documentation can be found here
Backends
The library was originally developed using the sparse_dot_topn library,
but has since been rewritten to use the sp_matmul_rs library, which is a Rust
implementation of the sparse matrix multiplication algorithm optimized with Claude Fable. To run the library with the
original sparse_dot_topn backend, set use_sp_matmul_rs to False.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file string_grouper-0.8.0.tar.gz.
File metadata
- Download URL: string_grouper-0.8.0.tar.gz
- Upload date:
- Size: 2.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cfd147f7c13ef0f181defc41b026ce468045a4c21f9aac7d20ffbd5b96d1d4e8
|
|
| MD5 |
be702db53215b4a99427b7ae43b14209
|
|
| BLAKE2b-256 |
696a59d613d3561bf1c97fbb938311ba9f49fd54cd6cb8288c46266d1d9823e3
|
Provenance
The following attestation bundles were made for string_grouper-0.8.0.tar.gz:
Publisher:
publish.yml on Bergvca/string_grouper
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
string_grouper-0.8.0.tar.gz -
Subject digest:
cfd147f7c13ef0f181defc41b026ce468045a4c21f9aac7d20ffbd5b96d1d4e8 - Sigstore transparency entry: 2256479804
- Sigstore integration time:
-
Permalink:
Bergvca/string_grouper@cfdddb2f6385b24f766d4268b4de736584d3caa7 -
Branch / Tag:
refs/tags/v0.8.0 - Owner: https://github.com/Bergvca
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@cfdddb2f6385b24f766d4268b4de736584d3caa7 -
Trigger Event:
push
-
Statement type:
File details
Details for the file string_grouper-0.8.0-py3-none-any.whl.
File metadata
- Download URL: string_grouper-0.8.0-py3-none-any.whl
- Upload date:
- Size: 31.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2607b28840c70d17b4a8f0f38e1ad050afab9c4d45cfef043a5f729e12192680
|
|
| MD5 |
4b7170e7b37414a290384b29a0d37d30
|
|
| BLAKE2b-256 |
7eee13e29384d04dfa5d006dbf25c1e7a3d9afbfc98cbbc15c5df1536b38d467
|
Provenance
The following attestation bundles were made for string_grouper-0.8.0-py3-none-any.whl:
Publisher:
publish.yml on Bergvca/string_grouper
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
string_grouper-0.8.0-py3-none-any.whl -
Subject digest:
2607b28840c70d17b4a8f0f38e1ad050afab9c4d45cfef043a5f729e12192680 - Sigstore transparency entry: 2256479810
- Sigstore integration time:
-
Permalink:
Bergvca/string_grouper@cfdddb2f6385b24f766d4268b4de736584d3caa7 -
Branch / Tag:
refs/tags/v0.8.0 - Owner: https://github.com/Bergvca
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@cfdddb2f6385b24f766d4268b4de736584d3caa7 -
Trigger Event:
push
-
Statement type: