Skip to main content

Python library for performing string similarity joins.

Project description

py_stringsimjoin

This project seeks to build a Python software package that provides scalable implementation of string similarity joins over two tables, for commonly used similarity measures such as Jaccard, Dice, cosine, overlap, overlap coefficient and edit distance. The package is free, open-source, and BSD-licensed.

Dependencies

py_stringsimjoin has been tested on each Python version between 3.7 and 3.12, inclusive.

The required dependencies to build the package are pandas 0.16.0 or higher, py_stringmatching 0.2.1 or higher, joblib, pyprind, six and a C++ compiler. For the development version, you will also need Cython.

Platforms

py_stringsimjoin has been tested on Linux, OS X and Windows.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

py-stringsimjoin-0.3.6.tar.gz (1.4 MB view details)

Uploaded Source

File details

Details for the file py-stringsimjoin-0.3.6.tar.gz.

File metadata

  • Download URL: py-stringsimjoin-0.3.6.tar.gz
  • Upload date:
  • Size: 1.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.2 CPython/3.11.7

File hashes

Hashes for py-stringsimjoin-0.3.6.tar.gz
Algorithm Hash digest
SHA256 50629a538f9b40a32b9d612ac3ee46d11fae1d0e2923a971ed54666ab7d2f31f
MD5 ef2cc5df4a2da7e1b066fc0d19dfc5a8
BLAKE2b-256 5230c3807065671c4c780eb3d1859d70ff27d907e282258c78bb65e60fb76778

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page