Skip to main content

Python library for performing string similarity joins.

Project description

py_stringsimjoin

This project seeks to build a Python software package that provides scalable implementation of string similarity joins over two tables, for commonly used similarity measures such as Jaccard, Dice, cosine, overlap, overlap coefficient and edit distance. The package is free, open-source, and BSD-licensed.

Dependencies

py_stringsimjoin has been tested on each Python version between 3.7 and 3.12, inclusive.

The required dependencies to build the package are pandas 0.16.0 or higher, py_stringmatching 0.2.1 or higher, joblib, pyprind, six and a C++ compiler. For the development version, you will also need Cython.

Platforms

py_stringsimjoin has been tested on Linux, OS X and Windows.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

py-stringsimjoin-0.3.5.tar.gz (1.4 MB view details)

Uploaded Source

File details

Details for the file py-stringsimjoin-0.3.5.tar.gz.

File metadata

  • Download URL: py-stringsimjoin-0.3.5.tar.gz
  • Upload date:
  • Size: 1.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.2 CPython/3.11.7

File hashes

Hashes for py-stringsimjoin-0.3.5.tar.gz
Algorithm Hash digest
SHA256 e8b5a805c991ceb7d72ecefb3d20808dd25bdb0efac14ac415c74cad56b52ce4
MD5 135e5a97f5208226df3a29b2c7d1efbb
BLAKE2b-256 357173a8bff32a36c28e02685090b4ed0e82c8e9d18a7092a1b3d8f2de8fe5ff

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page