Skip to main content

Python library for performing string similarity joins.

Project description

py_stringsimjoin

This project seeks to build a Python software package that provides scalable implementation of string similarity joins over two tables, for commonly used similarity measures such as Jaccard, Dice, cosine, overlap, overlap coefficient and edit distance. The package is free, open-source, and BSD-licensed.

Dependencies

py_stringsimjoin has been tested on Python 2.7, 3.5, 3.6, and 3.7.

The required dependencies to build the package are pandas 0.16.0 or higher, py_stringmatching 0.2.1 or higher, joblib, pyprind, six and a C++ compiler. For the development version, you will also need Cython.

Platforms

py_stringsimjoin has been tested on Linux, OS X and Windows.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

py_stringsimjoin-0.3.2.tar.gz (1.1 MB view details)

Uploaded Source

File details

Details for the file py_stringsimjoin-0.3.2.tar.gz.

File metadata

  • Download URL: py_stringsimjoin-0.3.2.tar.gz
  • Upload date:
  • Size: 1.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.2.0 pkginfo/1.6.0 requests/2.24.0 setuptools/50.3.0.post20201006 requests-toolbelt/0.9.1 tqdm/4.51.0 CPython/3.7.9

File hashes

Hashes for py_stringsimjoin-0.3.2.tar.gz
Algorithm Hash digest
SHA256 a41aa3d0b3520333b7a38132fd1e6f5f9b3421be64f964e644ff7544cacbc89b
MD5 f23d9da181c77b0d49fbceba5e1ee8af
BLAKE2b-256 69f8343a7277ce5952a923302bb29ac547a2e3eab45965fdf40a7dd43ed058ef

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page