hdbscan

Clustering based on density with variable density clusters

These details have not been verified by PyPI

Project links

Homepage

Project description

HDBSCAN - Hierarchical Density-Based Spatial Clustering of Applications with Noise. Performs DBSCAN over varying epsilon values and integrates the result to find a clustering that gives the best stability over epsilon. This allows HDBSCAN to find clusters of varying densities (unlike DBSCAN), and be more robust to parameter selection.

In practice this means that HDBSCAN returns a good clustering straight away with little or no parameter tuning – and the primary parameter, minimum cluster size, is intuitive and easy to select.

HDBSCAN is ideal for exploratory data analysis; it’s a fast and robust algorithm that you can trust to return meaningful clusters (if there are any).

Based on the paper:: R. Campello, D. Moulavi, and J. Sander, Density-Based Clustering Based on Hierarchical Density Estimates In: Advances in Knowledge Discovery and Data Mining, Springer, pp 160-172. 2013

Notebooks comparing HDBSCAN to other clustering algorithms, explaining how HDBSCAN works and comparing performance with other python clustering implementations are available.

How to use HDBSCAN

The hdbscan package inherits from sklearn classes, and thus drops in neatly next to other sklearn clusterers with an identical calling API. Similarly it supports input in a variety of formats: an array (or pandas dataframe, or sparse matrix) of shape (num_samples x num_features); an array (or sparse matrix) giving a distance matrix between samples.

import hdbscan

clusterer = hdbscan.HDBSCAN(min_cluster_size=10)
cluster_labels = clusterer.fit_predict(data)

Performance

Significant effort has been put into making the hdbscan implementation as fast as possible. It is orders of magnitude faster than the reference implementation in Java, and is currently faster than highly optimized single linkage implementations in C and C++. version 0.6 performance can be seen in this notebook . In particular performance on low dimensional data is better than sklearn’s DBSCAN , and via support for caching with joblib, re-clustering with different parameters can be almost free.

Additional functionality

The hdbscan package comes equipped with visualization tools to help you understand your clustering results. After fitting data the clusterer object has attributes for:

The condensed cluster hierarchy
The robust single linkage cluster hierarchy
The reachability distance minimal spanning tree

All of which come equipped with methods for plotting and converting to Pandas or NetworkX for further analysis. See the notebook on how HDBSCAN works for examples and further details.

The clusterer objects also have an attribute providing cluster membership strengths, resulting in optional soft clustering (and no further compute expense)

Robust single linkage

The hdbscan package also provides support for the robust single linkage clustering algorithm of Chaudhuri and Dasgupta. As with the HDBSCAN implementation this is a high performance version of the algorithm outperforming scipy’s standard single linkage implementation. The robust single linkage hierarchy is available as an attribute of the robust single linkage clusterer, again with the ability to plot or export the hierarchy, and to extract flat clusterings at a given cut level and gamma value.

Example usage:

import hdbscan

clusterer = hdbscan.RobustSingleLinkage(cut=0.125, k=7)
cluster_labels = clusterer.fit_predict(data)
hierarchy = clusterer.cluster_hierarchy_
alt_labels = hierarchy.get_clusters(0.100, 5)
hierarchy.plot()

Based on the paper:: K. Chaudhuri and S. Dasgupta. “Rates of convergence for the cluster tree.” In Advances in Neural Information Processing Systems, 2010.

Installing

Fast install, presuming you have sklearn and all its requirements installed:

pip install hdbscan

If pip is having difficulties pulling the dependencies then we’d suggest installing the dependencies manually using anaconda followed by pulling hdbscan from pip:

conda install cython
conda install sklearn
pip install hdbscan

For a manual install get this package:

wget https://github.com/lmcinnes/hdbscan/archive/master.zip
unzip master.zip
rm master.zip
cd hdbscan-master

Install the requirements

sudo pip install -r requirements.txt

conda install sklearn cython

Install the package

python setup.py install

Licensing

The hdbscan package is 3-clause BSD licensed. Enjoy.

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

0.8.41

Dec 12, 2025

0.8.40

Nov 18, 2024

0.8.39

Oct 12, 2024

0.8.38.post2

Oct 8, 2024

0.8.38.post1

Aug 5, 2024

0.8.37

Jun 17, 2024

0.8.36

May 24, 2024

0.8.34rc1 pre-release

Nov 20, 2023

0.8.33

Jul 18, 2023

0.8.32

Jul 17, 2023

0.8.31

Jul 17, 2023

0.8.30

Jul 5, 2023

0.8.29

Oct 31, 2022

0.8.28

Feb 8, 2022

0.8.27

Feb 3, 2021

0.8.26

Mar 19, 2020

0.8.25

Mar 10, 2020

0.8.24

Dec 8, 2019

0.8.23

Oct 8, 2019

0.8.22

May 14, 2019

0.8.21

May 13, 2019

0.8.20

Apr 16, 2019

0.8.19

Jan 24, 2019

0.8.18

Sep 11, 2018

0.8.17

Sep 9, 2018

0.8.16

Sep 7, 2018

0.8.15

Jul 21, 2018

0.8.14

Jul 18, 2018

0.8.13

Apr 29, 2018

0.8.12

Jan 20, 2018

0.8.11

Nov 8, 2017

0.8.10

Mar 30, 2017

0.8.8

Mar 5, 2017

0.8.7

Feb 2, 2017

0.8.6

Feb 1, 2017

0.8.5

Jan 28, 2017

0.8.4

Jan 5, 2017

0.8.3

Nov 24, 2016

0.8.2

Sep 15, 2016

0.8.1

Aug 18, 2016

0.8

Jun 5, 2016

0.7.3

Apr 30, 2016

0.7.2

Feb 29, 2016

0.7.1

Feb 26, 2016

0.7

Feb 22, 2016

0.6.5

Dec 16, 2015

0.6.4

Dec 5, 2015

0.6.2

Dec 3, 2015

This version

0.6.1

Dec 3, 2015

0.6

Dec 2, 2015

0.5

Nov 15, 2015

0.4.2

Nov 9, 2015

0.4.1

Nov 9, 2015

0.4

Nov 8, 2015

0.3.1

Nov 5, 2015

0.3

Oct 21, 2015

0.2

Oct 20, 2015

0.1

Oct 10, 2015

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hdbscan-0.6.1.tar.gz (640.9 kB view details)

Uploaded Dec 3, 2015 Source

File details

Details for the file hdbscan-0.6.1.tar.gz.

File metadata

Download URL: hdbscan-0.6.1.tar.gz
Upload date: Dec 3, 2015
Size: 640.9 kB
Tags: Source
Uploaded using Trusted Publishing? No

File hashes

Hashes for hdbscan-0.6.1.tar.gz
Algorithm	Hash digest
SHA256	`e047518e4a718e23627710fe005da74e89461849c33df79ef7a06210a7bf90ff`
MD5	`3b5dcd55872d475e39a47231041cf692`
BLAKE2b-256	`cd7600bde253896d4c5c48f7b3cc54890f1e4e0a0d55ea1ca1f8cc1855a5e971`

See more details on using hashes here.

hdbscan 0.6.1

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

How to use HDBSCAN

Performance

Additional functionality

Robust single linkage

Installing

Licensing

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

File details

File metadata

File hashes