Skip to main content

Clustering based on density with variable density clusters

Project description

HDBSCAN - Hierarchical Density-Based Spatial Clustering of Applications with Noise. Performs DBSCAN over varying epsilon values and integrates the result to find a clustering that gives the best stability over epsilon. This allows HDBSCAN to find clusters of varying densities (unlike DBSCAN), and be more robust to parameter selection.

Based on the paper:
R. Campello, D. Moulavi, and J. Sander, Density-Based Clustering Based on Hierarchical Density Estimates In: Advances in Knowledge Discovery and Data Mining, Springer, pp 160-172. 2013

Notebooks comparing HDBSCAN to other clustering algorithms, and explaining how HDBSCAN works are available.

How to use HDBSCAN

The hdbscan package inherits from sklearn classes, and thus drops in neatly next to other sklearn clusterers with an identical calling API. Similarly it supports input in a variety of formats: an array (or pandas dataframe, or sparse matrix) of shape (num_samples x num_features); an array (or sparse matrix) giving a distance matrix between samples.

import hdbscan

clusterer = hdbscan.HDBSCAN(min_cluster_size=10)
cluster_labels = clusterer.fit_predict(data)

Note that clustering larger datasets will require significant memory (as with any algorithm that needs all pairwise distances). Support for low memory/better scaling is planned but not yet implemented.


Fast install

pip install hdbscan

For a manual install get this package:

cd hdbscan-master

Install the requirements

sudo pip install -r requirements.txt

Install the package

python install


The hdbscan package is BSD licensed. Enjoy.

Project details

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Files for hdbscan, version 0.2
Filename, size File type Python version Upload date Hashes
Filename, size hdbscan-0.2.tar.gz (140.9 kB) File type Source Python version None Upload date Hashes View hashes

Supported by

Elastic Elastic Search Pingdom Pingdom Monitoring Google Google BigQuery Sentry Sentry Error logging AWS AWS Cloud computing DataDog DataDog Monitoring Fastly Fastly CDN SignalFx SignalFx Supporter DigiCert DigiCert EV certificate StatusPage StatusPage Status page