Skip to main content
Version License Test Status Test Coverage Code Health

kmodes

Description

Python implementations of the k-modes and k-prototypes clustering algorithms. Relies on numpy for a lot of the heavy lifting.

k-modes is used for clustering categorical variables. It defines clusters based on the number of matching categories between data points. (This is in contrast to the more well-known k-means algorithm, which clusters numerical data based on Euclidean distance.) The k-prototypes algorithm combines k-modes and k-means and is able to cluster mixed numerical / categorical data.

Implemented are:

The code is modeled after the clustering algorithms in scikit-learn and has the same familiar interface.

I would love to have more people play around with this and give me feedback on my implementation. If you come across any issues in running or installing kmodes, please submit a bug report.

Enjoy!

Installation

kmodes can be installed using pip:

pip install kmodes

To upgrade to the latest version (recommended), run it like this:

pip install --upgrade kmodes

Alternatively, you can build the latest development version from source:

git clone https://github.com/nicodv/kmodes.git
cd kmodes
python setup.py install

Usage

import numpy as np
from kmodes import kmodes

# random categorical data
data = np.random.choice(20, (100, 10))

km = kmodes.KModes(n_clusters=4, init='Huang', n_init=5, verbose=1)

clusters = km.fit_predict(data)

# Print the cluster centroids
print(km.cluster_centroids_)

More simple usage examples of both k-modes (‘soybean.py’) and k-prototypes (‘stocks.py’) are included in the examples directory.

Missing / unseen data

The k-modes algorithm accepts np.NaN values as missing values in the X matrix. However, users are strongly suggested to consider filling in the missing data themselves in a way that makes sense for the problem at hand. This is especially important in case of many missing values.

The k-modes algorithm currently handles missing data as follows. When fitting the model, np.NaN values are encoded into their own category (let’s call it “unknown values”). When predicting, the model treats any values in X that (1) it has not seen before during training, or (2) are missing, as being a member of the “unknown values” category. Simply put, the algorithm treats any missing / unseen data as matching with each other but mismatching with non-missing / seen data when determining similarity between points.

The k-prototypes also accepts np.NaN values as missing values for the categorical variables, but does not accept missing values for the numerical values. It is up to the user to come up with a way of handling these missing data that is appropriate for the problem at hand.

References

[HUANG97] (1,2)

Huang, Z.: Clustering large data sets with mixed numeric and categorical values, Proceedings of the First Pacific Asia Knowledge Discovery and Data Mining Conference, Singapore, pp. 21-34, 1997.

[HUANG98]

Huang, Z.: Extensions to the k-modes algorithm for clustering large data sets with categorical values, Data Mining and Knowledge Discovery 2(3), pp. 283-304, 1998.

[CAO09]

Cao, F., Liang, J, Bai, L.: A new initialization method for categorical data clustering, Expert Systems with Applications 36(7), pp. 10223-10228., 2009.

Metadata

Release files for kmodes 0.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for kmodes 0.6
File Size Uploaded
kmodes-0.6.tar.gz 11.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for kmodes 0.6
File Interpreter ABI Platform
kmodes-0.6-py2.py3-none-any.whl Python 2, Python 3 none any Details

Total release size: 27.4 kB

Release files / kmodes-0.6.tar.gz

Download URL kmodes-0.6.tar.gz
Size 11.4 kB
Tags Source
SHA-256 checksum
How to use checksums
14eab4ae818e143177b4e195d45ef260a6789f2c9f889b1a1e07cbf108e78be4
BLAKE2b-256 checksum
How to use checksums
44eb0b07bb150b8b725bfe997b4a1199c095191a1f433da1b0a10b04e2be497a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No

Release files / kmodes-0.6-py2.py3-none-any.whl

Download URL kmodes-0.6-py2.py3-none-any.whl
Size 16.0 kB
Tags Python 2 Python 3
SHA-256 checksum
How to use checksums
d5ba99bd39b1452e81de551cdfca689c05fe2b8db927831e1c9609c1353d3e80
BLAKE2b-256 checksum
How to use checksums
cbd1a025c0ef91af63f77b2e44b7bce0c1380c752a94aa42aa2fc9715b4d7ad0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No

Release history Release notifications | RSS feed

0.12.1

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.2

2 release files

0.10.1

2 release files

0.9

2 release files

0.8

2 release files

0.7

2 release files

This release

0.6 This release

2 release files

0.5

2 release files

0.4

2 release files

0.2

2 release files

0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page