Skip to main content

betula-cluster

PyPI Python CI Python coverage 100% License: MIT Rust core · PyO3 DOI

Rust-powered, memory-bounded clustering for large embeddings & tabular streams. It compresses raw data into numerically stable BETULA microclusters, then runs the clustering head on the compressed representation — k-means · GMM (diagonal & full) · Ward · spectral · Leiden community detection · directional (von Mises–Fisher / spherical) · HDBSCAN-CF · scale-space modes · Mapper — so cost scales with the microcluster count, not N. Streaming partial_fit, a scikit-learn API, from-scratch Rust core + PyO3, no LAPACK or SciPy at runtime.

pip install betula-cluster

Verified: a 202-case Python suite at 100% wrapper coverage + 177 Rust tests, clippy -D warnings + fmt clean across all feature sets, CI on CPython 3.11–3.14 (one abi3 wheel).

At a glance — honest benchmarks

Measured against scikit-learn on StandardScaler-normalized data, each method in its own subprocess with peak RSS sampled from /proc/self/statm. Full methodology, every metric, and all tables (wins and losses) live in bench/RESULTS.md.

  • ⚡🪶 Always faster — and lighter — at scale (the unconditional win). betula labels 1 M points in 0.22 s: 13× faster than scikit-learn KMeans, 23× vs GaussianMixture, 37× vs Birch — and streams 10 M in a flat ~60 MB where an in-core KMeans needs ~5 GB (≈83× less, and the gap grows without bound). This holds for every method at every size.
  • 🎯 Quality at parity — or better — on most tasks. betula's k-means is at exact parity with scikit-learn (blobs 0.861 = 0.861); full-covariance GMM matches it on anisotropic data (0.90 vs 0.90) and beats scikit-learn's GMM on real 64-D digits (0.51 vs 0.40, via the high-dimensional covariance floor); betula-ward clusters 1 M in 0.30 s where O(N²) sklearn-ward can't run past ~10 k; and on non-convex moons & circles the spectral and HDBSCAN heads hit ARI 1.00 — spectral matching scikit-learn's SpectralClustering at 3–5× the speed.
  • 🌍 Real data, honest trade-offs. On digits betula-kmeans leads (0.57 vs 0.47); given adequate leaf resolution its diagonal GMM overtakes scikit-learn on hard covtype too (0.096 vs 0.080); it clusters full covtype (581 k rows) ~5.8× faster at matching ARI; in 784-D MNIST normalize=True beats scikit-learn (0.33 vs 0.32). Where a compression method costs some quality — raw-Euclidean k-means in 784-D, HDBSCAN on overlapping blobs — bench/RESULTS.md reports it rather than hides it.
Fit time vs N Peak memory vs N
Phase-3 clusters only the ~2 000 leaf microclusters, not the raw points, so every head finishes 1 M points in under ½ s. The CF-tree is capped by max_leaves, so streaming memory stays flat — it clusters data larger than RAM.

Why

Who it's for: practitioners clustering large embedding or tabular datasets — in batch or as a stream — who need bounded memory and the numerical stability that classic BIRCH and in-core scikit-learn don't provide together.

Clustering libraries tend to either not scale (full GMM/HDBSCAN on raw points), lose precision (classic BIRCH computes variance as SS − ‖LS‖²/n, which catastrophically cancels far from the origin), or blow up in memory (BIRCH-family subcluster explosion in high dimensions). betula-cluster addresses all three:

  • Numerically stable — clustering features (n, μ, S) via Welford / Chan updates; the covariance is PSD by construction. Classic BIRCH loses all digits near coordinate 1e7; betula does not.
  • Memory-bounded by design — the CF-tree caps its leaves (max_leaves) and rebuilds, so it never explodes; streaming memory is flat in N and clusters data larger than RAM.
  • Complete — one stable engine spanning k-means / GMM (diag & full) / Ward / spectral / Leiden community detection / HDBSCAN-style / Mapper, with streaming partial_fit, a scikit-learn API, and dataset-structure inspection.

The math (stable CF, the expected-log GMM E-step, distance derivations, relation to BIRCH/BETULA) is written up — verified symbolically and numerically — in docs/MATH.md.

When to use it

Reach for betula-cluster when the data is large or streaming, memory must stay bounded, you want fast predict on new points, or you want one numerically stable engine spanning k-means / GMM / Ward / density / topology plus dedup / outliers / representatives — especially on embeddings and tabular streams.

Use raw scikit-learn instead when N fits comfortably in RAM and you want the exact point-level algorithm with no compression: at small N the two-phase overhead removes the speed edge, and raw HDBSCAN is stronger on overlapping density. betula-cluster trades a CF-compression approximation for scale and bounded memory — if you need neither, a plain in-core clusterer is simpler.

Installation

pip install betula-cluster            # prebuilt abi3 wheels, CPython 3.11–3.14
pip install 'betula-cluster[tune]'    # + Optuna backend for memory-aware tuning

NumPy is the only runtime dependency — no SciPy, LAPACK, or BLAS. Prebuilt wheels ship for Linux (x86-64 + aarch64), macOS (Intel + Apple Silicon), and Windows (x64); one abi3 wheel covers every supported Python. Building from source needs a Rust toolchain — maturin develop --release (or pip install .) in a clone.

Quick start

import numpy as np
import betula_cluster

X = np.random.default_rng(0).normal(size=(100_000, 10))

labels = betula_cluster.fit_predict(X, n_clusters=10, method="kmeans")
labels = betula_cluster.fit_predict(X, n_clusters=0, feature="full", method="gmm-full")  # auto-k via BIC
labels = betula_cluster.fit_predict(X, n_clusters=8, method="spectral", threshold=0.0)   # non-convex / manifold
labels = betula_cluster.fit_predict(X, method="leiden", threshold=0.4)                    # graph communities; count auto-discovered
labels = betula_cluster.fit_predict(X, method="hdbscan", min_cluster_size=25)            # HDBSCAN-CF; -1 = noise
labels = betula_cluster.fit_predict(X, n_clusters=10, method="vmf")                       # directional / cosine (input auto-L2-normalized)

Streaming / out-of-core — feed chunks, finalize, predict; memory stays bounded by max_leaves:

est = betula_cluster.Betula(method="gmm", memory_budget_mb=512)
for chunk in stream_of_arrays:        # each chunk is a 2-D float32/float64 array
    est.partial_fit(chunk)
est.partial_fit()                     # finalize the global clustering over everything seen
labels = est.predict(X_query)

Robustness — the CF-tree is insertion-order sensitive, so consensus clusters several random permutations and votes, returning a consensus labelling plus a per-point stability score (any partitional head — kmeans / gmm / ward / spectral):

res = betula_cluster.consensus(X, n_clusters=10, n_runs=5, method="kmeans", n_jobs=-1)  # -1 = all cores
res.labels           # (n,) consensus label per point
res.confidence       # (n,) in [0, 1] — per-point agreement across runs (1.0 = every order agrees)
res.mean_confidence  # scalar robustness summary

Memory-aware hyperparameter tuning (tune, optional Optuna), Mapper topology (mapper), semi-supervised constraints (COP-KMeans), mixed numeric+categorical (KPrototypes), streaming density (DenStream / DbStream), the O(nnz) sparse-native path (fit_predict_sparse), quantile sketches, scipy.sparse input, threshold="auto", soft assignment / coresets / diagnostics / drift snapshots / active-learning batches, the Rust API, and the CLI — all in the usage guide.

Capabilities

Stable core — production-ready:

  • Clustering heads — weighted k-means (Hamerly), GMM (diagonal & full covariance, BIC auto-k), exact Ward HAC, spectral (non-convex / manifold), Leiden graph community detection (auto community count, resolution / CPM, optional covariance/manifold-aware affinity), and directional spherical k-means / von Mises–Fisher mixtures for L2-normalized embeddings (cosine geometry), all over the numerically stable BETULA CF-tree.
  • Streamingpartial_fit at bounded memory (max_leaves / memory_budget_mb), EWMA decay.
  • scikit-learn APIfit / predict / fit_predict, get_params / set_params (works with Pipeline / clone / GridSearchCV); typed abi3 wheel, save / load + pickle, reusable Rust core.
  • Inspection & robustnesspredict_proba, coresets, microcluster/cluster geometry, outliers, near-duplicates, representatives, diagnostics, and consensus (per-point stability across insertion-order permutations).
  • Tuningtune: memory-aware hyperparameter search with a quality / memory / speed Pareto mode; NumPy-only, optional Optuna backend (pip install 'betula-cluster[tune]').

Experimental / evolving — useful today, API may still move:

  • Density & topology — HDBSCAN-CF (density over microclusters), scale-space Morse-persistence density-mode clustering (method="scale-space" — no k, no bandwidth), and a Mapper topological skeleton (mapper / mapper_stability).
  • Structured-covariance GMMmethod="gmm-toeplitz" (banded AR) and "gmm-toeplitz-full" (general positive-definite Toeplitz): covariance-shape clustering for ordered, stationary signals (time-series windows, trajectories, sensor waveforms), well-posed where full covariance is singular (N_k ≪ d) and a diagonal model ignores neighbour correlation; the general head also captures structure beyond a low-order AR (e.g. a long-lag echo).
  • More heads & dataDenStream / DbStream evolving-stream density, mergeable KllSketch / DdSketch quantiles, scipy.sparse (O(nnz), never densified), mixed numeric+categorical (KPrototypes), COP-KMeans constraints, robust (Huber) insertion, drift snapshots, dependency-free CLI.

Full reference: docs/FEATURES.md.

Examples

Sixteen executed, plotted notebooks — one per capability — live in examples/ (render on GitHub):

And six end-to-end use cases (each scored against ground truth):

Documentation

  • Usage guide — runnable snippets for every interface.
  • Features — full capability reference + crate architecture.
  • Math — stable CF, GMM E-step, distance derivations, relation to BIRCH/BETULA.
  • Benchmarks — methodology, every metric, all tables, honest wins & losses.
  • Design — internal design, invariants, and testing strategy.

Verified: 177 Rust unit + 4 integration tests + a 202-case Python suite at 100% wrapper coverage (Rust ≥95%, CI-enforced), clippy -D warnings + fmt clean across all feature sets, on Python 3.11–3.14 (single abi3 wheel).

Known limitations

Honest scope — inherent to a CF-compression + streaming design, not bugs:

  1. Insertion-order sensitive — like every BIRCH-family streaming method, the labels depend on the order points arrive (the parallel build differs from the serial one, as a different order would).
  2. threshold / max_leaves are real hyperparameters — they trade compression against resolution; n_rebuilds_ / threshold_ expose thrashing / over-coarsening.
  3. CF-level heads approximate raw-data clustering — Phase-3 runs on the M ≪ N microclusters; quality degrades when clusters overlap at the compression scale. Mitigation: more leaves.
  4. HDBSCAN-on-CF ≠ raw-point HDBSCAN — mass-aware HDBSCAN over microclusters: fast and close, but an approximation (weaker on overlapping blobs; see the benchmarks).
  5. The expected-log GMM optimizes a CF-level objective, not pointwise EM (a deliberate, measured choice for coarse CFs).
  6. Frequent-Directions is an approximate low-rank covariance (exact only up to its rank ).

How to cite

If betula-cluster supports your research, please cite the software and the underlying algorithms it implements. Machine-readable metadata (including the method references) lives in CITATION.cff — GitHub's "Cite this repository" renders it directly.

@software{gradina_betula_cluster,
  author  = {Gradina, Ilia},
  title   = {betula-cluster: numerically stable {BETULA} clustering with a {Rust} core},
  year    = {2026},
  version = {0.2.0},
  doi     = {10.5281/zenodo.21427331},
  license = {MIT},
  url     = {https://github.com/ilgrad/betula-cluster}
}

betula-cluster is an independent implementation (with extensions); the algorithms are due to BETULA — Lang & Schubert, Information Systems (2022), doi:10.1016/j.is.2021.101918 — building on BIRCH — Zhang, Ramakrishnan & Livny, SIGMOD (1996), doi:10.1145/233269.233324.

License

MIT © Ilia Gradina

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

betula_cluster-0.3.0.tar.gz (232.7 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

betula_cluster-0.3.0-cp311-abi3-win_amd64.whl (2.0 MB view details)

Uploaded CPython 3.11+Windows x86-64

betula_cluster-0.3.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.0 MB view details)

Uploaded CPython 3.11+manylinux: glibc 2.17+ x86-64

betula_cluster-0.3.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (1.6 MB view details)

Uploaded CPython 3.11+manylinux: glibc 2.17+ ARM64

betula_cluster-0.3.0-cp311-abi3-macosx_11_0_arm64.whl (1.7 MB view details)

Uploaded CPython 3.11+macOS 11.0+ ARM64

betula_cluster-0.3.0-cp311-abi3-macosx_10_12_x86_64.whl (2.0 MB view details)

Uploaded CPython 3.11+macOS 10.12+ x86-64

File details

Details for the file betula_cluster-0.3.0.tar.gz.

File metadata

  • Download URL: betula_cluster-0.3.0.tar.gz
  • Upload date:
  • Size: 232.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for betula_cluster-0.3.0.tar.gz
Algorithm Hash digest
SHA256 6b04d7694772db8406a7381e906e737da918e0275fe92baa84bbfe8a1bea084a
MD5 178b446a187b45a3aeeb874371e824a3
BLAKE2b-256 d922d2138023da81620cede99156c538ccd223f0ae6d83824fe8ac80fe860de0

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.3.0.tar.gz:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file betula_cluster-0.3.0-cp311-abi3-win_amd64.whl.

File metadata

File hashes

Hashes for betula_cluster-0.3.0-cp311-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 9c712c0f5d2559400b1412e083484ab7c9b8445bdd6ae9108c57417263a95b57
MD5 54937638c27b100b32442b5b9d2b10dc
BLAKE2b-256 3fb37a9e532e447f7137c3272d680233179284b9852db7f22c984590d50d7ed8

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.3.0-cp311-abi3-win_amd64.whl:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file betula_cluster-0.3.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for betula_cluster-0.3.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 f9dbb967c96e87ac2781f9e8c9eb836d06ecde075ab9a920255006a60856e866
MD5 020fd3ff4238700f95da46a0a671f7ec
BLAKE2b-256 99977618600cc962997daff3bb605a4357fc7d25e47f8239a3f5acd927c26ad8

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.3.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file betula_cluster-0.3.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for betula_cluster-0.3.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 7f8dea227cc007bab38c2ca08679b7366b2b65b9b64ef01db533854b7f7a6997
MD5 ce53fa0b985f9cbc3bd3c454b50e0212
BLAKE2b-256 0d6a589b21e264e7175f1edbd009619f09aa52b1102d4516e0be55331b01a6b9

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.3.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file betula_cluster-0.3.0-cp311-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for betula_cluster-0.3.0-cp311-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 61931062b5778528d462f257589964953b7b527fabc0cf5e24a51c1a8a6c5de5
MD5 cff963b973a295f36f8c0297d57eee72
BLAKE2b-256 3878c6a6440af1a430e8722eb041dfcb3821792cd300db95ac46d5ecfdf46f86

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.3.0-cp311-abi3-macosx_11_0_arm64.whl:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file betula_cluster-0.3.0-cp311-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for betula_cluster-0.3.0-cp311-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 24930c26dfd5f821ac22907d5039419035c601ff12b4d6f0bf414e0e721e53cd
MD5 41d727ed22a18cbf010aca27c79bdd1c
BLAKE2b-256 9e9d6fb9314c6bea0dc9d2105320c11528250f85c22175f94f4371dcd0327e1d

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.3.0-cp311-abi3-macosx_10_12_x86_64.whl:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.6.0

6 files

0.5.0

6 files

0.4.0

6 files

This release

0.3.0 This release

6 files

0.2.0

6 files

0.1.5

6 files

0.1.4

6 files

0.1.3

6 files

0.1.2

6 files

0.1.1

6 files

0.1.0

6 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page