Skip to main content

betula-cluster

PyPI Python CI Python coverage 100% License: MIT Rust core · PyO3 DOI

Rust-powered, memory-bounded clustering for large embeddings & tabular streams. It compresses raw data into numerically stable BETULA microclusters, then runs the clustering head on the compressed representation — k-means · GMM (diagonal & full) · Ward · spectral · Leiden community detection · directional (von Mises–Fisher / spherical) · HDBSCAN-CF · scale-space modes · Mapper — so cost scales with the microcluster count, not N. Streaming partial_fit, a scikit-learn API, from-scratch Rust core + PyO3, no LAPACK or SciPy at runtime.

pip install betula-cluster

Verified: a 213-case Python suite at 100% wrapper coverage + 185 Rust tests, clippy -D warnings + fmt clean across all feature sets, CI on CPython 3.11–3.14 (one abi3 wheel).

At a glance — honest benchmarks

Measured against scikit-learn on StandardScaler-normalized data, each method in its own subprocess with peak RSS sampled from /proc/self/statm. Full methodology, every metric, and all tables (wins and losses) live in bench/RESULTS.md.

  • ⚡🪶 Always faster — and lighter — at scale (the unconditional win). betula labels 1 M points in 0.22 s: 13× faster than scikit-learn KMeans, 23× vs GaussianMixture, 37× vs Birch — and streams 10 M in a flat ~60 MB where an in-core KMeans needs ~5 GB (≈83× less, and the gap grows without bound). This holds for every method at every size.
  • 🎯 Quality at parity — or better — on most tasks. betula's k-means is at exact parity with scikit-learn (blobs 0.861 = 0.861); full-covariance GMM matches it on anisotropic data (0.90 vs 0.90) and beats scikit-learn's GMM on real 64-D digits (0.51 vs 0.40, via the high-dimensional covariance floor); betula-ward clusters 1 M in 0.30 s where O(N²) sklearn-ward can't run past ~10 k; and on non-convex moons & circles the spectral and HDBSCAN heads hit ARI 1.00 — spectral matching scikit-learn's SpectralClustering at 3–5× the speed.
  • 🌍 Real data, honest trade-offs. On digits betula-kmeans leads (0.57 vs 0.47); given adequate leaf resolution its diagonal GMM overtakes scikit-learn on hard covtype too (0.096 vs 0.080); it clusters full covtype (581 k rows) ~5.8× faster at matching ARI; in 784-D MNIST normalize=True beats scikit-learn (0.33 vs 0.32). Where a compression method costs some quality — raw-Euclidean k-means in 784-D, HDBSCAN on overlapping blobs — bench/RESULTS.md reports it rather than hides it.
Fit time vs N Peak memory vs N
Phase-3 clusters only the ~2 000 leaf microclusters, not the raw points, so every head finishes 1 M points in under ½ s. The CF-tree is capped by max_leaves, so streaming memory stays flat — it clusters data larger than RAM.

Why

Who it's for: practitioners clustering large embedding or tabular datasets — in batch or as a stream — who need bounded memory and the numerical stability that classic BIRCH and in-core scikit-learn don't provide together.

Clustering libraries tend to either not scale (full GMM/HDBSCAN on raw points), lose precision (classic BIRCH computes variance as SS − ‖LS‖²/n, which catastrophically cancels far from the origin), or blow up in memory (BIRCH-family subcluster explosion in high dimensions). betula-cluster addresses all three:

  • Numerically stable — clustering features (n, μ, S) via Welford / Chan updates; the covariance is PSD by construction. Classic BIRCH loses all digits near coordinate 1e7; betula does not.
  • Memory-bounded by design — the CF-tree caps its leaves (max_leaves) and rebuilds, so it never explodes; streaming memory is flat in N and clusters data larger than RAM.
  • Complete — one stable engine spanning k-means / GMM (diag & full) / Ward / spectral / Leiden community detection / HDBSCAN-style / Mapper, with streaming partial_fit, a scikit-learn API, and dataset-structure inspection.

The math (stable CF, the expected-log GMM E-step, distance derivations, relation to BIRCH/BETULA) is written up — verified symbolically and numerically — in docs/MATH.md.

When to use it

Reach for betula-cluster when the data is large or streaming, memory must stay bounded, you want fast predict on new points, or you want one numerically stable engine spanning k-means / GMM / Ward / density / topology plus dedup / outliers / representatives — especially on embeddings and tabular streams.

Use raw scikit-learn instead when N fits comfortably in RAM and you want the exact point-level algorithm with no compression: at small N the two-phase overhead removes the speed edge, and raw HDBSCAN is stronger on overlapping density. betula-cluster trades a CF-compression approximation for scale and bounded memory — if you need neither, a plain in-core clusterer is simpler.

Installation

pip install betula-cluster            # prebuilt abi3 wheels, CPython 3.11–3.14
pip install 'betula-cluster[tune]'    # + Optuna backend for memory-aware tuning

NumPy is the only runtime dependency — no SciPy, LAPACK, or BLAS. Prebuilt wheels ship for Linux (x86-64 + aarch64), macOS (Intel + Apple Silicon), and Windows (x64); one abi3 wheel covers every supported Python. Building from source needs a Rust toolchain — maturin develop --release (or pip install .) in a clone.

Quick start

import numpy as np
import betula_cluster

X = np.random.default_rng(0).normal(size=(100_000, 10))

labels = betula_cluster.fit_predict(X, n_clusters=10, method="kmeans")
labels = betula_cluster.fit_predict(X, n_clusters=0, feature="full", method="gmm-full")  # auto-k via BIC
labels = betula_cluster.fit_predict(X, n_clusters=8, method="spectral", threshold=0.0)   # non-convex / manifold
labels = betula_cluster.fit_predict(X, method="leiden", threshold=0.4)                    # graph communities; count auto-discovered
labels = betula_cluster.fit_predict(X, method="hdbscan", min_cluster_size=25)            # HDBSCAN-CF; -1 = noise
labels = betula_cluster.fit_predict(X, n_clusters=10, method="vmf")                       # directional / cosine (input auto-L2-normalized)

Streaming / out-of-core — feed chunks, finalize, predict; memory stays bounded by max_leaves:

est = betula_cluster.Betula(method="gmm", memory_budget_mb=512)
for chunk in stream_of_arrays:        # each chunk is a 2-D float32/float64 array
    est.partial_fit(chunk)
est.partial_fit()                     # finalize the global clustering over everything seen
labels = est.predict(X_query)

Robustness — the CF-tree is insertion-order sensitive, so consensus clusters several random permutations and votes, returning a consensus labelling plus a per-point stability score (any partitional head — kmeans / gmm / ward / spectral):

res = betula_cluster.consensus(X, n_clusters=10, n_runs=5, method="kmeans", n_jobs=-1)  # -1 = all cores
res.labels           # (n,) consensus label per point
res.confidence       # (n,) in [0, 1] — per-point agreement across runs (1.0 = every order agrees)
res.mean_confidence  # scalar robustness summary

Memory-aware hyperparameter tuning (tune, optional Optuna), Mapper topology (mapper), semi-supervised constraints (COP-KMeans), mixed numeric+categorical (KPrototypes), streaming density (DenStream / DbStream), the O(nnz) sparse-native path (fit_predict_sparse), CF-weighted NMF for nonnegative data (projection="weighted-nmf", or "weighted-nmf-kl" for counts), quantile sketches, scipy.sparse input, threshold="auto", soft assignment / coresets / diagnostics / drift snapshots / active-learning batches, the Rust API, and the CLI — all in the usage guide.

Capabilities

Stable core — production-ready:

  • Clustering heads — weighted k-means (Hamerly), GMM (diagonal & full covariance, BIC auto-k), exact Ward HAC, spectral (non-convex / manifold), Leiden graph community detection (auto community count, resolution / CPM, optional covariance/manifold-aware affinity), and directional spherical k-means / von Mises–Fisher mixtures for L2-normalized embeddings (cosine geometry), all over the numerically stable BETULA CF-tree.
  • Streamingpartial_fit at bounded memory (max_leaves / memory_budget_mb), EWMA decay.
  • scikit-learn APIfit / predict / fit_predict, get_params / set_params (works with Pipeline / clone / GridSearchCV); typed abi3 wheel, save / load + pickle, reusable Rust core.
  • Inspection & robustnesspredict_proba, coresets, microcluster/cluster geometry, outliers, near-duplicates, representatives, diagnostics, and consensus (per-point stability across insertion-order permutations).
  • Tuningtune: memory-aware hyperparameter search with a quality / memory / speed Pareto mode; NumPy-only, optional Optuna backend (pip install 'betula-cluster[tune]').

Experimental / evolving — useful today, API may still move:

  • Density & topology — HDBSCAN-CF (density over microclusters), scale-space Morse-persistence density-mode clustering (method="scale-space" — no k, no bandwidth), and a Mapper topological skeleton (mapper / mapper_stability).
  • Structured-covariance GMM — a three-rung Toeplitz ladder: method="gmm-toeplitz" (banded AR), "gmm-toeplitz-full" (general positive-definite Toeplitz covariance), and "gmm-toeplitz-gs" (full-order Gohberg–Semencul MLE precision): covariance-shape clustering for ordered, stationary signals (time-series windows, trajectories, sensor waveforms), well-posed where full covariance is singular (N_k ≪ d) and a diagonal model ignores neighbour correlation; the -full head captures structure beyond a low-order AR (e.g. a long-lag echo), the -gs head fits a likelihood-optimal precision with a cheaper E-step than full at large d.
  • More heads & dataDenStream / DbStream evolving-stream density, mergeable KllSketch / DdSketch quantiles, scipy.sparse (O(nnz), never densified), mixed numeric+categorical (KPrototypes), COP-KMeans constraints, robust (Huber) insertion, drift snapshots, dependency-free CLI.

Full reference: docs/FEATURES.md.

Examples

Seventeen executed, plotted notebooks — one per capability — live in examples/ (render on GitHub):

And six end-to-end use cases (each scored against ground truth):

Documentation

  • Usage guide — runnable snippets for every interface.
  • Features — full capability reference + crate architecture.
  • Math — stable CF, GMM E-step, distance derivations, relation to BIRCH/BETULA.
  • Benchmarks — methodology, every metric, all tables, honest wins & losses.
  • Design — internal design, invariants, and testing strategy.

Verified: 185 Rust unit + 4 integration tests + a 213-case Python suite at 100% wrapper coverage (Rust ≥95%, CI-enforced), clippy -D warnings + fmt clean across all feature sets, on Python 3.11–3.14 (single abi3 wheel).

Known limitations

Honest scope — inherent to a CF-compression + streaming design, not bugs:

  1. Insertion-order sensitive — like every BIRCH-family streaming method, the labels depend on the order points arrive (the parallel build differs from the serial one, as a different order would).
  2. threshold / max_leaves are real hyperparameters — they trade compression against resolution; n_rebuilds_ / threshold_ expose thrashing / over-coarsening.
  3. CF-level heads approximate raw-data clustering — Phase-3 runs on the M ≪ N microclusters; quality degrades when clusters overlap at the compression scale. Mitigation: more leaves.
  4. HDBSCAN-on-CF ≠ raw-point HDBSCAN — mass-aware HDBSCAN over microclusters: fast and close, but an approximation (weaker on overlapping blobs; see the benchmarks).
  5. The expected-log GMM optimizes a CF-level objective, not pointwise EM (a deliberate, measured choice for coarse CFs).
  6. Frequent-Directions is an approximate low-rank covariance (exact only up to its rank ).

How to cite

If betula-cluster supports your research, please cite the software and the underlying algorithms it implements. Machine-readable metadata (including the method references) lives in CITATION.cff — GitHub's "Cite this repository" renders it directly.

@software{gradina_betula_cluster,
  author  = {Gradina, Ilia},
  title   = {betula-cluster: numerically stable {BETULA} clustering with a {Rust} core},
  year    = {2026},
  version = {0.5.0},
  doi     = {10.5281/zenodo.21427331},
  license = {MIT},
  url     = {https://github.com/ilgrad/betula-cluster}
}

betula-cluster is an independent implementation (with extensions); the algorithms are due to BETULA — Lang & Schubert, Information Systems (2022), doi:10.1016/j.is.2021.101918 — building on BIRCH — Zhang, Ramakrishnan & Livny, SIGMOD (1996), doi:10.1145/233269.233324.

License

MIT © Ilia Gradina

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

betula_cluster-0.5.0.tar.gz (243.2 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

betula_cluster-0.5.0-cp311-abi3-win_amd64.whl (2.2 MB view details)

Uploaded CPython 3.11+Windows x86-64

betula_cluster-0.5.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.1 MB view details)

Uploaded CPython 3.11+manylinux: glibc 2.17+ x86-64

betula_cluster-0.5.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (1.7 MB view details)

Uploaded CPython 3.11+manylinux: glibc 2.17+ ARM64

betula_cluster-0.5.0-cp311-abi3-macosx_11_0_arm64.whl (1.8 MB view details)

Uploaded CPython 3.11+macOS 11.0+ ARM64

betula_cluster-0.5.0-cp311-abi3-macosx_10_12_x86_64.whl (2.1 MB view details)

Uploaded CPython 3.11+macOS 10.12+ x86-64

File details

Details for the file betula_cluster-0.5.0.tar.gz.

File metadata

  • Download URL: betula_cluster-0.5.0.tar.gz
  • Upload date:
  • Size: 243.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for betula_cluster-0.5.0.tar.gz
Algorithm Hash digest
SHA256 492f4fe7c8e26bc1dd924e139af9b0259b9728ea06271d432e396bf3cd80bb7d
MD5 c21e7e484cc5ea6fd906836158c799d4
BLAKE2b-256 5ae91f3d79e9633d4db3b958d6032812b65eb7e42111aa82e19bada83c912a9c

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.5.0.tar.gz:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file betula_cluster-0.5.0-cp311-abi3-win_amd64.whl.

File metadata

File hashes

Hashes for betula_cluster-0.5.0-cp311-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 4a49a3f913f415abe62dff14ce1d26b9d3b1b05e701dee5ea99d7ead24c5ec92
MD5 9266b9edd3eecca86f3993abb843a53e
BLAKE2b-256 418a984f58c3fecceccd8d06b166fdbbd74f76cde7f31f638c693132c27ffdac

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.5.0-cp311-abi3-win_amd64.whl:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file betula_cluster-0.5.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for betula_cluster-0.5.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 2d079e8af263159ce7084bc833a4ec460ae6d156c47501462afd60365c6001fc
MD5 8357256536ed0500d0749ef8332c385f
BLAKE2b-256 701cfcc64ef051edace0c828e54a8b0c9bb1b0b8f64cbf283da2276d35c7ea6a

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.5.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file betula_cluster-0.5.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for betula_cluster-0.5.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 d126157d873eb39833ce9e803c47a67194e0916c0fa0453ef4a2bfa6da3681f4
MD5 32960258ce5f264410946dd281bfed3f
BLAKE2b-256 f847d40f0a5cca5493b45388e4feec91e22ec5bf82e2aa1b9698cd5a75d19abe

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.5.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file betula_cluster-0.5.0-cp311-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for betula_cluster-0.5.0-cp311-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 503c9d2e0a0ddf2ede054aeeaf63fdad0f7dcab576ed84f75731b806137a38da
MD5 ddd2da478f51ebd3483e87b9e37427d7
BLAKE2b-256 6aefd2796c75cfe3178be0cba7e6bfabbbd49958968e513a51fba96465493437

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.5.0-cp311-abi3-macosx_11_0_arm64.whl:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file betula_cluster-0.5.0-cp311-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for betula_cluster-0.5.0-cp311-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 ccc7a8a43525241dfc668683b2bb56b1bc015aafd98ce0ef14a3fa947e639d85
MD5 a548467c45012770f79b8f51f93d673c
BLAKE2b-256 9ed34ab4cf6272e606c1d01c50af47968ae309166548a0a72a7b881f8d95d194

See more details on using hashes here.

Provenance

The following attestation bundles were made for betula_cluster-0.5.0-cp311-abi3-macosx_10_12_x86_64.whl:

Publisher: release.yml on ilgrad/betula-cluster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.6.0

6 files

This release

0.5.0 This release

6 files

0.4.0

6 files

0.3.0

6 files

0.2.0

6 files

0.1.5

6 files

0.1.4

6 files

0.1.3

6 files

0.1.2

6 files

0.1.1

6 files

0.1.0

6 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page