fastumap
A UMAP and clustering library for environments where a JIT compiler at import time is not acceptable: serverless functions, small or CPU-limited containers, and autoscaled workers that pay a cold start on every new process.
umap-learn compiles its kernels with numba and LLVM the first time the package is imported. That costs seconds on a workstation and minutes on a CPU-throttled container, and it recurs on every cold start rather than once per machine. NumPy and SciPy are equally compiled code, but they ship their machine code prebuilt in the wheel, so they load immediately. fastumap reimplements the UMAP algorithm on top of them, with precompiled Rust kernels (also shipped in the wheel) for the parts that need to be fast. Nothing is left to compile at import.
| cold import | added to image | |
|---|---|---|
import umap (umap-learn) |
~21 s | ~172 MB |
import fastumap |
~12 ms | 0 MB |
Projection quality stays within 0.01 neighbour overlap of umap-learn; the measurements are below.
Installation
pip install fastumap # one wheel, native kernels included
pip install fastumap[ann] # adds an approximate kNN backend for large inputs (see below)
Usage
from fastumap import umap_project, spectral_project
xy = umap_project(x, 2) # (n, 2)
xyz = umap_project(x, 3) # (n, 3)
red = umap_project(x, 10) # ~10-D for clustering (not just 2/3)
cos = umap_project(x, 2, metric="cosine") # text / CLS embeddings
dm = umap_project(x, 2, densmap=True) # densMAP: keep dense/sparse regions distinct
metric="cosine"is recommended for encoder embeddings. Euclidean distance on unnormalised vectors is dominated by magnitude rather than by the direction that carries meaning.pca_dim=100pre-reduces very wide inputs (for example 1024-dimensional embeddings) before the nearest-neighbour search. Neighbour overlap is preserved to within about 0.01. Off by default.
umap_project also accepts n_neighbors, min_dist, spread, n_epochs,
negative_sample_rate, random_state, and chunk_count.
Supervised projection
Categorical labels passed as y inform the layout: same-label points attract, so structure that
is separable in the input space stays separable in 2-D instead of interleaving.
xy = umap_project(x, 2, y=labels) # labels: one int per point, -1 = unlabelled
xy = umap_project(x, 2, y=labels, target_weight=0.9) # lean harder on the labels
target_weight is in [0, 1] and defaults to 0.5. It trades geometry against labels: 0.0 barely
weakens inter-class edges, 1.0 cuts them entirely. Labels of -1 count as unlabelled
(semi-supervised). The default y=None is ordinary unsupervised UMAP, unchanged.
This ports umap-learn's categorical intersection. The UMAP paper only sketches the idea (arXiv:1802.03426, §7 Future Work), so umap-learn is the reference; see its supervised docs.
Clustering
cluster groups points directly, implemented here in numpy and scipy plus the native k-means
kernel, with no external clustering library. It does not compute a 2-D layout first. That layout
is an iterative SGD and takes minutes at a million points, and clustering does not need it.
from fastumap import cluster
labels = cluster(embeddings, method="kmeans", n_clusters=20) # k-means++ / Lloyd
labels = cluster(embeddings, method="spectral", n_clusters=20) # non-convex clusters
labels = cluster(embeddings, method="dbscan") # density: finds k, noise = -1
labels = cluster(embeddings, method="hdbscan") # density, no eps to pick
hdbscan is the hierarchical form of dbscan. It takes no eps, so clusters of different
densities are found together. The root of the hierarchy is never a cluster, so a dataset
containing a single group returns all -1.
kmeans and spectral take n_clusters and metric ("euclidean" or "cosine") and are
deterministic given random_state. k-means uses k-means++ seeding and Lloyd iterations with a
blocked assignment step, so the n×k distance matrix is never materialised. spectral reuses
fastumap's normalised-eigenvector machinery and runs k-means on the result. For a visualisation,
call umap_project separately, usually on a sample; clustering and visualisation are separate
operations.
Further options:
return_centroids=Truereturns aClusterResult(labels, centroids, inertia)instead of labels alone, which is what is needed to compare runs or choose k.choose_k(x, k_min=2, k_max=10)selectsn_clustersby the inertia elbow. It is a heuristic; it lands on the true k give or take one when the clusters are clear.assign_clusters(x_new, centroids)labels new points against already-fitted centroids, without re-clustering.
Above 50,000 points the native k-means kernel (matrixmultiply SIMD GEMM with rayon over row blocks, compiled into the wheel) replaces the inner loop. It clusters 1,000,000 × 128 into 64 groups in about 10 s on a 16-core machine, measured on SIFT1M, against minutes for the numpy path.
Above 200,000 points it also switches to mini-batch k-means, which samples batch_size rows per
iteration rather than sweeping all of them. Measured on SIFT1M inside a docker --cpus=2
container: 5.2 s against full Lloyd's 25.3 s, at 1.010× the inertia.
from fastumap import cluster, kmeans
labels = cluster(x, method="kmeans", n_clusters=64) # mini-batch above 200k rows
labels, cent = kmeans(x, 64, batch_size=0) # force full Lloyd at any size
labels, cent = kmeans(x, 64, batch_size=5000) # force mini-batch, with an explicit batch
The switch is on n alone and never on the core count, so the same input and seed produce the
same labels on any machine. One limitation: on perfectly separated clusters, mini-batch can leave
two centroids inside one group, measured at 4.8× worse inertia on synthetic blobs 8σ apart. Its
running-mean update cannot move a centroid back across an empty gap, where full Lloyd's recompute
can. Real embeddings overlap and do not trigger this (MNIST measures 1.013× Lloyd), and
batch_size=0 disables the switch.
Quality relative to umap-learn
fastumap stays within 0.01 neighbour overlap of umap-learn. Measured on MNIST (784-dimensional,
seed 42). These numbers are generated rather than transcribed: make bench-report regenerates them
into a dated report under
docs/bench/, which records the
machine, the load average and every package version they were measured on.
| n | method | wall time | overlap@15 | global corr |
|---|---|---|---|---|
| 5000 | fastumap | 22 s | 0.341 | 0.320 |
| 5000 | umap-learn | 79 s | 0.344 | 0.329 |
| 5000 | pca | 87 s | 0.058 | 0.523 |
| 10000 | fastumap | 46 s | 0.271 | 0.361 |
| 10000 | umap-learn | 50 s | 0.268 | 0.387 |
| 10000 | pca | 120 s | 0.038 | 0.487 |
| 20000 | fastumap | 57 s | 0.207 | 0.358 |
| 20000 | umap-learn | 69 s | 0.204 | 0.368 |
| 20000 | pca | 237 s | 0.024 | 0.504 |
The quality columns are the comparable ones. overlap@15 and global corr are deterministic; the wall times came from a loaded machine (recorded in the report) and are inflated for every method. PCA is included as a control: it preserves global structure best and local structure worst, which is the trade-off UMAP addresses.
umap-learn's times reuse the numba compilation from its first fit in the same process. A fresh
process pays roughly 20 s of compilation on every invocation, which fastumap does not pay. Above
about 50,000 points fastumap's exact O(n²) neighbour search becomes the bottleneck;
pip install fastumap[ann] addresses that, described below. Raising chunk_count (default 1)
recovers most of the remaining local-overlap gap at proportional cost, and remains deterministic.
overlap@15 is the fraction of each point's 15 input-space neighbours retained after projection, read against a random baseline. global corr is the Spearman correlation of all pairwise distances, before against after.
Speed per call
Import and cold-start cost are where fastumap differs. Per-call latency is close to umap-learn rather than better than it.
- The native kernel ships by default, so per-call speed on large batches is close to umap-learn. It runs umap-learn's in-place SGD, in Rust.
- The numpy fallback, which only an unbuilt source checkout uses, is roughly 2× slower at 1024 dimensions and n=5000 (about 40 s against 20 s). A vectorised numpy SGD cannot match numba's compiled loop.
fastumap is the appropriate choice when import and cold-start cost dominate. On per-call latency for large, high-dimensional batches it is roughly even with umap-learn on the kernel path.
BLAS thread cap under a CPU quota
Some containers cap CPU (docker --cpus, Kubernetes and EKS limits, EC2 cgroups) while still
reporting the host's full core count. BLAS then starts more threads than the quota supports and
thrashes under load.
fastumap caps the BLAS pool to the quota automatically, with no configuration. Measured under
docker --cpus=0.5 on a 16-core host with 4 workers, per-call time drops from about 70 s to about
43 s (roughly 1.6×). Where no quota is visible, nothing is capped.
This is a no-op on AWS Fargate, which meters CPU outside the cgroup, leaving the quota invisible. Fargate also reports a low core count, so BLAS does not oversubscribe there in the first place.
Server usage
umap_project is thread-safe. It holds no module-level mutable state and seeds a fresh RNG per
call, so it can be called from a worker thread (await asyncio.to_thread(umap_project, x, 2)).
Recomputing the whole layout on every request is unnecessary. Fit once, then place new points into the existing layout:
from fastumap import fit, transform
from fastumap.projection import UMAPModel
model = fit(window, 2) # cache it
xy, fit_distance = transform(model, pts, return_distances=True) # coords + per-point fit
model.to_npz("layout.npz") # persist across restarts (versioned)
model = UMAPModel.from_npz("layout.npz")
transformplaces new points without a refit. It retains about 72% of a full fit's local overlap, and the layout stays stable across requests.return_distances=Truereturns each point's distance to its nearest training neighbour, which serves as a fit score. A point 2 to 3× further out than the training mean is an extrapolation. A rising batch mean indicates that a refit is due.to_npzandfrom_npzpersist the model in a versioned numpy format. It survives releases, where a rawpicklewould break on any dataclass change.init=previousis an alternative:umap_project(window, 2, init=previous)reuses the previous coordinates so carried-over points start where they were, and new rows are placed by the caller. It refreshes a view without the fit/transform split.
A cached 5000×1024 model is about 20 MB, with training data stored as float32. Rolling windows and sparse input are not supported: densify sparse input first, and refit when the window slides.
The native accelerator
fastumap ships a compiled Rust extension, fastumap._accel, inside the wheel. It reimplements the
SGD layout kernel and the k-means kernel. The prebuilt wheels (abi3, Python 3.11+; Linux x86_64
and aarch64/Graviton, musllinux for Alpine, and Windows x86_64) carry it, so it is used
automatically. macOS has no prebuilt wheel yet and builds from the sdist, which requires a Rust
toolchain, as does any other platform without a prebuilt wheel.
The numpy path remains as a runtime fallback. If the extension is absent, for instance in an unbuilt source checkout, fastumap still works, more slowly. The SGD kernel is roughly 1.7 to 1.9× faster, with slightly higher overlap.
Because the kernel performs umap-learn's in-place walk rather than the fallback's per-epoch approximation, it produces a different, higher-quality layout for the same seed. The active path can be queried:
import fastumap
fastumap.accelerator_active() # True if the native kernel is present (the normal case)
Upgrading from fastumap-accel: that package no longer exists separately. Its kernels are built
into fastumap itself from 0.2.1 onward. Run pip install -U fastumap and drop any
fastumap-accel dependency; nothing else changes.
Large inputs (approximate kNN)
The exact neighbour search is O(n²) and dominates runtime above roughly 20,000 points.
pip install fastumap[ann] adds an approximate backend, faiss HNSW, which
ships prebuilt wheels for Linux x86_64 and aarch64, macOS and Windows, so it installs without a
compiler.
xy = umap_project(x, 2, knn="auto") # default: exact for small n, approximate above 16k
xy = umap_project(x, 2, knn="approx") # force approximate (needs fastumap[ann])
xy = umap_project(x, 2, knn="exact") # force the exact brute force
"auto", the default, switches to approximate only when the extra is installed and n is at least
16,384, so small inputs remain bit-identical. It is deterministic and retains at least 0.86
neighbour recall against exact. Measured on the neighbour search alone, at 256 dimensions:
| n | exact | approx | speedup | recall@15 |
|---|---|---|---|---|
| 20000 | 71 s | 29 s | 2.4× | 0.91 |
| 30000 | 142 s | 50 s | 2.8× | 0.86 |
fastumap.ann_available() reports whether the backend is installed.
For a given input and seed, output is bit-identical across processes and machines in the same environment. Two things change it across different environments:
- The native kernel against the numpy fallback. The wheels carry the kernel and use it by default; an unbuilt checkout falls back to numpy and differs.
fastumap[ann], which changes the neighbour graph above 16,384 points.
Both are properties of the environment. accelerator_active() and ann_available() report which
paths a run used, so a stored projection can record how it was produced.
Guarantees
Each of these is enforced by a test:
- Nothing compiles at import: no numba or llvmlite JIT. The base is NumPy, SciPy and the
pure-Python
threadpoolctl, and the native kernels ship precompiled in the wheel. - Import completes in under 200 ms, output is deterministic (bit-identical), and the API is thread-safe.
- Memory is bounded: the full n×n distance matrix is never materialised (blocked kNN), staying under 200 MB at 5000×1024.
- Any output dimension is supported, 2-D and 3-D for visualisation and around 10-D for clustering. The codebase is fully type-checked under pyright strict.
Development
make check # lint, type-check, and tests
make tox # the suite across Python 3.11, 3.12, and 3.13
make bench-report # regenerate the quality + speed tables into a dated docs/bench/ report
make fargate # import and fit timing under one CPU cap (docker --cpus=0.5 --memory=2g)
make constrained # fit and cluster timing under 0.5, 1 and 2 vCPU, into a dated report
License and attribution
fastumap is an independent reimplementation of the UMAP algorithm (McInnes, Healy & Melville, arXiv:1802.03426). It is not affiliated with or endorsed by the UMAP authors, and it is not a drop-in replacement; the public API is deliberately small. MIT licensed.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file fastumap-0.2.20.tar.gz.
File metadata
- Download URL: fastumap-0.2.20.tar.gz
- Upload date:
- Size: 54.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
913547f39e26d8511d79f816d95bd5183a2157f55b46fc44e01be379b85df035
|
|
| MD5 |
57c3df58c6300bdbc10824214530c04f
|
|
| BLAKE2b-256 |
f0877d1ddaafea874dca8762cad66cb8f9a978f5ee9730f391d503bc5d189ab4
|
File details
Details for the file fastumap-0.2.20-cp311-abi3-win_amd64.whl.
File metadata
- Download URL: fastumap-0.2.20-cp311-abi3-win_amd64.whl
- Upload date:
- Size: 263.0 kB
- Tags: CPython 3.11+, Windows x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5402627ffb9db84c4df42a4df91e0105abd1b6636cf080f320ead36a32ef0983
|
|
| MD5 |
31298af4e08b0d7b9323dd129c3599f2
|
|
| BLAKE2b-256 |
5c972d4f1153b362bd4997decc41129e73a5acb794ebefa21dd1b190cfe0ce15
|
File details
Details for the file fastumap-0.2.20-cp311-abi3-musllinux_1_2_x86_64.whl.
File metadata
- Download URL: fastumap-0.2.20-cp311-abi3-musllinux_1_2_x86_64.whl
- Upload date:
- Size: 361.3 kB
- Tags: CPython 3.11+, musllinux: musl 1.2+ x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3afda388277912f61233ce412a286bb17ca8b7e5307b62f5f84624de1101a985
|
|
| MD5 |
9d57606419d1765d4af171569ca86602
|
|
| BLAKE2b-256 |
1e4f59749135b2b488ef620a4cfe8dbb7d1b9c097beaaef5b02434c7dc844c23
|
File details
Details for the file fastumap-0.2.20-cp311-abi3-musllinux_1_2_aarch64.whl.
File metadata
- Download URL: fastumap-0.2.20-cp311-abi3-musllinux_1_2_aarch64.whl
- Upload date:
- Size: 331.2 kB
- Tags: CPython 3.11+, musllinux: musl 1.2+ ARM64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
297e66cd097696c74ea0ad6b3ab016eb7af1482f4c932781c125a76f131c1968
|
|
| MD5 |
5b8a8c3ff1316e536aafe2a13a220c27
|
|
| BLAKE2b-256 |
6c96d9c23f3f4b9ce219baffe81949ad4b9dc65bb9a8a97b6d7df667d7771102
|
File details
Details for the file fastumap-0.2.20-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.
File metadata
- Download URL: fastumap-0.2.20-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
- Upload date:
- Size: 362.6 kB
- Tags: CPython 3.11+, manylinux: glibc 2.17+ x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3a87ccb9f2fec7f9c0977ef73a452fcdddfeaf0435aae2e4760c44dff0627801
|
|
| MD5 |
9f55e81e992b7fdeebe97a4356694ad8
|
|
| BLAKE2b-256 |
11d8172497419fc210de5cd552eafef385a3ae4ec79132fb37848e07122740b0
|
File details
Details for the file fastumap-0.2.20-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.
File metadata
- Download URL: fastumap-0.2.20-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
- Upload date:
- Size: 332.4 kB
- Tags: CPython 3.11+, manylinux: glibc 2.17+ ARM64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1ed5fde464b52d7bf382b1f4be8b5f9f7d1eae332a9b6341e1b50ddcc82f5e9c
|
|
| MD5 |
5893b1e6636a8ddd365e47a15054ee2c
|
|
| BLAKE2b-256 |
af31477b031982ec203f9d17d7060c4f230d7daf2f64dc2b0a1fd6614c70f74a
|