sparse_convolution
Sparse 2D convolution in Python via Toeplitz matrix methods.
Fast when the kernel is small, the input is sparse, and/or many arrays share the same kernel.
Install
pip install sparse_convolution
Or from source:
git clone https://github.com/RichieHakim/sparse_convolution
cd sparse_convolution
pip install -e .
Usage
Single image
import numpy as np
import scipy.sparse
import sparse_convolution as sc
x = scipy.sparse.random(100, 100, density=0.01)
k = np.random.rand(5, 5)
conv = sc.Toeplitz_convolution2d(x_shape=x.shape, k=k, mode='same')
out = conv(x=x, batching=False).toarray()
Batched
Input: (n_images, H * W) sparse matrix. Output: (n_images, H_out * W_out).
x_batch = scipy.sparse.vstack([
scipy.sparse.random(100, 100, density=0.01).reshape(1, -1)
for _ in range(50)
]).tocsr()
conv = sc.Toeplitz_convolution2d(x_shape=(100, 100), k=k, mode='same')
out = conv(x=x_batch, batching=True)
Methods and backends
Four methods, each with selectable backends:
| Method | numpy | numba | torch |
|---|---|---|---|
direct |
n/a | yes | n/a |
precomputed |
yes | yes | yes |
lazy |
yes | n/a | yes |
gather_scatter |
yes | yes | yes |
direct: Batch-parallel scatter convolution with thread-local dense buffers (numba only). For each image in parallel, scatters kernel-weighted input values into a dense accumulator spanning only that image's output bounding box, then extracts nonzeros into CSR format. Cost therefore does not grow with frame size for localized inputs (e.g. small ROIs in a large field of view), and frames over 2³¹ pixels get int64 indices. Uses a two-phase approach: a lightweight boolean counting pass (1-byte flags, no float arithmetic) determines exact output sizes, then the scatter pass writes directly to right-sized arrays with zero waste. Interior pixels (~92-100%) skip bounds checking entirely via precomputed safe regions. O(nnz × K) per image with no init overhead. Fastest method across nearly all configurations. Requiresnumba.precomputed: Builds a sparse Toeplitz matrix at init; fast batched matmul. Best for large batches with the same kernel when numba is not available.lazy: COO broadcasting, no init cost. Best for very sparse inputs with small batches.gather_scatter: Per-kernel-position scatter into a dense accumulator. General-purpose method for sparse batched inputs. Usesnumbaautomatically when available, and falls back tonumpyotherwise.
Backend selection:
numpy: scipy/numpy ops. Always available.numba: JIT-compiled parallel loops. Fastest on CPU for batched inputs. Requiresnumba.torch: PyTorch ops with optional GPU. Requirestorch.
conv = sc.Toeplitz_convolution2d(
x_shape=(100, 100),
k=k,
mode='same',
method='direct', # default
backend=None, # numba for the default direct method
)
If backend=None (default), direct uses numba. For environments without numba, choose a numpy-capable method explicitly, such as method='gather_scatter', backend='numpy' or method='precomputed', backend='numpy'.
References
- Toeplitz convolution: stackoverflow.com/a/51865516, alisaaalehi/convolution_as_multiplication
- 1D convolution matrix: scipy.linalg.convolution_matrix
Benchmarks
All benchmarks run on CPU with 1s minimum measurement time per configuration (median reported). Nine method+backend combinations compared across six scaling sweeps.
Scaling overview
Six scaling sweeps varying batch size, density, image size, and kernel size. direct+numba (brown stars) is the fastest method in nearly all regimes.
Grid search: fastest method per configuration
Each cell shows the winning method and total time (init + call) for that batch size × density combination. direct+numba wins 28 of 36 configurations with an average 4.75× speedup over the second-fastest method.
Individual scaling curves
Batch size scaling — 100×100, 5×5, density=0.01
Density scaling — 100×100, 5×5, batch=100
Image size scaling — 5×5, density=0.01, batch=50
Kernel size scaling — 100×100, density=0.01, batch=50
Batch scaling — high density (0.1)
Batch scaling — very sparse (density=0.001, 200×200)
Metadata
Release files for sparse-convolution 0.4.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sparse_convolution-0.4.1.tar.gz | 30.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sparse_convolution-0.4.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 56.7 kB
Release files / sparse_convolution-0.4.1.tar.gz
| Download URL | sparse_convolution-0.4.1.tar.gz |
|---|---|
| Size | 30.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7a19d5d1994f78f1b8286029cc125338bd82439775fa5ffb8031c9ca2e2d4eae
|
|
BLAKE2b-256 checksum How to use checksums |
d87d69504878c517ac9cce8ceb1635d70bebca575ac2052f84607c0903c5b1bc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency logRelease files / sparse_convolution-0.4.1-py3-none-any.whl
| Download URL | sparse_convolution-0.4.1-py3-none-any.whl |
|---|---|
| Size | 26.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
38ec8dd0be366f28f04445eaa0e25a13fc8905016bfe21df36072814f1b22634
|
|
BLAKE2b-256 checksum How to use checksums |
681d2b6b3b7de371f59c0a5c6fc069900c0db37ca7bf47039001812cddba9b49
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency log