Skip to main content

sparse_convolution

Sparse 2D convolution in Python via Toeplitz matrix methods.

Fast when the kernel is small, the input is sparse, and/or many arrays share the same kernel.

Install

pip install sparse_convolution

Or from source:

git clone https://github.com/RichieHakim/sparse_convolution
cd sparse_convolution
pip install -e .

Usage

Single image

import numpy as np
import scipy.sparse
import sparse_convolution as sc

x = scipy.sparse.random(100, 100, density=0.01)
k = np.random.rand(5, 5)

conv = sc.Toeplitz_convolution2d(x_shape=x.shape, k=k, mode='same')
out = conv(x=x, batching=False).toarray()

Batched

Input: (n_images, H * W) sparse matrix. Output: (n_images, H_out * W_out).

x_batch = scipy.sparse.vstack([
    scipy.sparse.random(100, 100, density=0.01).reshape(1, -1)
    for _ in range(50)
]).tocsr()

conv = sc.Toeplitz_convolution2d(x_shape=(100, 100), k=k, mode='same')
out = conv(x=x_batch, batching=True)

Methods and backends

Four methods, each with selectable backends:

Method numpy numba torch
direct n/a yes n/a
precomputed yes yes yes
lazy yes n/a yes
gather_scatter yes yes yes
  • direct: Batch-parallel scatter convolution with thread-local dense buffers (numba only). For each image in parallel, scatters kernel-weighted input values into a dense accumulator spanning only that image's output bounding box, then extracts nonzeros into CSR format. Cost therefore does not grow with frame size for localized inputs (e.g. small ROIs in a large field of view), and frames over 2³¹ pixels get int64 indices. Uses a two-phase approach: a lightweight boolean counting pass (1-byte flags, no float arithmetic) determines output sizes, then the scatter pass writes directly to right-sized arrays (outputs that are exactly zero are dropped and compacted out). Interior pixels (~92-100%) skip bounds checking entirely via precomputed safe regions. O(nnz × K) per image with no init overhead. Fastest method across nearly all configurations. Requires numba.
  • precomputed: Builds a sparse Toeplitz matrix at init; fast batched matmul. Best for large batches with the same kernel when numba is not available.
  • lazy: COO broadcasting, no init cost. Best for very sparse inputs with small batches.
  • gather_scatter: Per-kernel-position scatter into a dense accumulator. General-purpose method for sparse batched inputs. Uses numba automatically when available, and falls back to numpy otherwise.

Backend selection:

  • numpy: scipy/numpy ops. Always available.
  • numba: JIT-compiled parallel loops. Fastest on CPU for batched inputs. Requires numba.
  • torch: PyTorch ops with optional GPU. Requires torch.
conv = sc.Toeplitz_convolution2d(
    x_shape=(100, 100),
    k=k,
    mode='same',
    method='direct',  # default
    backend=None,     # numba for the default direct method
)

If backend=None (default), direct uses numba. For environments without numba, choose a numpy-capable method explicitly, such as method='gather_scatter', backend='numpy' or method='precomputed', backend='numpy'.

References

Benchmarks

All benchmarks run on CPU with 1s minimum measurement time per configuration (median reported). Nine method+backend combinations compared across six scaling sweeps.

Scaling overview

Six scaling sweeps varying batch size, density, image size, and kernel size. direct+numba (brown stars) is the fastest method in nearly all regimes.

Scaling overview

Grid search: fastest method per configuration

Each cell shows the winning method and total time (init + call) for that batch size × density combination. direct+numba wins 28 of 36 configurations with an average 4.75× speedup over the second-fastest method.

Grid winners

Individual scaling curves

Batch size scaling — 100×100, 5×5, density=0.01

Batch scaling

Density scaling — 100×100, 5×5, batch=100

Density scaling

Image size scaling — 5×5, density=0.01, batch=50

Image size scaling

Kernel size scaling — 100×100, density=0.01, batch=50

Kernel size scaling

Batch scaling — high density (0.1)

Batch scaling high density

Batch scaling — very sparse (density=0.001, 200×200)

Batch scaling very sparse

Metadata

Release files for sparse-convolution 0.4.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sparse-convolution 0.4.2
File Size Uploaded
sparse_convolution-0.4.2.tar.gz 31.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sparse-convolution 0.4.2
File Interpreter ABI Platform
sparse_convolution-0.4.2-py3-none-any.whl Python 3 none any Details

Total release size: 57.1 kB

Release files / sparse_convolution-0.4.2.tar.gz

Download URL sparse_convolution-0.4.2.tar.gz
Size 31.0 kB
Tags Source
SHA-256 checksum
How to use checksums
2a95c2bfc34375e21345b95be913f3d67f6f1d040b51658a674df7b21c52737c
BLAKE2b-256 checksum
How to use checksums
092c100eee567d78cbdb1c4ba76ee6486975bb8eeacbf1d6bad6f12cd9012662
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release files / sparse_convolution-0.4.2-py3-none-any.whl

Download URL sparse_convolution-0.4.2-py3-none-any.whl
Size 26.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
132e953d6b039a160dc4058d15ec46c4a9e0defbff25c54b906f1f8c92142ceb
BLAKE2b-256 checksum
How to use checksums
0606220c5b4b1cbc29b38262926fad30fa1f619d58ebcf9e2ae3873b0b693d92
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.2 This release

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.2.0

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page