Skip to main content
fuzzgpu logo

fuzzgpu

Hardware-Accelerated Fuzzy String Matching & Sequence Alignment

Cross-platform GPU compute via WebGPU (wgpu) & Multi-Core CPU parallelism with Rayon. Zero CUDA dependencies.

PyPI Version License: MIT Rust Cross Platform


Overview

fuzzgpu is a high-throughput string distance and sequence alignment engine written in Rust with native Python and WebAssembly bindings. It leverages GPU compute shaders (wgpu / WGSL) and Rayon multi-threading to accelerate large-scale batch queries and distance matrix computations across:

  • Apple Silicon (Metal)
  • Linux (Vulkan)
  • Windows (DirectX 12 / Vulkan)
  • Integrated GPUs (Intel Iris Xe, AMD Radeon)
  • WebAssembly (In-browser execution)

No NVIDIA CUDA drivers or complex toolkits required.


Benchmark Results

Hardware: Intel(R) Iris(R) Xe Graphics (Vulkan) / Intel Core i7 CPU

1. Damerau-Levenshtein Batch (1 Query × N Candidates)

Batch Size fuzzgpu rapidfuzz python-Levenshtein Speedup vs RapidFuzz
100 0.21 ms 0.36 ms N/A 1.74×
1,000 0.96 ms 3.72 ms N/A 3.88×
5,000 3.67 ms 18.44 ms N/A 5.02×
10,000 5.95 ms 37.46 ms N/A 6.30×
50,000 31.74 ms 85.87 ms N/A 2.71×

2. Levenshtein Cross-Product Matrix (cdist $N \times M$)

Utilizing dedicated 2D Grid Workgroup Shaders with $O(N + M)$ memory bandwidth:

Matrix Size Total Pairs fuzzgpu rapidfuzz python-Levenshtein Speedup vs RF Speedup vs py-Lev
10 × 10 100 0.04 ms 0.04 ms 0.05 ms 1.00× 1.31×
50 × 50 2,500 2.22 ms 0.71 ms 0.99 ms 0.32× 0.45×
100 × 100 10,000 5.29 ms 3.51 ms 4.03 ms 0.66× 0.76×
200 × 200 40,000 15.05 ms 33.58 ms 46.82 ms 2.23× 3.11×

3. Jaro-Winkler Similarity Batch

Batch Size fuzzgpu rapidfuzz python-Levenshtein Speedup vs RapidFuzz
1,000 3.10 ms 0.81 ms 1.13 ms 0.26×
5,000 6.56 ms 5.12 ms 5.78 ms 0.78×
10,000 8.56 ms 7.80 ms 9.17 ms 0.91×
50,000 33.24 ms 47.11 ms 24.86 ms 1.42×

Installation

Python

pip install fuzzgpu

Rust (Cargo.toml)

[dependencies]
fuzzgpu-core = "0.1.2"

Quickstart

import fuzzgpu
from fuzzgpu.fuzz import ratio, partial_ratio, token_sort_ratio, token_set_ratio, extract, extractOne

# 1. Classical Distance Metrics
lev = fuzzgpu.levenshtein_distance("kitten", "sitting")         # 3
dam = fuzzgpu.damerau_levenshtein_distance("ab", "ba")          # 1 (transposition-aware)
jw  = fuzzgpu.jaro_winkler_similarity("MARTHA", "MARHTA", 0.1)  # 0.9611

# 2. High-Throughput Batch Processing (Auto-dispatched to GPU/CPU)
candidates = ["hallo", "hullo", "jello", "yellow", "hello world"] * 10_000
distances  = fuzzgpu.levenshtein_batch("hello", candidates)
jw_scores  = fuzzgpu.jaro_winkler_batch("hello", candidates, prefix_weight=0.1)

# 3. 2D Cross-Product Distance Matrix (Dedicated 2D Grid Shader)
matrix = fuzzgpu.levenshtein_cdist(["abc", "def", "xyz"], ["abd", "axy", "def"])

# 4. Global Sequence Alignment (Gotoh 1982 Linear & Affine Gap Penalties)
score_linear = fuzzgpu.needleman_wunsch_score("AGTACGCA", "TATGC", match=2, mismatch=-1, gap=-2)
score_affine = fuzzgpu.needleman_wunsch_affine("AGTACGCA", "TATGC", match=2, mismatch=-1, gap_open=-3, gap_extend=-1)

# 5. RapidFuzz-Compatible Scorer & Search API
score = ratio("fuzzy was a bear", "fuzzy was a bear")          # 100.0
part  = partial_ratio("hello", "oh hello there")              # 100.0
tsr   = token_sort_ratio("new york mets", "mets new york")     # 100.0
tset  = token_set_ratio("fuzzy was a bear", "fuzzy bear")      # 100.0

# 6. Top-K Best Match Search
best  = extractOne("hellp", ["hello", "world", "help"], score_cutoff=50.0)
# Output: ("hello", 80.0, 0)

top_3 = extract("apple", ["apply", "ape", "banana", "applesauce"], score_cutoff=50.0, limit=3)

# 7. Hardware Diagnostics
print(fuzzgpu.gpu_info())
# Output: Intel(R) Iris(R) Xe Graphics (Vulkan) / Apple M2 (Metal)

Technical Architecture

fuzzgpu combines a tiered execution pipeline to balance low-latency single queries and high-throughput batch workloads:

                          ┌──────────────────────────┐
                          │     User Query / API     │
                          └─────────────┬────────────┘
                                        │
                         Batch Size / Dataset Assessment
                                        │
                ┌───────────────────────┴───────────────────────┐
                ▼                                               ▼
     Small Workloads (< 500)                         Large Batches (≥ 500)
                │                                               │
   ┌───────────────────────────┐                 ┌───────────────────────────┐
   │    Rayon Multi-Threaded   │                 │     wgpu WebGPU Compute   │
   │      CPU Parallelism      │                 │  Shaders (Metal/Vulkan)   │
   │  - Myers 1999 Bit-Vector  │                 │  - 2D Workgroup Grids     │
   │  - Zero PCIe Latency      │                 │  - Streaming Chunking     │
   └───────────────────────────┘                 └───────────────────────────┘

Key Architectural Optimizations

  1. 2D Grid Matrix Shaders (levenshtein_matrix.wgsl & jaro_matrix.wgsl): Instead of duplicating string pairs across PCIe ($O(N \cdot M)$ transfers), List A and List B are uploaded once ($O(N + M)$ memory bandwidth). Workgroups execute on a 2D grid (@workgroup_size(16, 16)).
  2. Myers (1999) Bit-Parallel CPU Engine: For strings $\le 64$ characters, computes Levenshtein edit distance using bit-vector operations with zero inner dynamic programming loops ($O(N)$ execution).
  3. Lowrance & Wagner (1975) Unrestricted Damerau-Levenshtein: Full support for character insertions, deletions, substitutions, and arbitrary transpositions.
  4. Gotoh (1982) Affine Gap Sequence Alignment: Memory-efficient 3-state recurrence ($O(N)$ auxiliary space) for bioinformatics and long-sequence alignment.
  5. Streaming Chunk Partitioner: Datasets exceeding GPU buffer limits (>128MB or >500,000 pairs) are automatically streamed in chunks to prevent VRAM overflow.

Project Structure

fuzzgpu/
├── assets/
│   └── logo.svg               # Vector brand asset
├── crates/
│   ├── fuzzgpu-core/          # Core Rust engine & compute shaders
│   │   ├── src/
│   │   │   ├── gpu.rs         # wgpu instance and device singleton
│   │   │   ├── levenshtein.rs # Levenshtein kernel & 2D matrix dispatch
│   │   │   ├── damerau.rs     # Lowrance-Wagner Damerau-Levenshtein
│   │   │   ├── needleman.rs   # Needleman-Wunsch (Linear & Affine)
│   │   │   ├── jaro.rs        # Jaro / Jaro-Winkler GPU & CPU kernels
│   │   │   ├── fuzz.rs        # Fuzzy ratio, token sort/set, extract
│   │   │   ├── simd.rs        # Myers bit-vector algorithms
│   │   │   └── shaders/       # WGSL compute shaders (1D & 2D)
│   ├── fuzzgpu-python/        # PyO3 CPython C-extension module
│   └── fuzzgpu-wasm/          # wasm-bindgen WebAssembly module
├── python/
│   └── fuzzgpu/               # Python package wrapper & typing
├── tests/
│   └── test_basic.py          # Comprehensive test suite (50 tests)
└── benchmarks/
    └── bench_compare.py       # Comparative benchmarking harness

Building from Source

Prerequisites

Build Python Extension

# Clone the repository
git clone https://github.com/Flaxmbot/fuzzgpu.git
cd fuzzgpu

# Build and install into current virtual environment
maturin develop --release

Run Tests & Benchmarks

# Run pytest verification suite
pytest tests/ -v

# Run comparative benchmark harness
python benchmarks/bench_compare.py

Build WebAssembly (Browser Target)

cd crates/fuzzgpu-wasm
wasm-pack build --target web --release

License

This project is licensed under the MIT License.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

fuzzgpu-0.1.2.tar.gz (316.5 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

fuzzgpu-0.1.2-cp39-abi3-win_amd64.whl (2.5 MB view details)

Uploaded CPython 3.9+Windows x86-64

fuzzgpu-0.1.2-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.6 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ x86-64

fuzzgpu-0.1.2-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (2.7 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ ARM64

fuzzgpu-0.1.2-cp39-abi3-macosx_11_0_arm64.whl (2.0 MB view details)

Uploaded CPython 3.9+macOS 11.0+ ARM64

fuzzgpu-0.1.2-cp39-abi3-macosx_10_12_x86_64.whl (2.1 MB view details)

Uploaded CPython 3.9+macOS 10.12+ x86-64

File details

Details for the file fuzzgpu-0.1.2.tar.gz.

File metadata

  • Download URL: fuzzgpu-0.1.2.tar.gz
  • Upload date:
  • Size: 316.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: maturin/1.14.1

File hashes

Hashes for fuzzgpu-0.1.2.tar.gz
Algorithm Hash digest
SHA256 705e356b7bc245542d1891de3dc4cc571f2b0fdbaf0f72814a22c1ab6cb44254
MD5 703de98388c288ffd5edd26464852f79
BLAKE2b-256 5a4f2b41e360731c73a28aa84631b15e639d9355b4f171e3d9066adfd11bfbf5

See more details on using hashes here.

File details

Details for the file fuzzgpu-0.1.2-cp39-abi3-win_amd64.whl.

File metadata

  • Download URL: fuzzgpu-0.1.2-cp39-abi3-win_amd64.whl
  • Upload date:
  • Size: 2.5 MB
  • Tags: CPython 3.9+, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: maturin/1.14.1

File hashes

Hashes for fuzzgpu-0.1.2-cp39-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 0b58991f1f9525e97637adc95701535f8d7157665195b649fbf61d3a352eefb1
MD5 44c9a5b97c2ab875ef052b207b657167
BLAKE2b-256 4ca14c40ed744474766d062a3f73e6537f62a8fa5c98d77a385ed6677011d65b

See more details on using hashes here.

File details

Details for the file fuzzgpu-0.1.2-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for fuzzgpu-0.1.2-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 9b7c9a1e404e83be6e9a85d975ee01a7f84bc82166f98390268d9ca3fadacdf0
MD5 02671097e3324894e579b07483934d50
BLAKE2b-256 5b145482e3ae3078a1741cb557ccaa86d23bc1f6e83c54ca68527675982bb908

See more details on using hashes here.

File details

Details for the file fuzzgpu-0.1.2-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for fuzzgpu-0.1.2-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 a77f79ed02651c0d621ef215daf66e5338ee0da19a76125a23cbd744da4348e2
MD5 f42a3fed0f91772b997ac550bc3d1ac8
BLAKE2b-256 50ce84f671233a2a149469225ac126049b57cb2a79e2ca8f8e86e8163b16b746

See more details on using hashes here.

File details

Details for the file fuzzgpu-0.1.2-cp39-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for fuzzgpu-0.1.2-cp39-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 1e36b40f66df23ef2b2d85812e4b1df8a0c16e316954dcf4fcc8e78abb622603
MD5 20a6af4330c0ec539898a31d6b12b2fc
BLAKE2b-256 c55396578d532e3adc39713a10704431a95897c6b0bff9384858ad4d3d1c412b

See more details on using hashes here.

File details

Details for the file fuzzgpu-0.1.2-cp39-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for fuzzgpu-0.1.2-cp39-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 2d2f7e94015df950cc37a60d513b53cbb4e906573fd872c46a90778c926d221e
MD5 a85f8248abaacc5650b6b8cacd8dcb76
BLAKE2b-256 d7cca205ad37e47b33149c3507fa9a2183241aeba10192adfdae8128d05950e6

See more details on using hashes here.

Release history Release notifications | RSS feed

0.4.0

6 files

0.3.0

6 files

0.2.0

6 files

0.1.8

6 files

0.1.7

6 files

0.1.6

6 files

0.1.5

6 files

0.1.4

6 files

0.1.3

6 files

This release

0.1.2 This release

6 files

0.1.1

6 files

0.1.0

6 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page