GPU-first co-occurrence matrix builder for large text corpora
Project description
fast_cooc
GPU-first co-occurrence matrix builder for large text corpora. Uses CUDA kernels via CuPy to count word co-occurrences, with automatic dense/sparse mode selection based on vocabulary size and available VRAM.
Installation
pip install .
For HuggingFace tokenizer support:
pip install ".[hf]"
Requirements: CUDA-capable GPU, Python 3.10+, CUDA 12.x toolkit.
Quickstart
CLI
# Wikipedia with default settings (auto mode, regex tokenizer)
fast-cooc --dataset nlp/wikipedia --workdir wiki_out
# Small vocab → auto-selects dense mode (fast atomicAdd path)
fast-cooc --dataset nlp/wikipedia --max-vocab 30000 --workdir wiki_out
# Large vocab → auto-selects sparse mode
fast-cooc --dataset nlp/wikipedia --max-vocab 300000 --workdir wiki_out
# Force a specific mode
fast-cooc --dataset nlp/wikipedia --mode dense --max-vocab 30000
# Use a HuggingFace tokenizer
fast-cooc --dataset nlp/wikipedia --tokenizer hf --hf-model bert-base-uncased
Python API
from fast_cooc import run_warp_pipeline, VocabConfig, CountConfig
out_path = run_warp_pipeline(
dataset_id="nlp/wikipedia",
workdir="wiki_out",
row_limit=100_000, # None for full dataset
vocab_cfg=VocabConfig(min_freq=10, max_vocab=30_000),
count_cfg=CountConfig(window=5, max_vram_gb=20.0, mode="auto"),
)
# out_path is either wiki_out/embeddings.npy (dense) or wiki_out/cooc_final.npz (sparse)
Using the GPU kernel directly
import numpy as np
from fast_cooc import count_cooc_gpu, collapse
# corpus_ids: int32 array where each value is a word index (or -1 for OOV)
corpus_ids = np.array([0, 1, 2, 3, 1, 0, 2, 3, 1, 2], dtype=np.int32)
n_vocab = 4
# Dense mode — returns a CuPy (n_vocab, n_vocab) float32 array on GPU
matrix = count_cooc_gpu(corpus_ids, n_vocab, window=2, mode="dense")
# Collapse: log1p → center columns → L2 normalize rows → CPU numpy array
embeddings = collapse(matrix)
# Sparse mode — returns a scipy.sparse.csr_matrix
sparse_matrix = count_cooc_gpu(corpus_ids, n_vocab, window=2, mode="sparse")
Custom tokenizers
from fast_cooc import run_warp_pipeline, RegexTokenizer, HFTokenizer, make_tokenizer
# Regex (default)
tok = RegexTokenizer(pattern=r"[A-Za-z0-9_']+", lowercase=True)
# HuggingFace subword tokenizer
tok = HFTokenizer(model_name="bert-base-uncased")
# Factory function
tok = make_tokenizer("regex")
tok = make_tokenizer("hf", hf_model="bert-base-uncased")
tok = make_tokenizer("callable", fn=lambda text: text.lower().split())
out = run_warp_pipeline("nlp/wikipedia", "wiki_out", tokenizer=tok)
Dense vs Sparse mode
Dense (atomicAdd) |
Sparse (cp.unique + triplets) |
|
|---|---|---|
| VRAM | V * V * 4 bytes (30k vocab = 3.6 GB) |
Just chunk buffers |
| Speed | Fastest — no GPU sort, no CPU roundtrips | Slower — GPU radix sort per chunk |
| Vocab limit | ~50k on a 20 GB card | Unlimited |
| Output | embeddings.npy (after collapse) |
cooc_final.npz (raw counts) |
mode="auto" (default) picks dense when the V x V matrix fits in 80% of the VRAM budget, sparse otherwise.
CLI reference
| Flag | Default | Description |
|---|---|---|
--dataset |
nlp/wikipedia |
warpdata dataset ID |
--workdir |
wiki_cooc_work |
Output directory |
--row-limit |
0 (all) |
Max dataset rows to process |
--min-freq |
10 |
Minimum token frequency for vocabulary |
--max-vocab |
300000 |
Maximum vocabulary size |
--window |
5 |
Context window (tokens left and right) |
--max-vram-gb |
20.0 |
GPU VRAM budget in GB |
--mode |
auto |
auto, dense, or sparse |
--flush-triplets-every |
20000000 |
Sparse mode: flush to CPU every N accumulated entries |
--tokenizer |
regex |
regex or hf |
--hf-model |
bert-base-uncased |
HuggingFace model name (when --tokenizer hf) |
Pipeline stages
1. Build vocabulary — tokenize dataset, count frequencies, filter by min_freq/max_vocab
2. Write ID memmap — second pass: map tokens to int32 indices, write to disk
3. GPU co-occurrence — chunk corpus, launch CUDA kernels, accumulate counts
4. Collapse (dense) — log1p → column-center → L2 normalize → embeddings.npy
Save (sparse) — raw counts → cooc_final.npz
Output files
All outputs are written to --workdir:
| File | Description |
|---|---|
word2idx.json |
{word: index} vocabulary mapping |
corpus_ids.i32 |
Memory-mapped int32 token ID array |
embeddings.npy |
Dense mode: collapsed embeddings (V, V) float32 |
cooc_final.npz |
Sparse mode: raw co-occurrence counts (scipy CSR) |
API reference
run_warp_pipeline
run_warp_pipeline(
dataset_id: str,
workdir: str,
row_limit: Optional[int] = None,
vocab_cfg: VocabConfig = VocabConfig(),
count_cfg: CountConfig = CountConfig(),
tokenizer: Optional[TokenizerBackend] = None,
) -> Path
End-to-end pipeline. Returns path to the output file.
count_cooc_gpu
count_cooc_gpu(
corpus_ids: np.ndarray, # int32 token IDs (-1 = OOV)
n_vocab: int,
window: int = 5,
max_vram_gb: float = 20.0,
mode: Literal["auto", "dense", "sparse"] = "auto",
flush_triplets_every: int = 20_000_000,
progress_every_tokens: int = 0,
) -> Union[cp.ndarray, sp.csr_matrix]
GPU co-occurrence counting. Returns CuPy dense array or scipy sparse CSR matrix.
collapse
collapse(matrix: cp.ndarray) -> np.ndarray
Transforms a dense co-occurrence matrix into normalized embeddings on GPU: log1p → subtract column means → L2 normalize rows. Returns a CPU numpy array.
Config
VocabConfig(min_freq: int = 10, max_vocab: int = 300_000)
CountConfig(window: int = 5, max_vram_gb: float = 20.0, mode: str = "auto", flush_triplets_every: int = 20_000_000)
Tokenizers
All tokenizers implement the TokenizerBackend protocol (tokenize(text: str) -> Iterable[str]).
| Class | Args |
|---|---|
RegexTokenizer |
pattern=r"[A-Za-z0-9_']+", lowercase=True |
HFTokenizer |
model_name, lowercase=False, use_fast=True |
CallableTokenizer |
fn: Callable[[str], Iterable[str]] |
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file fast_cooc-0.1.0.tar.gz.
File metadata
- Download URL: fast_cooc-0.1.0.tar.gz
- Upload date:
- Size: 11.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0307457ee2e5c05e7012d92376212b0a99c3ce3966200df984a7d172284fa651
|
|
| MD5 |
cf08c79bc6f680d5603b48bebc4772de
|
|
| BLAKE2b-256 |
56e6db6c0810ed032ea5440f6f45fbf4057dbb3d23355c841efa7e9dc76c694f
|
File details
Details for the file fast_cooc-0.1.0-py3-none-any.whl.
File metadata
- Download URL: fast_cooc-0.1.0-py3-none-any.whl
- Upload date:
- Size: 10.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
aac678bafc61cc8c411a5e8c3f9024b05095aee4f185048e29c4befb5d1c8ff6
|
|
| MD5 |
574b7787a8ccda3e94a255dcf09e80b2
|
|
| BLAKE2b-256 |
23cf08973ac56945d1557a6c637ed35fbd85e17c4384e09b9cebbfae1f371ac0
|