Word2Vec with Matryoshka Representation Learning (Rust+PyO3), gensim-style API, numpy interop.
Project description
Word2Vec + Matryoshka
Rust + PyO3 backed Python library implementing Word2Vec (Skip‑gram/CBOW) with Negative Sampling (NS) and Hierarchical Softmax (HS), extended with Matryoshka multi‑level representations (prefix vectors). The public API mirrors gensim where sensible (Word2Vec, KeyedVectors, wv[...], wv.most_similar), and all vectors are returned as numpy.ndarray.
Quickstart (uv + maturin)
uv sync
uv add --dev ruff pytest maturin
uvx maturin develop
uv run -m pytest -q
To enable optional performance features at build time:
# AVX/SSE accelerated dot/AXPY (runtime‑detected)
uvx maturin develop --features simd
# Fast sigmoid approximation (minor accuracy trade‑off)
uvx maturin develop --features fast_sigmoid
# Combine multiple features
uvx maturin develop --features "simd,fast_sigmoid"
Usage
from word2vec_matryoshka import Word2Vec, KeyedVectors, set_seed
set_seed(42)
texts = [["hello", "world"], ["computer", "science"]]
model = Word2Vec(
sentences=texts,
vector_size=100, window=5, min_count=1, workers=4,
negative=5, sg=1, hs=0, levels=[25, 50, 100],
alpha=0.025, min_alpha=0.0001,
verbose=True, progress_interval=1.0,
)
"""
Word2Vec.save/load writes and reads a base path, producing:
<base>.vocab.json # vocabulary in ivocab order
<base>.npy # float32 matrix (rows=len(vocab), cols=vector_size)
The base can be any string (e.g., "word2vec.model"); the two files will be created next to it.
"""
model.save("word2vec.model")
model = Word2Vec.load("word2vec.model")
model.train([["hello", "world"]], total_examples=1, epochs=1)
vec = model.wv["computer"] # numpy.ndarray (full dimension)
sims = model.wv.most_similar("computer", topn=10, level=50) # prefix level
model.wv.save("word2vec.wordvectors")
wv = KeyedVectors.load("word2vec.wordvectors", mmap="r") # zero‑copy memmap
vec50 = wv.get_vector("computer", level=50)
# Inspect vocabulary and matrix
print(wv.index_to_key[:5]) # ['computer', 'hello', ...] in index order
M = wv.vectors # 2D numpy array (n_keys, vector_size)
print(M.shape)
Constructor notes:
- sentences (iterable of iterables, optional): The sentences iterable can be a simple list of token lists for small corpora. For large corpora, prefer an iterable that streams sentences directly from disk or network to avoid loading everything into memory.
- If you don’t supply sentences, the model is left uninitialized — use this if you plan to ingest/train later or initialize weights in another way.
- alpha (float, optional): Initial learning rate (default 0.025).
- min_alpha (float, optional): Learning rate linearly decays to this value over training (default 0.0001).
- verbose (bool, optional): Enable gensim-style training progress logging via Python
logging(default True; can be overridden pertrain). - progress_interval (float, optional): Seconds between progress updates when logging or callback is enabled (default 1.0; can be overridden per
train).
Deferred ingestion example:
from word2vec_matryoshka import Word2Vec
# Initialize without sentences (uninitialized model)
m = Word2Vec(vector_size=64, window=5, min_count=1, workers=2, levels=[16, 32, 64])
# Provide a restartable streaming iterable later
class Corpus:
def __iter__(self):
# stream from disk/network in real use
yield ["hello", "world"]
yield ["computer", "science"]
corpus = Corpus()
m.train(corpus, epochs=2)
# Now you can query vectors
vec = m.wv["hello"]
Important: for epochs > 1 or multiple train calls, the iterable must be restartable — implement __iter__ to return a fresh iterator each time. Avoid one‑shot generators that exhaust after the first pass.
Training Progress (gensim-style logging)
This library integrates with Python logging to emit gensim-like progress lines when verbose=True.
import logging
from word2vec_matryoshka import Word2Vec
# Configure logging to match gensim's default look
logging.basicConfig(
format="%(asctime)s: %(levelname)s: %(message)s",
datefmt="%Y-%m-%d %H:%M:%S",
level=logging.INFO,
)
corpus = [["hello", "world"], ["computer", "science"]] * 2000
model = Word2Vec(sentences=corpus, vector_size=64, workers=4)
# Emits lines like:
# 2025-09-12 21:06:40: INFO: PROGRESS: at 99.24% tokens, alpha 0.00019, 95350 tokens/s
model.train(corpus, epochs=2, verbose=True, progress_interval=0.5)
# You can also set defaults at construction and omit them here:
m2 = Word2Vec(sentences=corpus, vector_size=64, verbose=True, progress_interval=0.5)
m2.train(corpus, epochs=1) # uses model-level defaults
- Logger name:
word2vec_matryoshka(set level/handlers on this logger to control output). - Messages are rate-limited by
progress_intervalseconds and include an end-of-epoch flush. - Default behavior is silent (no logging) unless
verbose=Trueor aprogresscallback is supplied.
You can still use a Python callback if you prefer manual handling of progress (logger can be off or on):
def on_progress(done: int, total: int) -> None:
pass # e.g., update a progress bar
model.train(corpus, epochs=2, progress=on_progress, progress_interval=0.5)
# or combine with verbose logging
model.train(corpus, epochs=1, progress=on_progress, verbose=True)
Files created by the above code:
- Model save:
word2vec.model.vocab.json,word2vec.model.npy - KeyedVectors save:
word2vec.wordvectors.vocab.json,word2vec.wordvectors.npy
Streaming Training
sentencescan be any restartable iterable of token sequences: each epoch must be able to iterate again (define__iter__to return a fresh iterator; avoid one‑shot generators).- With
workers=1, training is truly streaming and does not materialize the full corpus. - With
workers>1, a bounded prefetch queue batches sentences per epoch for parallel compute (to avoid cross‑thread access to a Python iterator).
Example: restartable iterable
from word2vec_matryoshka import Word2Vec
class RestartableCorpus:
def __init__(self, data):
self._data = list(data)
def __iter__(self):
for sent in self._data:
yield list(sent)
corpus = RestartableCorpus([["hello", "world"], ["computer", "science"], ["hello", "computer"]])
model = Word2Vec(sentences=corpus, vector_size=32, window=5, min_count=1, workers=1, levels=[16, 32])
model.train(corpus, total_examples=3, epochs=2)
Run the built‑in example: uv run python -m word2vec_matryoshka._streaming_examples
Features
- Training: Skip‑gram (
sg=1) or CBOW (sg=0); choose NS (negative>0andhs=0) or HS (hs=1). - Matryoshka levels:
levels=[...]defines multiple prefix dimensions (defaults to[d/4, d/2, d]if omitted). Each prefix is optimized, enabling multi‑granularity queries. - Vector persistence:
KeyedVectors.save("base")producesbase.vocab.jsonandbase.npy(float32, C‑order).KeyedVectors.load("base", mmap='r')uses NumPy memmap for zero‑copy.most_similaron mmap uses NumPy vectorized dot and norms for performance. - Model persistence:
Word2Vec.save/load("base")writes and reads the samebase.vocab.json+base.npyformat asKeyedVectors, preservingivocaborder and using atomic writes. - Determinism:
set_seed(seed)for reproducible training; per‑thread RNGs derive from the base seed. - Parallelism: uses
rayonwith lock striping for safe, concurrent updates; optional SIMD acceleration available via--features simd.
Performance Options
--features simd: enables x86_64 AVX/SSE and aarch64 NEON intrinsics for core dot and AXPY (update) operations with runtime detection. Falls back to scalar on unsupported CPUs.--features fast_sigmoid: uses a smooth approximation0.5 + 0.5 * x / (1 + |x|)instead ofexp‑based sigmoid; improves throughput in activation‑heavy paths with small accuracy impact.- Thread‑local buffers: CBOW paths reuse per‑thread buffers to reduce allocations (enabled by default; no action required).
- Workers: tune
workersto available cores;workers>1uses parallel compute with a bounded prefetch. - Levels: fewer/lower levels reduce compute;
[d/4, d/2, d]is a balanced default.
Layout (excerpt)
src/ # Rust core (PyO3 module: word2vec_matryoshka._core)
python/word2vec_matryoshka # Python package (re‑exports _core)
tests/ # Pytests: API, IO, mmap, streaming, HS/NS, matryoshka
### Internal Rust module structure
src/ lib.rs # PyO3 entry, high‑level types, logging/RNG glue io/npy.rs # NPY read/write helpers ops/ # math kernels, SIMD paths, thread‑local buffers weights.rs # striped‑lock shared weights (SharedWeights, SHARDS) training/ ns.rs # SG/CBOW + Negative Sampling (seq + striped) hs.rs # Hierarchical Softmax + Huffman sampling/ alias.rs # alias method build/sample
This split keeps unsafe/SIMD and concurrency concerns localized and makes training variants easy to extend.
### Dev quality gates
- Rust: `cargo fmt` and `cargo clippy -- -D warnings` must pass.
- Python: `uv run ruff check .` and `uv run ruff format .` for lint/format.
Tips
- For production, prefer
mmap='r'when loading read‑only vectors to reduce memory. If you need a writable/contiguous copy, call.copy()on the NumPy array in Python.
License
This project is licensed under the MIT License. See the LICENSE file for details.
Micro‑benchmark and HS/NS switch
import time
from word2vec_matryoshka import Word2Vec
corpus = [[f"w{i}" for i in range(200)] for _ in range(200)]
def bench(**kw):
t0 = time.perf_counter()
m = Word2Vec(
sentences=corpus, vector_size=128, window=5, min_count=1,
workers=4, levels=[32, 64, 128], **kw
)
m.train(corpus, total_examples=len(corpus), epochs=1)
return time.perf_counter() - t0
print("NS (negative=5):\t", bench(negative=5, sg=1, hs=0))
print("HS: \t", bench(negative=0, sg=1, hs=1))
Rules of thumb: with large vocabularies/high dimensions NS is often faster with good accuracy; HS can be more stable for smaller vocabs or long‑tail focus. Matryoshka levels work with both.
Troubleshooting
- After
uvx maturin develop, run tests viauv run -m pytest -qso pytest sees the dev‑installed extension. - If you interact with Python objects from background threads, bind them under the GIL in that thread. Avoid holding the GIL while blocking on channels; wrap long or blocking sections in
Python::with_gil(|py| py.allow_threads(|| { ... })). - When cloning PyO3 handles, use
clone_ref(py)under the GIL rather thanclone()onPy<T>. - To run the test suite with features enabled, build the extension with the same features before invoking pytest, e.g.
uvx maturin develop --features simd && uv run -m pytest -q.
Migration note (save/load)
- If you previously saved a full model as a single JSON file, switch to the base‑path pattern shown above. Using the same base (even if it ends with
.json) is fine; the library will create<base>.vocab.jsonand<base>.npyalongside it. - For very large models, prefer loading vectors via
KeyedVectors.load(base, mmap='r')for minimal memory usage and fast startup.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file word2vec_matryoshka-0.1.0.tar.gz.
File metadata
- Download URL: word2vec_matryoshka-0.1.0.tar.gz
- Upload date:
- Size: 72.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
744ac6bdf91baf1578ea593ef74ae938522abcf826f8509f6fe679810d51d51b
|
|
| MD5 |
479a40f11c45317d14d3eab9727e8ed6
|
|
| BLAKE2b-256 |
7534231ced930040200db4422ff0d9b73601a79ebc6f3314bd5708541a8c3764
|
Provenance
The following attestation bundles were made for word2vec_matryoshka-0.1.0.tar.gz:
Publisher:
publish.yml on feisan/word2vec_matryoshka
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
word2vec_matryoshka-0.1.0.tar.gz -
Subject digest:
744ac6bdf91baf1578ea593ef74ae938522abcf826f8509f6fe679810d51d51b - Sigstore transparency entry: 517671750
- Sigstore integration time:
-
Permalink:
feisan/word2vec_matryoshka@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/feisan
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Trigger Event:
release
-
Statement type:
File details
Details for the file word2vec_matryoshka-0.1.0-cp310-abi3-win_amd64.whl.
File metadata
- Download URL: word2vec_matryoshka-0.1.0-cp310-abi3-win_amd64.whl
- Upload date:
- Size: 381.5 kB
- Tags: CPython 3.10+, Windows x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f04c7c1a83003aaba226daf7237c2574d06328bab4dd8cddf20155c234a02d22
|
|
| MD5 |
49f4d2bec07ecbf04958d65326440ea0
|
|
| BLAKE2b-256 |
0b330acd1adbd4a980cc8cd74f323a3d1feda87b541c6ddec0536c473a8a1037
|
Provenance
The following attestation bundles were made for word2vec_matryoshka-0.1.0-cp310-abi3-win_amd64.whl:
Publisher:
publish.yml on feisan/word2vec_matryoshka
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
word2vec_matryoshka-0.1.0-cp310-abi3-win_amd64.whl -
Subject digest:
f04c7c1a83003aaba226daf7237c2574d06328bab4dd8cddf20155c234a02d22 - Sigstore transparency entry: 517671869
- Sigstore integration time:
-
Permalink:
feisan/word2vec_matryoshka@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/feisan
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Trigger Event:
release
-
Statement type:
File details
Details for the file word2vec_matryoshka-0.1.0-cp310-abi3-musllinux_1_2_x86_64.whl.
File metadata
- Download URL: word2vec_matryoshka-0.1.0-cp310-abi3-musllinux_1_2_x86_64.whl
- Upload date:
- Size: 566.1 kB
- Tags: CPython 3.10+, musllinux: musl 1.2+ x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ec299c9a0c26c6e31b992c249dacb657409997f1f7a924f77cb2a9dbc9133fb8
|
|
| MD5 |
28c80f04753a5dd845d76162fd2de7ec
|
|
| BLAKE2b-256 |
4f1032c73a9126c51a9815b2d7c80e06f85ecd72bc662fefc496a306e32042cd
|
Provenance
The following attestation bundles were made for word2vec_matryoshka-0.1.0-cp310-abi3-musllinux_1_2_x86_64.whl:
Publisher:
publish.yml on feisan/word2vec_matryoshka
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
word2vec_matryoshka-0.1.0-cp310-abi3-musllinux_1_2_x86_64.whl -
Subject digest:
ec299c9a0c26c6e31b992c249dacb657409997f1f7a924f77cb2a9dbc9133fb8 - Sigstore transparency entry: 517671901
- Sigstore integration time:
-
Permalink:
feisan/word2vec_matryoshka@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/feisan
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Trigger Event:
release
-
Statement type:
File details
Details for the file word2vec_matryoshka-0.1.0-cp310-abi3-manylinux_2_28_x86_64.whl.
File metadata
- Download URL: word2vec_matryoshka-0.1.0-cp310-abi3-manylinux_2_28_x86_64.whl
- Upload date:
- Size: 488.1 kB
- Tags: CPython 3.10+, manylinux: glibc 2.28+ x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0a747dac70c4776fb7be39ef2a0524101bac94a6f309650ee9a2acb6f603c1e6
|
|
| MD5 |
7effeabdd25368ad457e2fe92d2e955d
|
|
| BLAKE2b-256 |
8f4706a795dd0289dd2aaf413a39ba5837e17fd3743e4a4acfce4830b6225b33
|
Provenance
The following attestation bundles were made for word2vec_matryoshka-0.1.0-cp310-abi3-manylinux_2_28_x86_64.whl:
Publisher:
publish.yml on feisan/word2vec_matryoshka
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
word2vec_matryoshka-0.1.0-cp310-abi3-manylinux_2_28_x86_64.whl -
Subject digest:
0a747dac70c4776fb7be39ef2a0524101bac94a6f309650ee9a2acb6f603c1e6 - Sigstore transparency entry: 517671795
- Sigstore integration time:
-
Permalink:
feisan/word2vec_matryoshka@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/feisan
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Trigger Event:
release
-
Statement type:
File details
Details for the file word2vec_matryoshka-0.1.0-cp310-abi3-macosx_11_0_universal2.whl.
File metadata
- Download URL: word2vec_matryoshka-0.1.0-cp310-abi3-macosx_11_0_universal2.whl
- Upload date:
- Size: 861.1 kB
- Tags: CPython 3.10+, macOS 11.0+ universal2 (ARM64, x86-64)
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9d07ce376e1b5c0d114f60fb20ed780db09bfe63e4ee63a13d2ee513a456ddc3
|
|
| MD5 |
2ced87a0a86fd67ea84d0aa45ae1558a
|
|
| BLAKE2b-256 |
fe3192fbf29b49c4fec7a5e692e9598b97a13aaad796b3fd36766b22ccd18cd1
|
Provenance
The following attestation bundles were made for word2vec_matryoshka-0.1.0-cp310-abi3-macosx_11_0_universal2.whl:
Publisher:
publish.yml on feisan/word2vec_matryoshka
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
word2vec_matryoshka-0.1.0-cp310-abi3-macosx_11_0_universal2.whl -
Subject digest:
9d07ce376e1b5c0d114f60fb20ed780db09bfe63e4ee63a13d2ee513a456ddc3 - Sigstore transparency entry: 517671845
- Sigstore integration time:
-
Permalink:
feisan/word2vec_matryoshka@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/feisan
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Trigger Event:
release
-
Statement type:
File details
Details for the file word2vec_matryoshka-0.1.0-cp310-abi3-macosx_11_0_arm64.whl.
File metadata
- Download URL: word2vec_matryoshka-0.1.0-cp310-abi3-macosx_11_0_arm64.whl
- Upload date:
- Size: 425.7 kB
- Tags: CPython 3.10+, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4a8fd2d998b9ab903539c9a8a6344dd60eddaef1805d4e61dbdbda319262f0c4
|
|
| MD5 |
d0d31102bc935bfcd251905ac314cb95
|
|
| BLAKE2b-256 |
0a4a3955f33017f83ca315723a4d9bfebed7ed74552488dd6d930d4b237f4057
|
Provenance
The following attestation bundles were made for word2vec_matryoshka-0.1.0-cp310-abi3-macosx_11_0_arm64.whl:
Publisher:
publish.yml on feisan/word2vec_matryoshka
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
word2vec_matryoshka-0.1.0-cp310-abi3-macosx_11_0_arm64.whl -
Subject digest:
4a8fd2d998b9ab903539c9a8a6344dd60eddaef1805d4e61dbdbda319262f0c4 - Sigstore transparency entry: 517671825
- Sigstore integration time:
-
Permalink:
feisan/word2vec_matryoshka@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/feisan
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a4dc076ce44292686afa5d9b9c248518072e9a80 -
Trigger Event:
release
-
Statement type: