VectorMesh
A PyTorch framework for embed once, reuse many times.
A large pretrained encoder — a BERT-family text model, a vision CNN/ViT, or a hand-written regex
feature extractor — is run over a dataset exactly once, by whoever has the hardware and the
judgement to pick it. The resulting vectors go to disk as a VectorCache: a versioned,
documented artefact. Everything after that trains a small head on those frozen vectors, cheaply
enough to run on a laptop CPU.
That split is the point. One embedding job, many models trained against its output — and the interesting design space (fusing representations, gating, mixtures of experts) lives entirely on the cheap side of it.
📚 Documentation
Full documentation is in docs/.
| Core concepts | The embed-once economics, thinking at the vector level, the 1D/2D/3D tensor-flow ladder |
| Tensor contracts | Shapes as a type system: jaxtyping + beartype, and how to read a shape error |
| The data layer | Vectorizers, VectorCache, metadata, DatasetSchema, collation |
| Components | Every building block, with shapes and when to reach for it |
| Architectures | Composition patterns from a two-line baseline to chunk-level MoE fusion |
| Training | Metrics, loss choice, mltrainer wiring, the batch scripts |
| API reference | Signature tables for everything exported |
| Teaching path | Notebook-by-notebook map and the questions each one raises |
Installation
uv sync
Requires Python ≥ 3.12. See pyproject.toml for the full dependency list.
Quick start
from pathlib import Path
import torch
from torch.utils.data import DataLoader
from vectormesh import VectorCache
from vectormesh.components import FixedPadding, MaskedMeanAggregator, NeuralNet, Serial
from vectormesh.data import Collate, OneHot
cache = VectorCache.load(path=Path("artefacts/my_dataset_train"))
hidden_size = cache.metadata["legal_dutch"]["hidden_size"]
data = cache.dataset.map(OneHot(num_classes=32, label_col="labels", target_col="onehot"))
loader = DataLoader(
data,
batch_size=32,
shuffle=True,
collate_fn=Collate(
embedding_col="legal_dutch",
target_col="onehot",
padder=FixedPadding(max_chunks=30),
),
)
pipeline = Serial([
MaskedMeanAggregator(), # (batch, chunks, dim) -> (batch, dim)
NeuralNet(hidden_size, out_size=32), # (batch, dim) -> (batch, 32)
])
Then hand pipeline and loader to mltrainer.Trainer — see
Training for the full wiring.
Building a cache instead of loading one, extending a cache with extra feature columns, and the image path are all covered in The data layer.
What's in the box
Data — VectorCache, Vectorizer (chunked text), ImageVectorizer, RegexVectorizer,
ChunkedRegexVectorizer, DatasetSchema, OneHot, Collate, CollateParallel, LabelEncoder.
Components — pipelines (Serial, Parallel), padding (FixedPadding, DynamicPadding),
aggregation (Mean, MaskedMean, Attention, RNN), neural blocks (NeuralNet, Projection,
Attention, TransformerBlock), connectors (Concatenate2D, Concatenate3D, Stack2D), gating
(Skip, Gate, Highway, MoE), augmentation (GaussianNoise), metrics (Accuracy, F1Score,
MAE, MASE).
Details and signatures: Components, API reference.
Runtime type checking
VectorMesh annotates every forward with jaxtyping shapes, checked at runtime by beartype:
@jaxtyped(typechecker=beartype)
def forward(self, tensors: Float[Tensor, "batch chunks dim"]) -> Float[Tensor, "batch dim"]:
return tensors.mean(dim=1)
This catches the failure mode that matters most here: nn.Linear accepts any leading shape, so
feeding it (batch, chunks, dim) when you meant (batch, dim) raises no error — it just trains
a model that quietly answers the wrong question. A BeartypeCallHintParamViolation naming a 3D
tensor where a 2D one was expected almost always means a missing aggregator.
How to read those errors: Tensor contracts.
Notebooks
| Notebook | Topic |
|---|---|
0_vectorizer.ipynb |
Creating vector caches; extending one with regex features |
1_training.ipynb |
Cache → padding → aggregation → pipeline → trained model |
2_design.ipynb |
Parallel branches, fusing two representations, skip connections |
3_moe.ipynb |
Mixture of experts |
4_image_vectorizer.ipynb |
The same pipeline on images; feature-space augmentation |
Walkthrough and exercises: Teaching path.
Scripts
Batch counterparts to the notebooks, for real dataset sizes:
uv run python scripts/create_cache_aktes.py # embeddings
uv run python scripts/add_chunked_regex.py # + chunk-aligned regex column
uv run python scripts/train_moe_parallel.py # train
Full list: Training §6.5.
Development
uv run pytest -m "not integration" # unit tests, no model downloads
uv run pytest # everything, downloads HF models
uv run ruff check src tests
uv run ty check
Project structure
src/vectormesh/
├── types.py # VectorMeshError, Cachable, BaseComponent, TensorInput
├── data/ # vectorizers, cache, schema, collation
└── components/ # pipelines, padding, aggregation, neural,
# connectors, gating, augmentation, metrics
docs/ # documentation (start at docs/README.md)
notebooks/ # tutorials, in order
scripts/ # batch pipelines
references/ # source papers
tests/
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file vectormesh-0.4.3.tar.gz.
File metadata
- Download URL: vectormesh-0.4.3.tar.gz
- Upload date:
- Size: 52.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cb861f8ae8c02b6a1c2de90dc23fba86fa0c2dbd02303fa6709ec6aff0b74dc0
|
|
| MD5 |
1570ca3215c0be0baccebbe4197f9ac4
|
|
| BLAKE2b-256 |
730b1fb5b390d088d27dd3fcb279d0e8f8b8b21df93abdd956adda5395b3fc0b
|
File details
Details for the file vectormesh-0.4.3-py3-none-any.whl.
File metadata
- Download URL: vectormesh-0.4.3-py3-none-any.whl
- Upload date:
- Size: 36.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e5a7097fec0a6de93c99bdab03936f5d829d6999050be3d1e90a9063d8a5ace5
|
|
| MD5 |
639fcb792f3144e9c65f23f74fd1cfc2
|
|
| BLAKE2b-256 |
40dd3201a6537f34b868d64d68381f337f9a65c26a365f82c4b8dd0f031769af
|