Skip to main content

VectorMesh

A PyTorch framework for embed once, reuse many times.

A large pretrained encoder — a BERT-family text model, a vision CNN/ViT, or a hand-written regex feature extractor — is run over a dataset exactly once, by whoever has the hardware and the judgement to pick it. The resulting vectors go to disk as a VectorCache: a versioned, documented artefact. Everything after that trains a small head on those frozen vectors, cheaply enough to run on a laptop CPU.

That split is the point. One embedding job, many models trained against its output — and the interesting design space (fusing representations, gating, mixtures of experts) lives entirely on the cheap side of it.

📚 Documentation

Full documentation is in docs/.

Core concepts The embed-once economics, thinking at the vector level, the 1D/2D/3D tensor-flow ladder
Tensor contracts Shapes as a type system: jaxtyping + beartype, and how to read a shape error
The data layer Vectorizers, VectorCache, metadata, DatasetSchema, collation
Components Every building block, with shapes and when to reach for it
Architectures Composition patterns from a two-line baseline to chunk-level MoE fusion
Training Metrics, loss choice, mltrainer wiring, the batch scripts
API reference Signature tables for everything exported
Teaching path Notebook-by-notebook map and the questions each one raises

Installation

uv sync

Requires Python ≥ 3.12. See pyproject.toml for the full dependency list.

Quick start

from pathlib import Path

import torch
from torch.utils.data import DataLoader

from vectormesh import VectorCache
from vectormesh.components import FixedPadding, MaskedMeanAggregator, NeuralNet, Serial
from vectormesh.data import Collate, OneHot

cache = VectorCache.load(path=Path("artefacts/my_dataset_train"))
hidden_size = cache.metadata["legal_dutch"]["hidden_size"]

data = cache.dataset.map(OneHot(num_classes=32, label_col="labels", target_col="onehot"))
loader = DataLoader(
    data,
    batch_size=32,
    shuffle=True,
    collate_fn=Collate(
        embedding_col="legal_dutch",
        target_col="onehot",
        padder=FixedPadding(max_chunks=30),
    ),
)

pipeline = Serial([
    MaskedMeanAggregator(),                  # (batch, chunks, dim) -> (batch, dim)
    NeuralNet(hidden_size, out_size=32),     # (batch, dim)         -> (batch, 32)
])

Then hand pipeline and loader to mltrainer.Trainer — see Training for the full wiring.

Building a cache instead of loading one, extending a cache with extra feature columns, and the image path are all covered in The data layer.

What's in the box

DataVectorCache, Vectorizer (chunked text), ImageVectorizer, RegexVectorizer, ChunkedRegexVectorizer, DatasetSchema, OneHot, Collate, CollateParallel, LabelEncoder.

Components — pipelines (Serial, Parallel), padding (FixedPadding, DynamicPadding), aggregation (Mean, MaskedMean, Attention, RNN), neural blocks (NeuralNet, Projection, Attention, TransformerBlock), connectors (Concatenate2D, Concatenate3D, Stack2D), gating (Skip, Gate, Highway, MoE), augmentation (GaussianNoise), metrics (Accuracy, F1Score, MAE, MASE).

Details and signatures: Components, API reference.

Runtime type checking

VectorMesh annotates every forward with jaxtyping shapes, checked at runtime by beartype:

@jaxtyped(typechecker=beartype)
def forward(self, tensors: Float[Tensor, "batch chunks dim"]) -> Float[Tensor, "batch dim"]:
    return tensors.mean(dim=1)

This catches the failure mode that matters most here: nn.Linear accepts any leading shape, so feeding it (batch, chunks, dim) when you meant (batch, dim) raises no error — it just trains a model that quietly answers the wrong question. A BeartypeCallHintParamViolation naming a 3D tensor where a 2D one was expected almost always means a missing aggregator.

How to read those errors: Tensor contracts.

Notebooks

Notebook Topic
0_vectorizer.ipynb Creating vector caches; extending one with regex features
1_training.ipynb Cache → padding → aggregation → pipeline → trained model
2_design.ipynb Parallel branches, fusing two representations, skip connections
3_moe.ipynb Mixture of experts
4_image_vectorizer.ipynb The same pipeline on images; feature-space augmentation

Walkthrough and exercises: Teaching path.

Scripts

Batch counterparts to the notebooks, for real dataset sizes:

uv run python scripts/create_cache_aktes.py    # embeddings
uv run python scripts/add_chunked_regex.py     # + chunk-aligned regex column
uv run python scripts/train_moe_parallel.py    # train

Full list: Training §6.5.

Development

uv run pytest -m "not integration"   # unit tests, no model downloads
uv run pytest                        # everything, downloads HF models
uv run ruff check src tests
uv run ty check

Project structure

src/vectormesh/
├── types.py               # VectorMeshError, Cachable, BaseComponent, TensorInput
├── data/                  # vectorizers, cache, schema, collation
└── components/            # pipelines, padding, aggregation, neural,
                           # connectors, gating, augmentation, metrics
docs/                      # documentation (start at docs/README.md)
notebooks/                 # tutorials, in order
scripts/                   # batch pipelines
references/                # source papers
tests/

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vectormesh-0.4.0.tar.gz (34.9 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

vectormesh-0.4.0-py3-none-any.whl (31.2 kB view details)

Uploaded Python 3

File details

Details for the file vectormesh-0.4.0.tar.gz.

File metadata

  • Download URL: vectormesh-0.4.0.tar.gz
  • Upload date:
  • Size: 34.9 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for vectormesh-0.4.0.tar.gz
Algorithm Hash digest
SHA256 e3d0a8e05b81255e8835cafc5b22ae61462f81ef639035837c13c54260a5d93e
MD5 768f7561af951cfac668fb3d3e42c18b
BLAKE2b-256 c960dc9acccfe2a7591bc82e853e776b0cf0356d560059d9a701439fd3fe0a89

See more details on using hashes here.

File details

Details for the file vectormesh-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: vectormesh-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 31.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for vectormesh-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 9cb9940183d986169315dff32db46952529ffda8a309571d09f0c84d25ab598d
MD5 65402a686ba852e93f51a3aa4cb192d9
BLAKE2b-256 521ea1f50b5bcae0196e0b8def2d8760e0f5951c5e40b00f355ea985fe9fd2ee

See more details on using hashes here.

Release history Release notifications | RSS feed

0.10.0

2 files

0.9.2

2 files

0.9.1

2 files

0.9.0

2 files

0.8.2

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

0.4.4

2 files

0.4.3

2 files

0.4.2

2 files

0.4.1

2 files

This release

0.4.0 This release

2 files

0.2.0

2 files

0.1.0

1 file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page