Skip to main content

VectorMesh

A PyTorch framework for embed once, reuse many times.

A large pretrained encoder — a BERT-family text model, a vision CNN/ViT, or a hand-written regex feature extractor — is run over a dataset exactly once, by whoever has the hardware and the judgement to pick it. The resulting vectors go to disk as a VectorCache: a versioned, documented artefact. Everything after that trains a small head on those frozen vectors, cheaply enough to run on a laptop CPU.

That split is the point. One embedding job, many models trained against its output — and the interesting design space (fusing representations, gating, mixtures of experts) lives entirely on the cheap side of it.

📚 Documentation

Full documentation is in docs/.

Core concepts The embed-once economics, thinking at the vector level, the 1D/2D/3D tensor-flow ladder
Tensor contracts Shapes as a type system: jaxtyping + beartype, and how to read a shape error
The data layer Vectorizers, VectorCache, metadata, DatasetSchema, collation
Components Every building block, with shapes and when to reach for it
Architectures Composition patterns from a two-line baseline to chunk-level MoE fusion
Training Metrics, loss choice, mltrainer wiring, the batch scripts
API reference Signature tables for everything exported
Teaching path Notebook-by-notebook map and the questions each one raises

Installation

uv sync

Requires Python ≥ 3.12. See pyproject.toml for the full dependency list.

Quick start

from pathlib import Path

import torch
from torch.utils.data import DataLoader

from vectormesh import VectorCache
from vectormesh.components import FixedPadding, MaskedMeanAggregator, NeuralNet, Serial
from vectormesh.data import Collate, OneHot

cache = VectorCache.load(path=Path("artefacts/my_dataset_train"))
hidden_size = cache.metadata["legal_dutch"]["hidden_size"]

data = cache.dataset.map(OneHot(num_classes=32, label_col="labels", target_col="onehot"))
loader = DataLoader(
    data,
    batch_size=32,
    shuffle=True,
    collate_fn=Collate(
        embedding_col="legal_dutch",
        target_col="onehot",
        padder=FixedPadding(max_chunks=30),
    ),
)

pipeline = Serial([
    MaskedMeanAggregator(),                  # (batch, chunks, dim) -> (batch, dim)
    NeuralNet(hidden_size, out_size=32),     # (batch, dim)         -> (batch, 32)
])

Then hand pipeline and loader to mltrainer.Trainer — see Training for the full wiring.

Building a cache instead of loading one, extending a cache with extra feature columns, and the image path are all covered in The data layer.

What's in the box

DataVectorCache, Vectorizer (chunked text), ImageVectorizer, RegexVectorizer, ChunkedRegexVectorizer, DatasetSchema, OneHot, Collate, CollateParallel, LabelEncoder.

Components — pipelines (Serial, Parallel), padding (FixedPadding, DynamicPadding), aggregation (Mean, MaskedMean, Attention, RNN), neural blocks (NeuralNet, Projection, Attention, TransformerBlock), connectors (Concatenate2D, Concatenate3D, Stack2D), gating (Skip, Gate, Highway, MoE), augmentation (GaussianNoise), metrics (Accuracy, F1Score, MAE, MASE).

Details and signatures: Components, API reference.

Runtime type checking

VectorMesh annotates every forward with jaxtyping shapes, checked at runtime by beartype:

@jaxtyped(typechecker=beartype)
def forward(self, tensors: Float[Tensor, "batch chunks dim"]) -> Float[Tensor, "batch dim"]:
    return tensors.mean(dim=1)

This catches the failure mode that matters most here: nn.Linear accepts any leading shape, so feeding it (batch, chunks, dim) when you meant (batch, dim) raises no error — it just trains a model that quietly answers the wrong question. A BeartypeCallHintParamViolation naming a 3D tensor where a 2D one was expected almost always means a missing aggregator.

How to read those errors: Tensor contracts.

Notebooks

Notebook Topic
0_vectorizer.ipynb Creating vector caches; extending one with regex features
1_training.ipynb Cache → padding → aggregation → pipeline → trained model
2_design.ipynb Parallel branches, fusing two representations, skip connections
3_moe.ipynb Mixture of experts
4_image_vectorizer.ipynb The same pipeline on images; feature-space augmentation

Walkthrough and exercises: Teaching path.

Scripts

Batch counterparts to the notebooks, for real dataset sizes:

uv run python scripts/create_cache_aktes.py    # embeddings
uv run python scripts/add_chunked_regex.py     # + chunk-aligned regex column
uv run python scripts/train_moe_parallel.py    # train

Full list: Training §6.5.

Development

uv run pytest -m "not integration"   # unit tests, no model downloads
uv run pytest                        # everything, downloads HF models
uv run ruff check src tests
uv run ty check

Project structure

src/vectormesh/
├── types.py               # VectorMeshError, Cachable, BaseComponent, TensorInput
├── data/                  # vectorizers, cache, schema, collation
└── components/            # pipelines, padding, aggregation, neural,
                           # connectors, gating, augmentation, metrics
docs/                      # documentation (start at docs/README.md)
notebooks/                 # tutorials, in order
scripts/                   # batch pipelines
references/                # source papers
tests/

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vectormesh-0.4.4.tar.gz (52.4 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

vectormesh-0.4.4-py3-none-any.whl (39.4 kB view details)

Uploaded Python 3

File details

Details for the file vectormesh-0.4.4.tar.gz.

File metadata

  • Download URL: vectormesh-0.4.4.tar.gz
  • Upload date:
  • Size: 52.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for vectormesh-0.4.4.tar.gz
Algorithm Hash digest
SHA256 ed408d11040ed235585bb0b4c1b1fc7eaf3c9d69abbdcc04cef96f67e355445d
MD5 3e7a90311ecea71545a9eadfd1f9c1b4
BLAKE2b-256 a5338d9fff8645ea3c9179a493688f23cbbaca40811219422170258d8b3f28e3

See more details on using hashes here.

File details

Details for the file vectormesh-0.4.4-py3-none-any.whl.

File metadata

  • Download URL: vectormesh-0.4.4-py3-none-any.whl
  • Upload date:
  • Size: 39.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for vectormesh-0.4.4-py3-none-any.whl
Algorithm Hash digest
SHA256 8c8aeac8e3b28abfba15c4b20f1c4e08af5369ca25995772ce83bd4e31a6d265
MD5 957994c3e6757d432c6e211a4cee3461
BLAKE2b-256 13fce62cdef78e25c8680525aff2a1f21e1539aee6740729fbb47170df425fd3

See more details on using hashes here.

Release history Release notifications | RSS feed

0.10.0

2 files

0.9.2

2 files

0.9.1

2 files

0.9.0

2 files

0.8.2

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

This release

0.4.4 This release

2 files

0.4.3

2 files

0.4.2

2 files

0.4.1

2 files

0.4.0

2 files

0.2.0

2 files

0.1.0

1 file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page