BabelVec
Position-aware, cross-lingually aligned word embeddings built on FastText.
Features
- Cross-Lingual Alignment: Procrustes alignment for multilingual compatibility
- Position-Aware Embeddings: Optional positional encoding (RoPE, sinusoidal, decay)
- FastText Foundation: Handles OOV words through subword information
Installation
pip install babelvec
For visualization support:
pip install babelvec[viz]
Quick Start
from babelvec import BabelVec
# Load a model
model = BabelVec.load('path/to/model.bin')
# Get word vector
vec = model.get_word_vector("hello")
# Position-aware sentence embedding
vec1 = model.get_sentence_vector("The dog bites the man", method='rope')
vec2 = model.get_sentence_vector("The man bites the dog", method='rope')
# vec1 != vec2 because word order is encoded
# Simple averaging (no position encoding)
vec = model.get_sentence_vector("Hello world", method='average')
Training
Monolingual Training
from babelvec.training import train_monolingual
model = train_monolingual(
lang='en',
corpus_path='corpus.txt',
dim=300,
epochs=5,
threads=8 # Optional: specify number of threads
)
model.save('en_300d.bin')
Parallel Multi-Language Training (v0.1.4+)
Train multiple languages simultaneously for faster training on multi-core servers:
from babelvec.training import train_multiple_languages, get_cpu_count
# Auto-detects CPU cores
print(f"Using {get_cpu_count()} cores")
models = train_multiple_languages(
languages={'en': 'en_corpus.txt', 'ar': 'ar_corpus.txt'},
parallel=True, # Train languages simultaneously
max_workers=2, # Number of parallel training jobs
)
Multilingual Training with Alignment
from babelvec.training import train_multilingual
models = train_multilingual(
languages=['en', 'ar'],
corpus_paths={'en': 'en.txt', 'ar': 'ar.txt'},
parallel_data={('en', 'ar'): parallel_pairs},
alignment='procrustes',
threads=8 # Optional: specify number of threads
)
Post-hoc Alignment
from babelvec.training import align_models
aligned = align_models(
models={'en': model_en, 'ar': model_ar},
parallel_data={('en', 'ar'): parallel_pairs},
method='procrustes'
)
Model Save/Load (v0.1.3+)
Models save projection matrices alongside the FastText binary:
# Save model
model.save('model.bin')
# Creates: model.bin, model.projection.npy (if aligned), model.meta.json
# Load model - projection is automatically restored
model = BabelVec.load('model.bin')
print(model.is_aligned) # True if projection was loaded
Encoding Methods
| Method | Description |
|---|---|
rope |
Rotary Position Embedding |
decay |
Exponential position decay |
sinusoidal |
Transformer-style positional encoding |
average |
Simple averaging (no position encoding) |
Evaluation
from babelvec.evaluation import cross_lingual_retrieval
metrics = cross_lingual_retrieval(
model_src=model_en,
model_tgt=model_ar,
parallel_sentences=test_pairs,
method='rope'
)
print(f"Recall@1: {metrics['recall@1']:.3f}")
Language Families for Joint Training
BabelVec includes a curated family assignment system for 355 Wikipedia languages, optimized for joint multilingual training.
from babelvec.families import get_family_key, get_family_languages, get_training_groups
# Get family for a language
get_family_key("ary") # -> "arabic"
get_family_key("fr") # -> "romance_galloitalic"
# Get all languages in a family
get_family_languages("arabic") # -> ["ar", "ary", "arz"]
# Create training groups (hybrid strategy)
groups = get_training_groups(
languages=["en", "ar", "ary", "arz"],
article_counts={"en": 6000000, "ar": 840000, "ary": 17000, "arz": 40000},
low_resource_threshold=50000
)
# -> {"separate": ["en", "ar"], "joint": {"arabic": ["ary", "arz"]}}
Joint training dramatically improves low-resource languages (+200-600% for Arabic dialects) while high-resource languages should be trained separately.
Examples
See the examples/ directory:
01_basic_usage.py- Getting started
Citation
@misc{babelvec2025,
title = {BabelVec: Position-Aware Cross-Lingual Word Embeddings},
author = {Kamali, Omar},
doi = {10.5281/zenodo.18065206},
publisher = {Zenodo},
year = {2025},
url = {https://github.com/omarkamali/babelvec}
}
License
MIT License - see LICENSE for details.
Copyright © 2025 Omar Kamali
Metadata
Release files for babelvec 0.1.7
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| babelvec-0.1.7.tar.gz | 39.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| babelvec-0.1.7-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 87.1 kB
Release files / babelvec-0.1.7.tar.gz
| Download URL | babelvec-0.1.7.tar.gz |
|---|---|
| Size | 39.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6bde24d3b231b4564457f7331d51aeab735fb799c6a335a17d57ead12073bbfc
|
|
BLAKE2b-256 checksum How to use checksums |
d86af17302ada4be9417bffc1d06df5c67e98901f36e42dab45709b5058d5493
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jan 9, 2026.
Transparency logRelease files / babelvec-0.1.7-py3-none-any.whl
| Download URL | babelvec-0.1.7-py3-none-any.whl |
|---|---|
| Size | 48.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7c6dc78df63860fec1b078b919800278529da86dc6349d91bc18842281d0faf9
|
|
BLAKE2b-256 checksum How to use checksums |
9f341bea0866d3c8c63c91731d76a4377a791d95436008691af4f28b1be38a72
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jan 9, 2026.
Transparency log