libembedding
Fast ONNX-based text, image, and sparse embeddings for Python. 5-8x faster than fastembed with 3.5x less memory.
Built on a C/C++ backend using ONNX Runtime, exposed to Python via zero-overhead cffi bindings. Supports 44 text embedding models, 5 image models, 2 sparse models, and 4 rerankers with automatic model downloading from HuggingFace Hub.
Installation
pip install libembedding
Requirements: ONNX Runtime must be installed on your system.
# macOS
brew install onnxruntime
# Ubuntu/Debian
apt install libonnxruntime-dev
# Or set ONNXRUNTIME_ROOT to your installation path
Quick Start
Text Embeddings
from libembedding import TextEmbedding
model = TextEmbedding("BAAI/bge-small-en-v1.5")
embeddings = model.embed(["Hello world", "How are you?"])
print(embeddings.shape) # (2, 384)
print(embeddings.dtype) # float32
Sparse Embeddings
from libembedding import SparseTextEmbedding
model = SparseTextEmbedding()
results = model.embed(["machine learning algorithms"])
for r in results:
print(r.indices.shape, r.values.shape)
Image Embeddings
from libembedding import ImageEmbedding
model = ImageEmbedding()
embeddings = model.embed_files(["photo.jpg", "diagram.png"])
Reranking
from libembedding import Reranker
reranker = Reranker("BAAI/bge-reranker-base")
results = reranker.rerank(
"What is deep learning?",
[
"Deep learning uses neural networks with many layers",
"The weather is sunny today",
"Neural networks are inspired by biological brains",
],
)
for r in results:
print(f"doc[{r.index}] score={r.score:.4f}")
Model Discovery
import libembedding
for m in libembedding.list_text_models():
print(f"{m.model_name:45} dim={m.dim:<5} {m.pooling}")
API Reference
TextEmbedding
TextEmbedding(
model_name="BAAI/bge-small-en-v1.5", # HuggingFace model name, repo code, or local dir path
provider="cpu", # "cpu", "cuda", "coreml", "directml", "tensorrt"
device_id=0,
cache_dir=None, # None = ~/.cache/libembedding
max_length=0, # 0 = model default
threads=0, # 0 = auto
batch_size=256, # internal batch size for embedding
offline=False, # True = use cache only, never download
show_download_progress=True,
dim=0, # embedding dim for local models without config.json
pooling="mean", # "cls" or "mean" for local models
num_threads=0, # deprecated, use threads
)
| Method | Returns | Description |
|---|---|---|
embed(texts, batch_size=None) |
np.ndarray (n, dim) |
L2-normalized dense embeddings |
embed_stream(texts, batch_size=None) |
generator | Yields one embedding at a time (low memory) |
dim |
int |
Embedding dimension |
name |
str |
Model name or local path |
info() |
ModelDesc |
Runtime model descriptor |
max_length() |
int |
Max token length for the model |
stats() |
Stats |
Runtime statistics (texts, batches, latency) |
close() |
None |
Release resources |
SparseTextEmbedding
SparseTextEmbedding(model_name="prithvida/SPLADE_PP_en_v1", ...)
| Method | Returns | Description |
|---|---|---|
embed(texts, batch_size=0) |
list[SparseEmbedding] |
Sparse vectors with .indices and .values |
embed_stream(texts, batch_size=None) |
generator | Yields one embedding at a time |
dim |
int |
Embedding dimension (0 = dynamic) |
name |
str |
Model name or local path |
info() |
ModelDesc |
Runtime model descriptor |
max_length() |
int |
Max token length |
stats() |
Stats |
Runtime statistics |
ImageEmbedding
ImageEmbedding(model_name="Qdrant/clip-ViT-B-32-vision", ...)
| Method | Returns | Description |
|---|---|---|
embed_files(paths, batch_size=0) |
np.ndarray (n, dim) |
Embed from file paths |
embed_bytes(images, batch_size=0) |
np.ndarray (n, dim) |
Embed from raw bytes |
dim |
int |
Embedding dimension |
name |
str |
Model name or local path |
info() |
ModelDesc |
Runtime model descriptor |
stats() |
Stats |
Runtime statistics |
Reranker
Reranker(model_name="BAAI/bge-reranker-base", ...)
| Method | Returns | Description |
|---|---|---|
rerank(query, documents, batch_size=0) |
list[RerankResult] |
Sorted by score descending |
name |
str |
Model name or local path |
info() |
ModelDesc |
Runtime model descriptor |
max_length() |
int |
Max token length |
stats() |
Stats |
Runtime statistics |
All classes support context managers (with TextEmbedding(...) as model:).
Similarity Functions
from libembedding import cosine_similarity, dot_product, euclidean_distance
import numpy as np
a = np.array([1.0, 2.0, 3.0], dtype=np.float32)
b = np.array([1.0, 2.0, 3.0], dtype=np.float32)
print(cosine_similarity(a, b)) # 1.0 (identical)
print(dot_product(a, b)) # 14.0
print(euclidean_distance(a, b)) # 0.0
Streaming Embeddings
Process large document sets without allocating a single result array:
from libembedding import TextEmbedding
with TextEmbedding("BAAI/bge-small-en-v1.5") as model:
for embedding in model.embed_stream(
["doc1", "doc2", ...], batch_size=32
):
# Each iteration yields a single (dim,) numpy array
process(embedding)
Runtime Statistics
with TextEmbedding("BAAI/bge-small-en-v1.5") as model:
model.embed(["text 1", "text 2", "text 3"])
stats = model.stats()
print(f"Embedded {stats.texts_embedded} texts "
f"({stats.batches_run} batches), "
f"avg latency {stats.avg_latency_ms:.2f}ms")
Data Types
ModelDesc — runtime model descriptor (returned by info()):
| Field | Type | Description |
|---|---|---|
name |
str |
Model name or local path |
dimension |
int |
Embedding dimension |
max_length |
int |
Max token length |
pooling |
str |
"cls" or "mean" |
num_threads |
int |
Threads configured |
batch_size |
int |
Batch size configured |
provider |
str |
Execution provider ("cpu", "cuda", "directml", "coreml") |
device_id |
int |
Device ID |
Stats — runtime statistics (returned by stats()):
| Field | Type | Description |
|---|---|---|
texts_embedded |
int |
Total texts processed |
batches_run |
int |
Total ONNX inference batches |
avg_latency_ms |
float |
Average milliseconds per embed call |
Local Model Loading
from libembedding import TextEmbedding
# Load from a local directory containing model.onnx + tokenizer.json (+ config.json)
model = TextEmbedding("/path/to/model_dir")
Available Models
44 text models including BGE, MiniLM, Nomic, E5, CLIP, Jina, GTE, Snowflake, ModernBERT (with quantized variants).
5 image models including CLIP ViT-B/32, ResNet-50, Unicom, Nomic Vision.
2 sparse models: SPLADE++, BGE-M3.
4 reranker models: BGE Reranker, Jina Reranker.
Benchmarks
Measured on Apple M-series with all-MiniLM-L6-v2 (384-dim). Median of 10 runs.
| Metric | libembedding | fastembed | Speedup |
|---|---|---|---|
| Single text latency (ms) | 4.4 | 38.0 | 8.6x |
| Batch 8 (texts/sec) | 641 | 92 | 7.0x |
| Batch 32 (texts/sec) | 581 | 89 | 6.5x |
| Peak RSS (MB) | 567 | 1,981 | 3.5x less |
Configuration
| Environment Variable | Purpose |
|---|---|
LIBEMBEDDING_CACHE_DIR |
Override model cache directory |
FASTEMBED_CACHE_DIR |
Alternative cache dir (fastembed compatibility) |
HF_ENDPOINT |
Custom HuggingFace Hub endpoint |
Building & Publishing
Prerequisites
pip install build twine
Build the shared library + wheel
cd python/
# Step 1: Build the C/C++ shared library and copy it into the package
./setup.sh --build-only
# Step 2: Build sdist and wheel
python -m build
This produces files in dist/:
dist/
libembedding-0.2.0.tar.gz # source distribution
libembedding-0.2.0-py3-none-any.whl # wheel (includes bundled .dylib/.so)
Upload to PyPI
# Upload to TestPyPI first to verify
twine upload --repository testpypi dist/*
# Install from TestPyPI to verify
pip install --index-url https://test.pypi.org/simple/ libembedding
# Upload to production PyPI
twine upload dist/*
One-liner (build + upload)
./setup.sh --build-only && python -m build && twine upload dist/*
Platform-specific wheels
The default wheel is py3-none-any and bundles the shared library for the build platform. To build platform-tagged wheels for distribution:
# macOS (current arch)
./setup.sh --build-only
python -m build
# For other platforms, build on that platform or use cibuildwheel:
pip install cibuildwheel
cibuildwheel --platform linux # builds manylinux wheels
cibuildwheel --platform macos # builds macOS wheels
License
MIT
Release files for libembedding-ng 1.0.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distributions (wheels)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| libembedding_ng-1.0.2-py3-none-win_amd64.whl | Python 3 | none | Windows x86-64 | Details |
| libembedding_ng-1.0.2-py3-none-manylinux_2_34_x86_64.whl | Python 3 | none | Linux glibc 2.34+ x86-64 | Details |
| libembedding_ng-1.0.2-py3-none-macosx_14_0_universal2.whl | Python 3 | none | macOS 14.0+ universal2 (ARM64, x86-64) | Details |
Total release size: 18.2 MB
Release files / libembedding_ng-1.0.2-py3-none-win_amd64.whl
| Download URL | libembedding_ng-1.0.2-py3-none-win_amd64.whl |
|---|---|
| Size | 5.0 MB |
| Tags | Python 3 Windows x86-64 |
|
SHA-256 checksum How to use checksums |
8380fdae0aeaeeb7c021d8ac926e2bf315dd4c71f3335a44a78e6ea702098b7b
|
|
BLAKE2b-256 checksum How to use checksums |
632fc825e41e8f2e82aa0b014ed7ccae6f67c991bde0b9c7d7c00fba0797f341
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / libembedding_ng-1.0.2-py3-none-manylinux_2_34_x86_64.whl
| Download URL | libembedding_ng-1.0.2-py3-none-manylinux_2_34_x86_64.whl |
|---|---|
| Size | 13.0 MB |
| Tags | Linux glibc 2.34+ x86-64 Python 3 |
|
SHA-256 checksum How to use checksums |
8e0d9c9d1c93ae86203eff8da93b7a6326dbb636d897a65b43e1f0d5a26cc4ab
|
|
BLAKE2b-256 checksum How to use checksums |
91e1d0be6de2c686dbd2376459b94d73534c2b23b24d7e65ef2e0b4038f3334d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / libembedding_ng-1.0.2-py3-none-macosx_14_0_universal2.whl
| Download URL | libembedding_ng-1.0.2-py3-none-macosx_14_0_universal2.whl |
|---|---|
| Size | 187.0 kB |
| Tags | Python 3 macOS 14.0+ universal2 (ARM64, x86-64) |
|
SHA-256 checksum How to use checksums |
792c5f5725b6f9950c61eacd223a0589b868b6dbe499a2b1f93087ae0cdd5afb
|
|
BLAKE2b-256 checksum How to use checksums |
8345fdb6ea924e74c7722ec618d867fcec091ebfe01ce2da4c149f9daede0bc3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|