Skip to main content

libembedding

Fast ONNX-based text, image, and sparse embeddings for Python. 5-8x faster than fastembed with 3.5x less memory.

Built on a C/C++ backend using ONNX Runtime, exposed to Python via zero-overhead cffi bindings. Supports 44 text embedding models, 5 image models, 2 sparse models, and 4 rerankers with automatic model downloading from HuggingFace Hub.

Installation

pip install libembedding

Requirements: ONNX Runtime must be installed on your system.

# macOS
brew install onnxruntime

# Ubuntu/Debian
apt install libonnxruntime-dev

# Or set ONNXRUNTIME_ROOT to your installation path

Quick Start

Text Embeddings

from libembedding import TextEmbedding

model = TextEmbedding("BAAI/bge-small-en-v1.5")
embeddings = model.embed(["Hello world", "How are you?"])

print(embeddings.shape)  # (2, 384)
print(embeddings.dtype)  # float32

Sparse Embeddings

from libembedding import SparseTextEmbedding

model = SparseTextEmbedding()
results = model.embed(["machine learning algorithms"])

for r in results:
    print(r.indices.shape, r.values.shape)

Image Embeddings

from libembedding import ImageEmbedding

model = ImageEmbedding()
embeddings = model.embed_files(["photo.jpg", "diagram.png"])

Reranking

from libembedding import Reranker

reranker = Reranker("BAAI/bge-reranker-base")
results = reranker.rerank(
    "What is deep learning?",
    [
        "Deep learning uses neural networks with many layers",
        "The weather is sunny today",
        "Neural networks are inspired by biological brains",
    ],
)
for r in results:
    print(f"doc[{r.index}] score={r.score:.4f}")

Model Discovery

import libembedding

for m in libembedding.list_text_models():
    print(f"{m.model_name:45} dim={m.dim:<5} {m.pooling}")

API Reference

TextEmbedding

TextEmbedding(
    model_name="BAAI/bge-small-en-v1.5",  # HuggingFace model name, repo code, or local dir path
    provider="cpu",                         # "cpu", "cuda", "coreml", "directml", "tensorrt"
    device_id=0,
    cache_dir=None,                         # None = ~/.cache/libembedding
    max_length=0,                           # 0 = model default
    threads=0,                              # 0 = auto
    batch_size=256,                         # internal batch size for embedding
    offline=False,                          # True = use cache only, never download
    show_download_progress=True,
    dim=0,                                  # embedding dim for local models without config.json
    pooling="mean",                         # "cls" or "mean" for local models
    num_threads=0,                          # deprecated, use threads
)
Method Returns Description
embed(texts, batch_size=None) np.ndarray (n, dim) L2-normalized dense embeddings
embed_stream(texts, batch_size=None) generator Yields one embedding at a time (low memory)
dim int Embedding dimension
name str Model name or local path
info() ModelDesc Runtime model descriptor
max_length() int Max token length for the model
stats() Stats Runtime statistics (texts, batches, latency)
close() None Release resources

SparseTextEmbedding

SparseTextEmbedding(model_name="prithvida/SPLADE_PP_en_v1", ...)
Method Returns Description
embed(texts, batch_size=0) list[SparseEmbedding] Sparse vectors with .indices and .values
embed_stream(texts, batch_size=None) generator Yields one embedding at a time
dim int Embedding dimension (0 = dynamic)
name str Model name or local path
info() ModelDesc Runtime model descriptor
max_length() int Max token length
stats() Stats Runtime statistics

ImageEmbedding

ImageEmbedding(model_name="Qdrant/clip-ViT-B-32-vision", ...)
Method Returns Description
embed_files(paths, batch_size=0) np.ndarray (n, dim) Embed from file paths
embed_bytes(images, batch_size=0) np.ndarray (n, dim) Embed from raw bytes
dim int Embedding dimension
name str Model name or local path
info() ModelDesc Runtime model descriptor
stats() Stats Runtime statistics

Reranker

Reranker(model_name="BAAI/bge-reranker-base", ...)
Method Returns Description
rerank(query, documents, batch_size=0) list[RerankResult] Sorted by score descending
name str Model name or local path
info() ModelDesc Runtime model descriptor
max_length() int Max token length
stats() Stats Runtime statistics

All classes support context managers (with TextEmbedding(...) as model:).

Similarity Functions

from libembedding import cosine_similarity, dot_product, euclidean_distance
import numpy as np

a = np.array([1.0, 2.0, 3.0], dtype=np.float32)
b = np.array([1.0, 2.0, 3.0], dtype=np.float32)

print(cosine_similarity(a, b))   # 1.0 (identical)
print(dot_product(a, b))         # 14.0
print(euclidean_distance(a, b))  # 0.0

Streaming Embeddings

Process large document sets without allocating a single result array:

from libembedding import TextEmbedding

with TextEmbedding("BAAI/bge-small-en-v1.5") as model:
    for embedding in model.embed_stream(
        ["doc1", "doc2", ...], batch_size=32
    ):
        # Each iteration yields a single (dim,) numpy array
        process(embedding)

Runtime Statistics

with TextEmbedding("BAAI/bge-small-en-v1.5") as model:
    model.embed(["text 1", "text 2", "text 3"])
    stats = model.stats()
    print(f"Embedded {stats.texts_embedded} texts "
          f"({stats.batches_run} batches), "
          f"avg latency {stats.avg_latency_ms:.2f}ms")

Data Types

ModelDesc — runtime model descriptor (returned by info()):

Field Type Description
name str Model name or local path
dimension int Embedding dimension
max_length int Max token length
pooling str "cls" or "mean"
num_threads int Threads configured
batch_size int Batch size configured
provider str Execution provider ("cpu", "cuda", "directml", "coreml")
device_id int Device ID

Stats — runtime statistics (returned by stats()):

Field Type Description
texts_embedded int Total texts processed
batches_run int Total ONNX inference batches
avg_latency_ms float Average milliseconds per embed call

Local Model Loading

from libembedding import TextEmbedding

# Load from a local directory containing model.onnx + tokenizer.json (+ config.json)
model = TextEmbedding("/path/to/model_dir")

Available Models

44 text models including BGE, MiniLM, Nomic, E5, CLIP, Jina, GTE, Snowflake, ModernBERT (with quantized variants).

5 image models including CLIP ViT-B/32, ResNet-50, Unicom, Nomic Vision.

2 sparse models: SPLADE++, BGE-M3.

4 reranker models: BGE Reranker, Jina Reranker.

Benchmarks

Measured on Apple M-series with all-MiniLM-L6-v2 (384-dim). Median of 10 runs.

Metric libembedding fastembed Speedup
Single text latency (ms) 4.4 38.0 8.6x
Batch 8 (texts/sec) 641 92 7.0x
Batch 32 (texts/sec) 581 89 6.5x
Peak RSS (MB) 567 1,981 3.5x less

Configuration

Environment Variable Purpose
LIBEMBEDDING_CACHE_DIR Override model cache directory
FASTEMBED_CACHE_DIR Alternative cache dir (fastembed compatibility)
HF_ENDPOINT Custom HuggingFace Hub endpoint

Building & Publishing

Prerequisites

pip install build twine

Build the shared library + wheel

cd python/

# Step 1: Build the C/C++ shared library and copy it into the package
./setup.sh --build-only

# Step 2: Build sdist and wheel
python -m build

This produces files in dist/:

dist/
  libembedding-0.2.0.tar.gz              # source distribution
  libembedding-0.2.0-py3-none-any.whl    # wheel (includes bundled .dylib/.so)

Upload to PyPI

# Upload to TestPyPI first to verify
twine upload --repository testpypi dist/*

# Install from TestPyPI to verify
pip install --index-url https://test.pypi.org/simple/ libembedding

# Upload to production PyPI
twine upload dist/*

One-liner (build + upload)

./setup.sh --build-only && python -m build && twine upload dist/*

Platform-specific wheels

The default wheel is py3-none-any and bundles the shared library for the build platform. To build platform-tagged wheels for distribution:

# macOS (current arch)
./setup.sh --build-only
python -m build

# For other platforms, build on that platform or use cibuildwheel:
pip install cibuildwheel
cibuildwheel --platform linux   # builds manylinux wheels
cibuildwheel --platform macos   # builds macOS wheels

License

MIT

Release files for libembedding-ng 1.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for libembedding-ng 1.0.2
File Interpreter ABI Platform
libembedding_ng-1.0.2-py3-none-win_amd64.whl Python 3 none Windows x86-64 Details
libembedding_ng-1.0.2-py3-none-manylinux_2_34_x86_64.whl Python 3 none Linux glibc 2.34+ x86-64 Details
libembedding_ng-1.0.2-py3-none-macosx_14_0_universal2.whl Python 3 none macOS 14.0+ universal2 (ARM64, x86-64) Details

Total release size: 18.2 MB

Release files / libembedding_ng-1.0.2-py3-none-win_amd64.whl

Download URL libembedding_ng-1.0.2-py3-none-win_amd64.whl
Size 5.0 MB
Tags Python 3 Windows x86-64
SHA-256 checksum
How to use checksums
8380fdae0aeaeeb7c021d8ac926e2bf315dd4c71f3335a44a78e6ea702098b7b
BLAKE2b-256 checksum
How to use checksums
632fc825e41e8f2e82aa0b014ed7ccae6f67c991bde0b9c7d7c00fba0797f341
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / libembedding_ng-1.0.2-py3-none-manylinux_2_34_x86_64.whl

Download URL libembedding_ng-1.0.2-py3-none-manylinux_2_34_x86_64.whl
Size 13.0 MB
Tags Linux glibc 2.34+ x86-64 Python 3
SHA-256 checksum
How to use checksums
8e0d9c9d1c93ae86203eff8da93b7a6326dbb636d897a65b43e1f0d5a26cc4ab
BLAKE2b-256 checksum
How to use checksums
91e1d0be6de2c686dbd2376459b94d73534c2b23b24d7e65ef2e0b4038f3334d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / libembedding_ng-1.0.2-py3-none-macosx_14_0_universal2.whl

Download URL libembedding_ng-1.0.2-py3-none-macosx_14_0_universal2.whl
Size 187.0 kB
Tags Python 3 macOS 14.0+ universal2 (ARM64, x86-64)
SHA-256 checksum
How to use checksums
792c5f5725b6f9950c61eacd223a0589b868b6dbe499a2b1f93087ae0cdd5afb
BLAKE2b-256 checksum
How to use checksums
8345fdb6ea924e74c7722ec618d867fcec091ebfe01ce2da4c149f9daede0bc3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

1.9.0

3 release files

1.8.0

3 release files

1.5.9

3 release files

1.2.1

3 release files

This release

1.0.2 This release

3 release files

1.0.1

3 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page