Skip to main content

High-performance video memory encoding library with semantic search (Rust bindings)

Project description

memvid-rs (Python Bindings)

🇰🇷 한국어 (Korean) PyPI version

Memvid-rs combines the performance of Rust with the ease of use of Python. It encodes text data into QR codes and then into video frames, with semantic search capabilities powered by FAISS and SentenceTransformers.

This project exposes the core logic written in Rust (MemvidEncoder, MemvidDecoder) as a Python module via PyO3, with a high-level Python wrapper for semantic search.

Attribution: This project is a port reimplemented in Rust based on the ideas and design of Olow304/memvid. We deeply appreciate the innovative approach of the original project.

🚀 Key Features

  • High Performance: Leverages Rust's parallel processing and efficient memory management for fast text-to-video conversion.
  • Semantic Search: FAISS vector index with SentenceTransformer embeddings for <100ms search across millions of chunks.
  • Pythonic: Easy to install via pip and use naturally as a Python object.
  • Strong Compression: Significantly reduces storage space via Text -> QR -> Video (H.264/H.265) conversion.

📦 Installation

From PyPI (Recommended)

# Basic installation (Rust encoder/decoder only)
pip install memvid-rs

# Full installation with semantic search
pip install memvid-rs[full]

From Source (Development)

Requires Rust and Python development environments.

# 1. Install Rust & FFmpeg
brew install rust ffmpeg

# 2. Create and activate virtual environment
python3 -m venv .venv
source .venv/bin/activate

# 3. Install Maturin and dependencies
pip install maturin sentence-transformers faiss-cpu

# 4. Build and install
maturin develop --release

💻 Usage

Basic Usage (Rust Encoder Only)

import memvid_rs

# Initialize Encoder
encoder = memvid_rs.MemvidEncoder()

# Add Text (chunk_size=200, overlap=20)
text_data = "Memvid-rs is fast and efficient. " * 100
encoder.add_text(text_data, chunk_size=200, overlap=20)

# Build Video
encoder.build("output.mp4", "index.json")
print("Video created: output.mp4")

Full Usage with Semantic Search

from memvid import MemvidEncoder, MemvidRetriever

# === ENCODING ===
encoder = MemvidEncoder()

# Add text chunks
texts = [
    "Python is a high-level programming language.",
    "Machine learning uses algorithms to learn from data.",
    "Rust is focused on safety and performance.",
]
for text in texts:
    encoder.add_text(text, chunk_size=512)

# Build video + FAISS index
encoder.build_video("memory.mp4", "index")
# Creates: memory.mp4, index.faiss, index.json

# === RETRIEVAL ===
retriever = MemvidRetriever("memory.mp4", "index")

# Semantic search
results = retriever.search("programming language", top_k=3)
for text, score, metadata in results:
    print(f"[{score:.3f}] {text}")

# Search with context
results = retriever.search_with_context(
    "artificial intelligence",
    top_k=2,
    context_before=1,
    context_after=1
)

🔧 API Reference

memvid.MemvidEncoder

High-level encoder with semantic search support.

encoder = MemvidEncoder(
    model_name="all-MiniLM-L6-v2",  # SentenceTransformer model
    index_type="Flat"               # FAISS index type: "Flat" or "IVF"
)

encoder.add_text(text, chunk_size=512, overlap=50)
encoder.add_pdf("document.pdf")  # Requires PyPDF2
encoder.build_video("output.mp4", "index")

memvid.MemvidRetriever

Semantic search and retrieval from QR videos.

retriever = MemvidRetriever(
    video_path="memory.mp4",
    index_path="index",
    cache_size=100  # Frame cache size
)

# Search methods
results = retriever.search(query, top_k=5)
results = retriever.search_with_context(query, top_k=5, context_before=1, context_after=1)
text = retriever.get_chunk(chunk_id)
all_text = retriever.get_all_text()
stats = retriever.get_stats()

memvid_rs (Low-level Rust API)

import memvid_rs

# Encoder
encoder = memvid_rs.MemvidEncoder()
encoder.add_text(text, chunk_size, overlap)
encoder.build(video_path, index_path)
chunks = encoder.get_chunks()

# Decoder
decoder = memvid_rs.MemvidDecoder(video_path, index_path)
text = decoder.get_chunk_text(chunk_id)
texts = decoder.get_chunks_text([0, 1, 2])
all_chunks = decoder.get_all_chunks()
info = decoder.get_video_info()

⚠️ Requirements

  • System: macOS (Tested), Linux, Windows
  • Runtime: ffmpeg (Required for video encoding)
    • macOS: brew install ffmpeg
  • Python: >= 3.8
  • Optional: sentence-transformers, faiss-cpu for semantic search

📄 License

MIT License

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

memvid_rs-0.2.0.tar.gz (256.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

memvid_rs-0.2.0-cp314-cp314-macosx_11_0_arm64.whl (2.0 MB view details)

Uploaded CPython 3.14macOS 11.0+ ARM64

File details

Details for the file memvid_rs-0.2.0.tar.gz.

File metadata

  • Download URL: memvid_rs-0.2.0.tar.gz
  • Upload date:
  • Size: 256.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: maturin/1.10.2

File hashes

Hashes for memvid_rs-0.2.0.tar.gz
Algorithm Hash digest
SHA256 ca926de836d330c7618f742058f0f302bac594bd72ea2c2f0486c50731950c70
MD5 57946f9692924092a617c3c710043461
BLAKE2b-256 5551e18ed8aff0a8bef9fe6c0f0fb9ec5e4f85cb5a155cc9402e1addeb21b225

See more details on using hashes here.

File details

Details for the file memvid_rs-0.2.0-cp314-cp314-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for memvid_rs-0.2.0-cp314-cp314-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 131183235526a9eb42b107bc15e048196b71e860795c5d9f0f864645ab3d6cfb
MD5 e3922e8b253f7f2db5defa8f81815484
BLAKE2b-256 a95767216081fadaf3d1b34b274a6fb7aee7628fd2dfdce3740ec320dbdaa36b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page