Skip to main content

Note: This repository is not actively maintained, but it is used to publish the pypi package to allow us to maintain the model within mteb (e.g. to solve issues such as 4777.

Banner

EBind: Multi-Modal Embeddings

EBind is a multi-modal embedding model that supports image, video, audio, text, and 3D point cloud inputs. All modalities are projected into a shared embedding space, enabling cross-modal similarity computation.

Installation

Option 1
If you want to work within the repository, use uv to install the necessary dependencies.

uv sync

Option 2
You can also install it as an external dependency for another project:

# Option 2.a
python -m pip install git@https://github.com/encord-team/ebind
# Option 2.b; or install a local, editable version
git clone https://github.com/encord-team/ebind
cd /path/to/your/project
python -m pip install -e /path/to/ebind

Loading the Model

import torch
from ebind import EBindModel, EBindProcessor

model = EBindModel.from_pretrained("encord-team/ebind-full")
processor = EBindProcessor.from_pretrained("encord-team/ebind-full")

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = model.to(device).eval()
processor = processor.to(device)

Processing Multi-Modal Inputs

inputs = {
    "image": ["examples/dog.png", "examples/cat.png"],
    "video": ["examples/dog.mp4", "examples/cat.mp4"],
    "audio": ["examples/dog.mp4", "examples/cat.mp4"],
    "text": ["A dog is howling in the street", "A cat is sleeping on the couch"],
    "points": ["examples/dog_point_cloud.npy", "examples/cat_point_cloud.npy"],
}

with torch.inference_mode():
    batch = processor(inputs, return_tensors="pt")  # set text_file_paths=True if passing text file paths instead of strings
    outputs = model.forward(**batch)

Computing Cross-Modal Similarities

keys = list(outputs.keys())
for i, modality in enumerate(keys):
    for j, modality2 in enumerate(keys[i + 1:]):
        result = outputs[modality] @ outputs[modality2].T
        print(f"{modality} x {modality2}:")
        print(result.cpu().detach().numpy())
        print('='*26)

Expected Output:

image x video similarity: 
[[0.48 0.42]
 [0.41 0.6 ]]
==========================
image x audio similarity: 
[[0.07 0.05]
 [0.02 0.12]]
==========================
image x text similarity: 
[[0.16 0.07]
 [0.08 0.14]]
==========================
image x points similarity: 
[[0.2  0.19]
 [0.18 0.19]]
==========================
video x audio similarity: 
[[0.19 0.08]
 [0.03 0.16]]
==========================
video x text similarity: 
[[0.26 0.05]
 [0.11 0.14]]
==========================
video x points similarity: 
[[0.24 0.15]
 [0.17 0.26]]
==========================
audio x text similarity: 
[[ 0.12 -0.  ]
 [ 0.07  0.09]]
==========================
audio x points similarity: 
[[0.13 0.06]
 [0.1  0.12]]
==========================
text x points similarity: 
[[0.19 0.14]
 [0.05 0.18]]
==========================

Note: The image/video similarity is significantly higher because they share the same vision encoder.

Compile PointNet2 CUDA ops (optional)

If you have CUDA available, consider building the PointNet2 custom ops used for embedding point clouds to get faster inference:

cd src/ebind/models/uni3d/pointnet2_ops && \
    uv run python -c "import torch,sys; sys.exit(0 if torch.cuda.is_available() else 1)" && \
    MAX_JOBS=$(nproc) uv run python setup.py build_ext --inplace

We have modified the code slightly in src/ebind/models/uni3d/pointnet2_ops/pointnet2_utils.py to have a fallback torch implementation in order for the model to be executable on no-GPU hardware.

Contributing

We welcome contributions! If you have suggestions for improvements, new features, or bug fixes, feel free to open an issue or pull request. Please follow the standard GitHub workflow and adhere to our code style and guidelines. For major changes, we recommend discussing them in an issue before submitting a PR.

How to contribute

  1. Fork the repository.
  2. Create your feature branch: git checkout -b my-feature
  3. Commit your changes: git commit -m 'Add some feature'
  4. Push to the branch: git push origin my-feature
  5. Open a pull request describing your changes.

Citation

If you use this codebase in your research or work, please cite it as follows (replace with your own citation when available):

@misc{encord-bind,
  author       = {The Encord Team},
  title        = {{EBind}: Multi-modal binding and inference},
  year         = {2025},
  howpublished = {\url{https://github.com/encord-team/ebind}},
}

License

This project is licensed under the Attribution-NonCommercial-ShareAlike 4.0 International. See the LICENSE file for details.

Metadata

Release files for ebind 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ebind 0.1.0
File Size Uploaded
ebind-0.1.0.tar.gz 12.8 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for ebind 0.1.0
File Interpreter ABI Platform
ebind-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 14.2 MB

Release files / ebind-0.1.0.tar.gz

Download URL ebind-0.1.0.tar.gz
Size 12.8 MB
Tags Source
SHA-256 checksum
How to use checksums
1dec6169abcaed25956ce97f487489232ccee44e2077d5e74d3844fb55cf19bc
BLAKE2b-256 checksum
How to use checksums
319657064d83b2f3f378724086604531a52d33400d09c9ca14ca6a28bdbfb8b3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.16 {"installer":{"name":"uv","version":"0.11.16","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / ebind-0.1.0-py3-none-any.whl

Download URL ebind-0.1.0-py3-none-any.whl
Size 1.4 MB
Tags Python 3
SHA-256 checksum
How to use checksums
79aa42936f58cb1e60fb2d4fd63204d11a12127f7d0ba7252155a61aa78355c9
BLAKE2b-256 checksum
How to use checksums
f99bd8395559010fdd32de2d9405211e1aa89f42549cb10f788e70ddfff712f7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.16 {"installer":{"name":"uv","version":"0.11.16","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page