dashAI Vision-Language plugin
Zero-shot image-classification plugin for dashAI. Classify images into any set of labels you choose, with no training or fine-tuning, using pretrained Hugging Face vision-language models.
It adds seven components to dashAI's image classification task, all Zero-Shot: CLIP (three sizes), SigLIP, ALIGN, AltCLIP, and MetaCLIP 2.
Installation
From dashAI (recommended)
Open the Plugins page in dashAI, search for dashai-vision-language-plugin, and install it with one click. No restart is needed.
With pip
Install it in the same Python environment where dashAI runs:
pip install dashai-vision-language-plugin
Restart dashAI afterwards. It discovers the plugin automatically through the dashai.plugins entry point.
Requirements
- Python 3.10 or newer
- dashAI 0.9.7.post2 or newer
- Internet access the first time you use a model, to download its checkpoint from the Hugging Face Hub (see System Requirements)
Training downloads the selected Hugging Face checkpoint unless it is already present in the local Hugging Face cache. After loading a saved checkpoint (see Persistence), the checkpoint is instead downloaded lazily before the first prediction.
Use in dashAI
- Create or load an image dataset and select its image column.
- Create an
ImageClassificationTaskwith the dataset's categorical label column. - In the task's component selector, choose one of the Zero-Shot components (a CLIP size, SigLIP, ALIGN, AltCLIP, or MetaCLIP 2).
- Configure the model, prompt template, batch size, and device.
- Run the task, then evaluate its predictions with the normal dashAI evaluation flow.
The labels in the training dataset define the zero-shot classes. For labels cat, dog, and car, the default template produces the prompts a photo of a cat, a photo of a dog, and a photo of a car.
train() does not fine-tune the backbone. It records the dataset class labels and computes their text embeddings; image inference compares each image embedding with those prompts.
Components
CLIP Zero-Shot
Unlike the other components below, the CLIP checkpoint is not a free-text parameter. Following dashAI's own convention for same-family, different-size models (e.g. ResNet18ImageClassifier/ResNet50ImageClassifier), each CLIP size is its own component with the Hugging Face checkpoint fixed in code, so users can pick a named option instead of typing a model ID:
| Component | Fixed checkpoint | Trade-off |
|---|---|---|
CLIPViTB32ZeroShotClassifier |
openai/clip-vit-base-patch32 |
Fastest, smallest; the default choice. |
CLIPViTB16ZeroShotClassifier |
openai/clip-vit-base-patch16 |
Mid-sized: more accurate than B/32, lighter than L/14. |
CLIPViTL14ZeroShotClassifier |
openai/clip-vit-large-patch14 |
Largest, most accurate, slowest, heaviest download. |
All three share the same other parameters:
| Parameter | Default | Meaning |
|---|---|---|
prompt_template |
a photo of a {} |
Text prompt with exactly one {} label placeholder. |
batch_size |
32 |
Number of images processed per inference batch; must be at least 1. |
device |
auto |
auto uses CUDA when available, otherwise CPU; cpu forces CPU; cuda requires CUDA to be available. |
They score each image-label pair with CLIP's joint softmax over labels (softmax(logit_scale.exp() * cosine_similarity)), so the reported probabilities are a standard categorical distribution. Need a different CLIP checkpoint than these three (e.g. a fine-tuned or community variant)? Subclass CLIPZeroShotClassifier and set MODEL_NAME; that's exactly what these three components do.
SigLIP Zero-Shot (SigLIPZeroShotClassifier)
| Parameter | Default | Meaning |
|---|---|---|
model_name |
google/siglip-base-patch16-224 |
Hugging Face SigLIP checkpoint ID. |
prompt_template |
a photo of a {} |
Text prompt with exactly one {} label placeholder. |
batch_size |
32 |
Number of images processed together during inference; must be at least 1. |
device |
auto |
auto uses CUDA when available, otherwise CPU; cpu forces CPU; cuda requires CUDA to be available. |
SigLIP was trained with an independent sigmoid loss per image-label pair (sigmoid(logit_scale.exp() * cosine_similarity + logit_bias)), not a joint softmax, so out of the box its per-label scores don't sum to 1 across labels. To stay compatible with dashAI's classification metrics (which expect a per-sample categorical distribution), this component renormalizes the sigmoid scores to sum to 1. This changes only the reported non-argmax probabilities, not the predicted class.
SigLIP2: the fixed-resolution SigLIP2 checkpoints (e.g. google/siglip2-base-patch16-224) declare model_type: "siglip" in their config and load correctly through this same SigLIPZeroShotClassifier component; just set model_name to a SigLIP2 checkpoint ID, no separate component needed. Only the variable resolution "NaFlex" SigLIP2 checkpoints require the newer Siglip2Model/Siglip2Processor classes with different image handling (pixel_attention_mask, spatial_shapes), which this plugin does not implement.
ALIGN Zero-Shot (ALIGNZeroShotClassifier)
| Parameter | Default | Meaning |
|---|---|---|
model_name |
kakaobrain/align-base |
Hugging Face ALIGN checkpoint ID. |
prompt_template |
a photo of a {} |
Text prompt with exactly one {} label placeholder. |
batch_size |
32 |
Number of images processed together during inference; must be at least 1. |
device |
auto |
auto uses CUDA when available, otherwise CPU; cpu forces CPU; cuda requires CUDA to be available. |
ALIGN pairs an EfficientNet vision encoder with a BERT text encoder. Unlike CLIP's logit_scale.exp() multiplier, ALIGN divides cosine similarity by a learned scalar temperature before a joint softmax over labels (softmax(cosine_similarity / temperature)), so its reported probabilities are already a standard categorical distribution.
AltCLIP Zero-Shot (AltCLIPZeroShotClassifier)
| Parameter | Default | Meaning |
|---|---|---|
model_name |
BAAI/AltCLIP |
Hugging Face AltCLIP checkpoint ID. |
prompt_template |
a photo of a {} |
Text prompt with exactly one {} label placeholder. |
batch_size |
32 |
Number of images processed together during inference; must be at least 1. |
device |
auto |
auto uses CUDA when available, otherwise CPU; cpu forces CPU; cuda requires CUDA to be available. |
AltCLIP swaps CLIP's text encoder for a multilingual XLM-R encoder while keeping CLIP's exact joint softmax scoring, making it a good fit for class labels or prompts written in languages other than English. It is a larger checkpoint than the others (CLIP ViT-L/14 + XLM-R-large).
MetaCLIP 2 Zero-Shot (MetaCLIP2ZeroShotClassifier)
| Parameter | Default | Meaning |
|---|---|---|
model_name |
facebook/metaclip-2-worldwide-s16 |
Hugging Face MetaCLIP 2 checkpoint ID. |
prompt_template |
a photo of a {} |
Text prompt with exactly one {} label placeholder. |
batch_size |
32 |
Number of images processed together during inference; must be at least 1. |
device |
auto |
auto uses CUDA when available, otherwise CPU; cpu forces CPU; cuda requires CUDA to be available. |
MetaCLIP 2 uses the same joint softmax scoring and processor as CLIP, but is trained on 300+ languages, so it is a multilingual alternative to AltCLIP. The default checkpoint is the smallest published MetaCLIP 2 variant (ViT-S/16); larger facebook/metaclip-2-worldwide-* checkpoints are available for higher accuracy at a larger download/RAM cost.
System Requirements
Measured against CLIPViTB32ZeroShotClassifier's checkpoint (openai/clip-vit-base-patch32). SigLIP, ALIGN, and MetaCLIP 2's default checkpoints are similar in size or smaller; CLIPViTL14ZeroShotClassifier and AltCLIP's checkpoint (BAAI/AltCLIP) are notably larger, so budget more disk and RAM if you use them.
| Resource | Requirement |
|---|---|
| Disk | ~2GB total for one default checkpoint: ~700MB for torch + transformers (already included in dashAI), plus ~1.1GB cached. Each additional cached checkpoint adds its own size; AltCLIP alone can add several GB. |
| RAM / CPU | Runs comfortably on CPU with 4-8GB RAM for the small default checkpoints (~150-200M parameters); AltCLIP needs more. |
| GPU | Optional. device="cuda"/"auto" use CUDA when available (including ROCm builds of PyTorch, which expose the same CUDA API). Apple Silicon (MPS) is not supported; this matches dashAI's own models, which only distinguish CUDA vs. CPU. |
| Network | Required on first use to download the checkpoint from the Hugging Face Hub, unless it is already cached locally. |
| Software | Python >=3.10, dashAI >=0.9.7.post2 (see pyproject.toml). |
Persistence
Saved checkpoints are intentionally lightweight: they keep component configuration and class labels, not the backbone's model weights. Loading a checkpoint requires the named Hugging Face checkpoint to be available again; it will be re-downloaded when absent from the local cache.
Contributing
Issues and pull requests are welcome at github.com/diegoolguinw/dashai_vision_language_plugin.
Set up a development environment:
git clone https://github.com/diegoolguinw/dashai_vision_language_plugin.git
cd dashai_vision_language_plugin
python -m venv .venv
# Windows: .venv\Scripts\activate macOS/Linux: source .venv/bin/activate
pip install -e '.[dev]'
Run the ordinary unit suite (the real-model integration test is excluded by default):
python -m pytest
Lint the project:
python -m ruff check .
Build source and wheel distributions:
python -m build
Run the opt-in smoke tests. This downloads and executes all seven real checkpoints if needed (including the larger CLIP ViT-L/14 and AltCLIP checkpoints):
RUN_CLIP_INTEGRATION=1 python -m pytest -m integration -v
Author
Diego Olguin-Wende - dolguin at dim dot uchile dot cl
Bug reports and questions: GitHub Issues.
License
Released under the MIT License.
Metadata
Release files for dashai-vision-language-plugin 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dashai_vision_language_plugin-0.1.1.tar.gz | 20.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dashai_vision_language_plugin-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 47.9 kB
Release files / dashai_vision_language_plugin-0.1.1.tar.gz
| Download URL | dashai_vision_language_plugin-0.1.1.tar.gz |
|---|---|
| Size | 20.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
24b4152dbd0be10706377681ce67f75353368ea0deab50cd884886a353fa7133
|
|
BLAKE2b-256 checksum How to use checksums |
d557e4ff635fc72fa6710b833780480b81640f320b6ae0868c6099625faf8fca
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / dashai_vision_language_plugin-0.1.1-py3-none-any.whl
| Download URL | dashai_vision_language_plugin-0.1.1-py3-none-any.whl |
|---|---|
| Size | 27.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9bf5df8b22f965ca596a0d62f94850c98c367d6cbbfd92058cd7704e0f1a0f08
|
|
BLAKE2b-256 checksum How to use checksums |
5da5d38562f16574d9ebd43696c78288010e1078298b28eb62416455a01371a7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log