Intent-aware, multi-model captioning pipeline for LoRA training, dataset curation, and generative AI workflows
Project description
Argus Lens
One image. A hundred perspectives.
Multi-model captioning pipeline for LoRA training, dataset curation, and generative AI workflows. Unlike traditional captioners that describe an image, Argus Lens produces intent-aware caption subsets -- structured, filtered, and optimised for how captions are actually used downstream.
Looking for the web UI? See argus-vision-demo -- a thin Next.js frontend for exploring argus-lens interactively.
Quick Start
pip install argus-lens[openai]
from argus_lens import ArgusLens
engine = ArgusLens(backend="openai", api_key="sk-...")
result = engine.caption("photo.jpg", trigger_word="sks_person")
print(result.final_caption)
print(result.caption_variants["training"])
print(result.caption_variants["zeroshot"])
Why Argus Lens?
Most captioning tools answer "what's in this image?" -- Argus Lens answers "how will this caption be used?"
BLIP gives you a sentence. WD14 gives you a flat tag list. Neither knows whether you're training a LoRA, curating a dataset, or generating zero-shot. Argus Lens produces multiple structured caption variants from the same image, each optimised for a different downstream task.
| Typical captioner | Argus Lens | |
|---|---|---|
| Output | Single flat string | Structured, category-bucketed variants |
| Intent awareness | None -- describes the image | Training, zero-shot, and per-category variants |
| Multi-model fusion | One model at a time | Hybrid pipelines (e.g. WD14 tags + GPT-4o prose) |
| Token budget management | Manual / none | Auto-tuned for SDXL (CLIP, 60 tokens) vs Flux/SD3 (T5, 200 tokens) |
| LoRA training support | Raw output, hope for the best | Identity suppression, omission cycles, tiered tag protection |
| Dataset workflow focus | Captioning only | Batch processing, deduplication, export to txt/JSON/JSONL/CSV |
| Transparency | Black box | Every removed phrase and compaction decision is reported |
What the assembly pipeline does
The model is just an input source. The real value is what happens after inference:
- Classify -- every tag and prose clause is bucketed into identity, wardrobe, pose/composition, setting, lighting, or action
- Filter -- redundant prose (>50% overlap with tags), filler prefixes ("The image shows..."), and training noise (rating tags, meta tokens) are stripped
- Specialise -- the training variant suppresses identity traits (the LoRA learns these visually; stating them conflicts with the base model's prior), while the zero-shot variant puts identity first (no LoRA to supply it)
- Budget -- fragments are assembled under the correct CLIP/T5 token limit with tiered protection (framing tags are never dropped, short generic tags are shed first)
- Diversify -- omission cycles systematically suppress different category buckets across images, creating stronger concept disentanglement across the dataset
- Enrich -- novel compound nouns from prose output ("gray sweater", "wooden door") are extracted and appended at lowest priority to fill remaining token budget
Features
- Multi-model backends: WD14, Florence-2, BLIP-2 (local GPU/CPU) + OpenAI, HuggingFace, Replicate, NVIDIA NIM, and any OpenAI-compatible server — Ollama, vLLM, LM Studio (cloud or self-hosted API)
- Structured captions: Category-bucketed variants (identity, wardrobe, pose, setting, lighting, action)
- Training-optimised: Tiered tag protection, omission cycles, CLIP/T5 token budgets, identity suppression
- Zero-shot variant: Identity-first, prose-preferred captions for generation without LoRA
- Hybrid pipelines: Mix local + cloud backends (e.g. WD14 tags + GPT-4o prose)
- Backend-aware budgets: Automatic token limits for SDXL (60), Flux (200), SD3 (200)
- Deterministic and filterable: Category-aware outputs you can audit, filter, and override
- CLI + Server: Command-line tool and optional FastAPI micro-server
- Export formats:
.txtsidecars, JSON, JSONL, CSV
Installation
pip handles all Python dependencies through extras. Pick the extras that match your use case:
# Assembly engine only (no model deps)
pip install argus-lens
# Local backends (GPU inference)
pip install argus-lens[local] # WD14 + Florence-2 + BLIP-2
pip install argus-lens[wd14] # WD14 only (CPU, no torch)
pip install argus-lens[wd14-gpu] # WD14 only, CUDA onnxruntime (no torch)
pip install argus-lens[torch] # Florence-2 / BLIP-2 only
# Cloud backends (no GPU needed)
pip install argus-lens[openai] # GPT-4o vision
pip install argus-lens[replicate] # Replicate API
# openai-compat, hf-inference, nvidia-nim need no extra — only the core httpx dep
# CLI (the `argus-lens` command; combine with your backend extras)
pip install argus-lens[cli,openai]
# Server (FastAPI + uvicorn; add [cli] for the `argus-lens serve` command)
pip install argus-lens[cli,server,local,openai]
# Everything
pip install argus-lens[all]
If you're adding argus-lens to an existing project, just add e.g. argus-lens[openai] to your requirements.txt -- pip resolves all transitive deps automatically.
System dependencies for local GPU backends
Cloud-only users ([openai], [replicate]) need no system packages -- skip this section.
Local backends ([local], [wd14], [torch]) require system libraries for image processing and (optionally) CUDA for GPU acceleration. On Ubuntu/Debian:
sudo apt install -y \
libgl1 libglib2.0-0 libxcb1 libsm6 libxext6 libxrender1
For GPU inference, you also need:
- NVIDIA GPU drivers (check with
nvidia-smi) - CUDA runtime (the
Dockerfile.gpu-basein this repo usesnvidia/cuda:12.4.1-runtime-ubuntu22.04as a reference) - NVIDIA Container Toolkit (for Docker deployment only)
If you already have torch and CUDA working in your environment, you're set -- the pip extras handle the rest.
Usage
Python API
Import and use directly in your code. This is the primary interface.
from argus_lens import ArgusLens
# Cloud backend -- works anywhere, no GPU
engine = ArgusLens(backend="openai", api_key="sk-...")
result = engine.caption("photo.jpg", trigger_word="sks_person")
# Local backend -- needs torch + GPU/CPU
engine = ArgusLens(backend="hybrid")
result = engine.caption("photo.jpg", trigger_word="sks_person")
# Self-hosted OpenAI-compatible server (Ollama, vLLM, LM Studio, ...)
# No API key needed for local servers; base_url defaults to Ollama localhost.
engine = ArgusLens(
backend="openai-compat",
base_url="http://localhost:11434/v1",
model_id="llama3.2-vision",
)
result = engine.caption("photo.jpg", trigger_word="sks_person")
# Batch processing
results = engine.caption_directory("./images/", output_format="txt")
CLI
The argus-lens command requires the [cli] extra (pip install argus-lens[cli,<backend>]).
# Caption a single image
argus-lens caption photo.jpg --trigger sks_person --backend openai
# Caption a directory, output as txt sidecars
argus-lens caption ./images/ --format txt --backend hybrid
# Self-hosted Ollama (or any OpenAI-compatible server)
argus-lens caption photo.jpg --backend openai-compat \
--base-url http://localhost:11434/v1 --model-id llama3.2-vision
# List available backends
argus-lens backends
HTTP Server
Run the built-in FastAPI server for frontend consumers (e.g. argus-vision-demo):
pip install argus-lens[cli,server,local]
argus-lens serve --cors --port 8080
Endpoints:
POST /caption-- multipart file uploadPOST /caption/url-- JSON body with image URLPOST /caption/batch-- multiple file uploadPOST /caption/stream-- NDJSON streaming for batchPOST /caption/manifest-- batch-caption an argus-curator JSONL manifest (sharedtarget_profile, writes.txtsidecars)POST /caption/manifest/stream-- streaming variant of/caption/manifest: one NDJSON progress line per image, then a completion summaryPOST /caption/folder-- batch-caption every image in a folder under the source root (optionally recursive, writes.txtsidecars)GET /folders?path=<rel>-- browse folders under--source-root/LENS_SOURCE_PATH(for the UI folder picker)GET /backends-- list available backendsPOST /v1/chat/completions-- OpenAI-compatible endpoint (always mounted; usable as a Frigate GenAI provider)
Immich
Immich is strong at CLIP search and faces but weak at descriptive keywords and captions -- the gap Argus Lens fills. Immich has no in-process ML plugin hook, so Argus Lens runs as a companion service: pull assets via the Immich API, caption them, and push keywords + description back.
Create an API key in Immich under Account Settings -> API Keys, then:
from argus_lens import ArgusLens
from argus_lens.connectors import ImmichSink, ImmichSource
IMMICH_URL = "http://immich.local:2283"
API_KEY = "..."
source = ImmichSource(IMMICH_URL, API_KEY)
sink = ImmichSink(IMMICH_URL, API_KEY)
engine = ArgusLens(backend="hybrid")
# Pass `since` (ISO 8601) to only process assets changed after your last run.
for ref in source.list_assets(since="2026-07-01T00:00:00Z"):
result = engine.caption(source.fetch_image(ref))
keywords = [t.strip() for t in result.raw_tags.split(",") if t.strip()]
sink.write(ref, keywords=keywords, description=result.caption_variants["zeroshot"])
Keywords are upserted as Immich tags (existing tags are reused) and attached to the asset; the description lands in the asset's description field, both searchable in the Immich UI. Writes are idempotent, so re-running over the same assets is safe.
If you'd rather keep Argus Lens decoupled from the Immich API entirely, use XmpSink instead: it writes standard .xmp sidecars next to your originals, which Immich (as well as Lightroom and digiKam) ingests on library scan.
Docker
For fresh hosts or isolated deployment with GPU passthrough. No pip install needed on the host.
# Build and run
./build-docker.sh
docker compose up
This builds a CUDA 12.4 base image, installs all extras into it, and runs argus-lens serve on port 8080.
Standalone image (GHCR)
A self-contained image built from Dockerfile.standalone is published to GHCR on each release — no local build needed:
docker run -p 8100:8100 -v /path/to/images:/data ghcr.io/smk762/argus-lens:latest
Folder browsing and captioning are confined to the mounted /data (override with LENS_SOURCE_PATH).
Configuration
Copy or create a .env file for the Docker deployment:
| Variable | Default | Description |
|---|---|---|
ARGUS_BACKEND |
hybrid |
Captioning backend (hybrid, wd14, florence2, openai, etc.) |
OPENAI_API_KEY / REPLICATE_API_TOKEN / HF_TOKEN / NVIDIA_API_KEY |
-- | API key for the matching cloud backend (each backend reads its own variable) |
ARGUS_PORT |
8080 |
Host port for the server |
WD14_MODEL_DIR |
~/.cache/wd14_tagger/ |
WD14 ONNX model directory (auto-downloads on first use) |
HF_HOME |
~/.cache/huggingface |
HuggingFace model cache (auto-downloads on first use) |
HF_TRUST_REMOTE_CODE |
false |
Only needed for legacy microsoft/Florence-2-* weights. See Security |
GPU prerequisites
# Verify NVIDIA driver
nvidia-smi
# Install container toolkit (if not already)
sudo apt install nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
Model caching
The docker-compose.yaml bind-mounts ~/.cache/wd14_tagger and ~/.cache/huggingface from the host so models persist across container rebuilds. Models auto-download on first use if not already cached.
Security
trust_remote_code and Florence-2
By default, the Florence-2 backend uses florence-community/Florence-2-base weights which are natively supported in transformers -- no trust_remote_code needed.
The legacy microsoft/Florence-2-base weights require HF_TRUST_REMOTE_CODE=true, which executes arbitrary Python from the model repository at load time. Only enable this for models you trust. WD14 uses a static ONNX model and never runs remote code.
Related projects
- argus-vision-demo -- a thin Next.js web UI for exploring argus-lens interactively.
- awesome-immich -- a curated list of Immich plugins, tools, and community projects. Argus Lens integrates with Immich as a companion tagging/captioning service -- see Immich.
License
MIT
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file argus_lens-0.3.0.tar.gz.
File metadata
- Download URL: argus_lens-0.3.0.tar.gz
- Upload date:
- Size: 177.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e9fc43b87858eec1be37ffa778acd30dc64b004dfcaf5579530625d18174ec27
|
|
| MD5 |
b0bd52e8a7b9b07f4bbb29d7195ca9c3
|
|
| BLAKE2b-256 |
0a0948bc8bd16b027b14c9e65522bc9e2f735e3c5d37d4ca09f7576729528c3b
|
Provenance
The following attestation bundles were made for argus_lens-0.3.0.tar.gz:
Publisher:
release.yml on smk762/argus-lens
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
argus_lens-0.3.0.tar.gz -
Subject digest:
e9fc43b87858eec1be37ffa778acd30dc64b004dfcaf5579530625d18174ec27 - Sigstore transparency entry: 2046732554
- Sigstore integration time:
-
Permalink:
smk762/argus-lens@8e19a299228526a4799f37e8b1479bb91e520ed3 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/smk762
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@8e19a299228526a4799f37e8b1479bb91e520ed3 -
Trigger Event:
push
-
Statement type:
File details
Details for the file argus_lens-0.3.0-py3-none-any.whl.
File metadata
- Download URL: argus_lens-0.3.0-py3-none-any.whl
- Upload date:
- Size: 78.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c16d07ab0ca6768e3246a66ac1718be44a57736168e9a47f6894f9571cae9468
|
|
| MD5 |
fbe595c1ade1a7b69e74cc7a9a153a69
|
|
| BLAKE2b-256 |
b1828bf1f302eb76e08e2ff761b6d98d3eb3486c934b578b89cc14b79d194935
|
Provenance
The following attestation bundles were made for argus_lens-0.3.0-py3-none-any.whl:
Publisher:
release.yml on smk762/argus-lens
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
argus_lens-0.3.0-py3-none-any.whl -
Subject digest:
c16d07ab0ca6768e3246a66ac1718be44a57736168e9a47f6894f9571cae9468 - Sigstore transparency entry: 2046732556
- Sigstore integration time:
-
Permalink:
smk762/argus-lens@8e19a299228526a4799f37e8b1479bb91e520ed3 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/smk762
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@8e19a299228526a4799f37e8b1479bb91e520ed3 -
Trigger Event:
push
-
Statement type: