Skip to main content

CaptionFlow

codecov PyPI version

scalable, fault-tolerant image captioning with vLLM or contributed API workers.

a fast websocket-based orchestrator paired with lightweight gpu workers achieves exceptional performance for batched requests through vLLM.

CaptionFlow is also integrated in bghira/SimpleTuner, where it powers an end-to-end caption-to-training workflow through the SimpleTuner WebUI. Use CaptionFlow directly when you want a standalone distributed captioning system, or use it through SimpleTuner when you want dataset captioning, caption review/export, and model training managed as one suite.

  • orchestrator: hands out work in chunked shards, collects captions, checkpoints progress, and keeps simple stats.
  • workers (vLLM or BYOK API): connect to the orchestrator, stream in image samples, and generate 1..N captions per image using prompts supplied by the orchestrator.
  • config-driven: all components read YAML config; flags can override.

no conda. just venv + pip.


install from pypi

python -m venv .venv
source .venv/bin/activate  # windows: .venv\Scripts\activate
pip install --upgrade pip
pip install "caption-flow[vllm]"

An OpenAI-compatible API worker does not need the GPU dependencies:

pip install "caption-flow[openai]"

For an orchestrator or monitor-only install, use pip install -e .. .[captioning] is an alias for .[vllm] for integrations such as SimpleTuner. Terminal image previews remain optional because the current term-image release requires an older Pillow major than CaptionFlow uses.

On Apple Silicon, install the MPS-compatible PyTorch chain and pinned Metal plugin with:

pip install -e ".[apple]"

CaptionFlow can use vLLM on Apple Silicon too, but the normal Linux vllm wheel is not the Apple install path. The pinned vllm-metal dependency is selected automatically on native arm64 Python 3.12+. For a ready-to-run Metal worker, use the upstream installer, which also builds/installs the Apple-specific vLLM core:

curl -fsSL https://raw.githubusercontent.com/vllm-project/vllm-metal/main/install.sh | bash
source ~/.venv-vllm-metal/bin/activate
pip install -e .

Do not combine .[apple] with the Linux .[vllm] extra.

For native macOS CPU vLLM instead, follow the official source-build instructions.

quickstart (single box)

for a full caption-to-training workflow with a web interface, use the SimpleTuner WebUI integration. the standalone flow below is best when you want to run CaptionFlow directly, contribute workers to a cluster, or export captions for your own downstream training pipeline.

  1. copy + edit the sample configs
cp examples/orchestrator/local_image_files.yaml my-orchestrator.yaml
cp examples/worker.yaml my-worker.yaml
cp examples/monitor.yaml my-monitor.yaml   # optional terminal interface

set a unique shared token in both my-orchestrator.yaml and my-worker.yaml (see auth.worker_tokens in the orchestrator config and worker.token in the worker config).

if you use private hugging face datasets/models, export HUGGINGFACE_HUB_TOKEN before starting anything.

  1. start the orchestrator
caption-flow orchestrator --config my-orchestrator.yaml
  1. start one or more vLLM workers
# gpu 0 on the same host
caption-flow worker --config my-worker.yaml --gpu-id 0

# your second GPU
caption-flow worker --config my-worker.yaml --gpu-id 1

# on a remote host
caption-flow worker --config my-worker.yaml --server ws://your.hostname.address:8765
  1. (optional) start the monitor
caption-flow monitor --config my-monitor.yaml
  1. export the data
% caption-flow export --help                                                                                                                                      
Usage: caption-flow export [OPTIONS]

  Export caption data to various formats.

Options:
  --format [jsonl|json|csv|txt|parquet|webshart|lance|huggingface_hub|all] Export format (default: jsonl)
  • jsonl: create JSON line file in the specified --output path
  • csv: exports CSV-compatible data columns to the --output path containing incomplete metadata
  • json: creates a .json file for each sample inside the --output subdirectory containing complete metadata; useful for webdatasets
  • txt: creates .txt file for each sample inside the --output subdirectory containing ONLY captions
  • webshart: updates an existing per-shard metadata .json file by writing captions under the plural captions key. for this format, pass --output as the path to the existing shard metadata JSON file when exporting one shard. if you export multiple shards, pass --output as a directory containing one existing {shard_name}.json file per shard.
  • huggingface_hub: creates a dataset on Hugging Face Hub, possibly --private and --nsfw where necessary
  • all: creates the directory/file-generating export formats in a specified --output directory. prefer a directory here; webshart is a special case that expects existing per-shard metadata .json files rather than creating new metadata files.

note: --output paths ending in .json are treated specially for webshart. use a directory for normal multi-format exports and an existing shard metadata JSON file only when intentionally updating a webshart shard.


how it’s wired

orchestrator

  • websocket server (default 0.0.0.0:8765) with three client roles: workers, data-feeders, and admin.
  • dataset control: the orchestrator centrally defines the dataset (huggingface or local) and version/name. it chunk-slices shards and assigns work.
  • data serving to remote workers: local files can be captioned by remote workers that don't have access to the same files, automatically.
  • vLLM config broadcast: model, tp size, dtype, max seq len, memory targets, batching, sampling params, and inference prompts are all pushed to workers; workers can apply many changes without a model reload.
  • storage + checkpoints: captions buffer to disk with periodic checkpoints. chunk state is tracked so restarts don’t double-work.
  • auth: token lists for worker, monitor, and admin roles.

vLLM worker

  • one process per gpu. select the device with --gpu-id (or worker.gpu_id in YAML).
  • gets its marching orders from the orchestrator: dataset info, model, prompts, batch size, and sampling.
  • resilient: detects disconnects, abandons the current chunk cleanly, clears queues, reconnects, and resumes.
  • batched generate(): images are resized down for consistent batching; each image can get multiple captions (one per prompt).

OpenAI-compatible BYOK worker

The API backend is a separate worker process, not an orchestrator integration. Provider keys remain in local environment variables; the CaptionFlow server sees only the normal worker token, submitted outputs, and non-secret capacity metrics.

cp examples/worker.openai-compatible.yaml my-api-worker.yaml
export ZAI_API_KEY="..."
caption-flow worker --config my-api-worker.yaml --openai-compatible

The worker sends OpenAI Chat Completions-compatible multimodal requests with an inline image data URL. Every configured endpoint can override the shared stage model with its own model, or use model_map when several orchestrator model names need explicit aliases.

Capacity is discovered independently for each endpoint:

  • it starts at initial_concurrency, adds a slot after sustained successful calls, and never exceeds max_concurrency;
  • a 429 or recognizable concurrency/capacity response halves the active limit and honors Retry-After or common rate-reset headers;
  • requests_per_minute adds conservative start-time pacing when a plan publishes an RPM limit;
  • retries can spill onto another endpoint, so multiple accounts or providers form one local pool without exposing their credentials.

The sample is ready for Z.AI's OpenAI-compatible Coding Plan endpoint and glm-5.3-flash. Change base_url and model to use any other compatible multimodal service; no orchestrator changes are needed.

The orchestrator may use the backend-neutral inference: key for shared stages, prompts, sampling, and output fields. Existing vllm: configurations are also accepted by API workers for backward compatibility; endpoint-local model values take precedence over the broadcast model name.


dataset formats

  • huggingface hub or local based URL list datasets that are compatible with the datasets library
  • huggingface hub datasets that are simple containers of raw image files
  • webdatasets shards containing full image data; also can be hosted on the hub
  • local folder filled with images; orchestrator will serve the data to workers

configuration path

config discovery order

for any component, the CLI looks for config in this order (first match wins):

  1. --config /path/to/file.yaml
  2. ./<component>.yaml (current directory)
  3. ~/.caption-flow/<component>.yaml
  4. $XDG_CONFIG_HOME/caption-flow/<component>.yaml
  5. /etc/caption-flow/<component>.yaml
  6. any $XDG_CONFIG_DIRS entries under caption-flow/
  7. ./examples/<component>.yaml (fallback)

tls / certificates

use the built-in helpers during development:

# self-signed certs for quick local testing
caption-flow generate_cert --self-signed --domain localhost --output-dir ./certs

# inspect any certificate file
caption-flow inspect_cert ./certs/fullchain.pem

then point the orchestrator at the resulting cert/key (or run --no-ssl for dev-only ws://).


tips & notes

  • multi-gpu: start one worker process per gpu (set --gpu-id or worker.gpu_id).
  • throughput: tune vllm.batch_size in the orchestrator config (or override with --batch-size at worker start). higher isn’t always better; watch VRAM.
  • prompts: add more strings under vllm.inference_prompts to get multiple captions per image; the worker returns only non-empty generations.
  • private HF: if your dataset/model needs auth, export HUGGINGFACE_HUB_TOKEN before caption-flow worker ....
  • self-signed ssl: pass --no-verify-ssl to workers/monitors in dev.
  • recovery: if you hard-crash mid-run, caption-flow scan_chunks --fix can reset abandoned chunks so the orchestrator can reissue them cleanly.

roadmap

  • hot config reload via the admin websocket path.
  • dedicated data-feeder clients (separate from gpu workers) that push samples into the orchestrator.
  • richer monitor TUI.

PRs welcome. keep it simple and fast.

architecture

┌─────────────┐     WebSocket      ┌─────────────┐
│   Worker    │◄──────────────────►│             │
│             │                    │             │     ┌──────────────┐
│             │◄───────────────────│             │────►│Arrow/Parquet │
└─────────────┘   HTTP (img data)  │ Orchestrator│     │   Storage    │
                                   │             │     └──────────────┘
┌─────────────┐                    │             │
│   Worker    │◄──────────────────►│             │
│             │                    │             │
│             │◄───────────────────│             │
└─────────────┘   HTTP (img data)  └─────────────┘
                                           ▲
┌─────────────┐                           │
│   Monitor   │◄──────────────────────────┘
└─────────────┘

Community Clusters

To contribute compute to a cluster:

  1. Install caption-flow: pip install "caption-flow[vllm]" for a GPU worker or pip install "caption-flow[openai]" for a BYOK API worker
  2. Get a worker token from the project maintainer
  3. Run: caption-flow worker --server wss://project.domain.com:8765 --token YOUR_TOKEN

Your contributions will be tracked and attributed in the final dataset!

License

AGPLv3

Metadata

Release files for caption-flow 0.5.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for caption-flow 0.5.5
File Size Uploaded
caption_flow-0.5.5.tar.gz 251.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for caption-flow 0.5.5
File Interpreter ABI Platform
caption_flow-0.5.5-py3-none-any.whl Python 3 none any Details

Total release size: 399.0 kB

Release files / caption_flow-0.5.5.tar.gz

Download URL caption_flow-0.5.5.tar.gz
Size 251.0 kB
Tags Source
SHA-256 checksum
How to use checksums
451c476b7258d71ede1d85271b405a7e8890fb9f9048c1d14646afd74492659a
BLAKE2b-256 checksum
How to use checksums
d3855a92348d599a338df44c8d275b749593b17001c0d520deb181dede9497a8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release files / caption_flow-0.5.5-py3-none-any.whl

Download URL caption_flow-0.5.5-py3-none-any.whl
Size 148.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0769d9f1896fb6fc1c465b828256e2574b5dd67c1e49a07ff8c67fed57fa96f8
BLAKE2b-256 checksum
How to use checksums
3f3bb1d5f6c7c4477f15a2cd4bc398c770270dc37849fef4f74ac97951a1ed1a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release history Release notifications | RSS feed

0.5.6

2 release files

This release

0.5.5 This release

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page