cuvis-ai-dataloader
Pluggable hyperspectral DataModules for the cuvis-ai ecosystem.
Overview
cuvis-ai-dataloader ships the concrete hyperspectral DataModules for the
Cuvis.AI ecosystem. A
DataModule is the unit the framework uses for both training and inference: it
bundles the data, the labels, the splits, and the train / val / test /
predict dataloaders.
This single plugin holds every concrete loader, with per-format heavy deps gated
behind optional extras. The cuvis SDK lives only here, behind [cu3s]; no
other Cuvis.AI repo pins it.
Module (data_module_name) |
Reads | Labels | Extra |
|---|---|---|---|
cu3s |
one .cu3s session (or a folder of them) via cuvis |
COCO JSON | [cu3s, coco] |
cu3s_multi |
many .cu3s, one frame per CSV row |
per-day COCO JSON | [cu3s, coco] |
npz_multi |
many .npz, one frame per CSV row |
baked mask array |
none |
tiff_paired |
a folder of *.tif / *.tiff cubes via tifffile |
paired PNG | [tiff] |
Key points:
- One plugin, three DataModules, per-format heavy deps behind extras
- The
cuvisSDK is isolated here, behind[cu3s] - Composable split selectors over an attributed sample universe
cuvis/tifffile/pycocotoolsimport lazily, so the plugin registers with any subset of extras
Splits
Splits are defined in one of two ways:
- Selectors (
splits.json) is the general mechanism, shared by every module. Composable selectors over an attributed sample universe are resolved into a committablesplits.jsonby theresolve-splitsCLI, then referenced from aDataConfig.splits. - One
universe.csvvocabulary (source, index+ optionalmaterialized_path, split, annotation, format, group) is read by bothcu3s_multiandnpz_multithrough a shared parser; each module keeps its own reader.cu3s_multimay carry an inlinesplitcolumn (present → module-owned; absent → needs asplits.json), andresolve-splits --from-csvturns that column into a committablesplits.json.npz_multiis selector-only (it rejects asplitcolumn) and requiresmaterialized_path(the.npz); forcu3s_multi,materialized_pathdefaults tosource(a raw.cu3sis its own file).sourceis the posix identity asplits.jsonselector keys on, so one split resolves against both the raw cu3s data and the converted npz.
Installation
uv pip install "cuvis-ai-dataloader[cu3s,coco]" # cu3s + COCO
uv pip install "cuvis-ai-dataloader[tiff]" # TIFF + paired PNG
uv pip install "cuvis-ai-dataloader[all]" # every format
Extras:
cu3s:.cu3ssession reading via thecuvisSDK bindingcoco: COCO-JSON mask labels (pycocotools,scikit-image)tiff: TIFF cube reading (tifffile)all: All formatsdev: Development dependencies
The cu3s extra carries the Windows cuvis-il<3.5.3 pin (the last build with a
win_amd64 wheel).
Cuvis SDK (system install, required for cu3s)
The [cu3s] extra installs the cuvis binding (with the Windows cuvis-il<3.5.3 pin noted
above), but that binding needs the system-wide C++ Cuvis SDK too, or any .cu3s read fails at
runtime. See the
Cuvis.AI installation guide for OS
support (Windows / Linux; not macOS), the SDK download, and verification. Quick check once
installed:
uv run python -c "import cuvis; print(cuvis.__version__)"
Usage
Inference (restore-pipeline) selects a module and its params on the CLI:
restore-pipeline \
--pipeline-path X.yaml \
--plugins-dir <this-repo>/configs/plugins \
--data-module cu3s \
--data-arg cu3s_file_path=X.cu3s \
--data-arg annotation_json_path=Y.json
Training (Train / RestoreTrainRun) selects the same module via the yaml
DataConfig:
data:
data_module: cu3s
splits:
train:
- { kind: file_indices, source: X.cu3s, ids: [0, 2, 3] }
val:
- { kind: file_indices, source: X.cu3s, ids: [1, 5] }
batch_size: 4
params:
cu3s_file_path: X.cu3s
annotation_json_path: Y.json
processing_mode: Reflectance
In-process / notebooks construct the DataModule directly and run it through
the Predictor:
from cuvis_ai_dataloader.data import Cu3sDataModule
from cuvis_ai_core.training import Predictor
from cuvis_ai_core.utils.restore import restore_pipeline
pipeline = restore_pipeline("X.yaml", plugins_dirs=[...])
dm = Cu3sDataModule(cu3s_file_path="X.cu3s", batch_size=1)
Predictor(pipeline, dm).predict()
NPZ (npz_multi)
npz_multi loads one frame per compressed .npz, selected by a splits.json over a
universe_csv (a universe.csv). It needs no extras (numpy is a core dep) and no Cuvis SDK. Each .npz carries:
cube:[H, W, C]float32wavelengths:[C](cast to int32)mask(optional):[H, W]int32 ground truth (zeros are emitted when absent)class_mask(optional):[H, W]uint8 per-pixel COCO category id (0 = background)
The universe_csv requires source, index plus materialized_path (the .npz, required for npz;
optional annotation, format, group; extra columns are ignored); materialized_path is relative
to the CSV and must not escape it via ... A split column is rejected here (npz is
selector-only). Each sample is
{cube, mask, class_mask, wavelengths, mesu_index, frame_id}. Unlike the cu3s modules, npz_multi
honors pin_memory / persistent_workers / worker_multiprocessing_context (pure-CPU numpy loads
benefit from them).
from cuvis_ai_dataloader.data import MultiNpzDataModule
from cuvis_ai_schemas.training.data import DataSplitConfig, Selector, SelectorKind
splits = DataSplitConfig(
train=[Selector(kind=SelectorKind.FILE_INDICES, source="X.cu3s", ids=[0, 2, 3])],
val=[Selector(kind=SelectorKind.FILE_INDICES, source="X.cu3s", ids=[1, 5])],
)
dm = MultiNpzDataModule(splits=splits, universe_csv="universe.csv", batch_size=4, num_workers=4)
dm.setup("fit")
batch = next(iter(dm.train_dataloader())) # cube [B,H,W,C], mask [B,H,W], ...
In a DataConfig (training / restore-trainrun):
data:
data_module: npz_multi
batch_size: 4
splits:
train:
- { kind: file_indices, source: X.cu3s, ids: [0, 2, 3] }
val:
- { kind: file_indices, source: X.cu3s, ids: [1, 5] }
params:
universe_csv: universe.csv
GUI-authored splits over a cu3s folder (contract)
External split authors (e.g. the CuvisNEXT split designer) write a frozen splits.json
(a serialized DataSplitConfig with file_indices selectors) against a folder of cu3s
files with per-measurement granularity. That contract is cu3s folder mode with
frames: measurements:
data:
data_module: cu3s
batch_size: 1
num_workers: 0
splits:
splits_path: <absolute path to the frozen splits.json>
params:
files: # the recordings to use, when the author knows them
- <absolute path to a .cu3s>
data_dir: <folder holding the .cu3s files>
frames: measurements
recursive: true # walk per-day subfolders (fallback only)
processing_mode: Reflectance
The frozen rules both sides implement:
- Universe = the recordings the run actually uses, one sample per measurement
0..N-1, ordered by(source, index). Which recordings those are is answered in this order: thefileslist when given (nothing is walked, and the list may point outsidedata_dir, e.g. a split spanning two drives); otherwise the sources the split's selectors name, when every selector names its sources (files/file_indices, or set operations over those); otherwise every*.cu3sunderdata_dir, recursive whenrecursive: true. A positional or attribute-driven selector (dir_indices,stems,glob,tag,categories,all) can only be answered by the full universe, so it keeps the walk. Enumeration opens each recording once for its measurement count only (no processing context), so a folder holding recordings the split does not name costs nothing. - Source identity is canonical: the absolute path with forward slashes and
filesystem-true case — Python
Path(p).resolve().as_posix(), C++/QtQFileInfo::canonicalFilePath(). Selectors in the authoredsplits.jsonmust carry exactly this form; a moved or renamed member file fails loud, naming the recording it could not read, rather than silently shrinking a split. Sources reached only through a set operation may be absent (except(files[a], files[gone])is legitimate, and core resolves a set operation's operands without its zero-match check). - An empty
predictstage serves the whole universe, which with an explicitfileslist or a fully source-naming split means the recordings that split uses, not the whole folder. uid=<source>#<index>(the sibling COCO image id equals the read position, so it never extends the uid).universe_hash= sha256 over the ordered uids, each followed by\n(cuvis_ai_core.data.splits_io.universe_hash). Forfile_indicessplits the server treats the hash as informational (only positionaldir_indicessplits are hash-verified); staleness detection is the author's concern.- Annotations are the sibling
<stem>.jsonCOCO next to each cu3s (attached automatically); an emptypredictstage means the whole universe. - Training stages require splits.
cu3sdoes not own split semantics:fit/validate/testwith noDataConfig.splitsraise instead of silently iterating the whole universe (which would contaminate statistical initialization with anomalous frames). Split-lesspredictover the whole universe stays valid.
The golden fixture tests/cuvis_ai_dataloader/fixtures/gui_authored_splits.json is the
byte-level reference of the authored shape (the {DATA_DIR} token stands in for the
machine-specific folder); the same file is committed in the CuvisNEXT test suite and its
universe_hash doubles as the shared sha256 test vector. Changing it is a cross-repo
contract change.
Architecture
Concrete DataModules subclass cuvis_ai_core.data.datamodule.BaseCuvisAIDataModule
and implement validate_params plus the selector hooks enumerate(required_attrs)
(the module's attributed sample universe) and build_dataset_from_refs(refs);
a module that owns its own splits also implements build_stage_dataset(stage).
Per-format cube readers and labelers are internal helpers (data/readers/,
data/labelers/), reused but not a plugin contract. Module-top imports stay free
of heavy deps; cuvis / tifffile / pycocotools / scikit-image load lazily
on first use (data/_extras.py).
Development
uv sync --extra dev
uv run pytest tests/ -v
uv run ruff check cuvis_ai_dataloader/ tests/
uv run ruff format cuvis_ai_dataloader/ tests/
uv run mypy cuvis_ai_dataloader/
Git hooks
Enable the repo's hooks once per clone:
git config core.hooksPath .githooks
- pre-commit:
ruff format+ruff check --fixon staged Python, then re-stages. - pre-push:
ruff format --check,ruff check, docstring coverage (uvx interrogate, ≥95%, configured in[tool.interrogate]), andpytest -m "not slow and not gpu".
Skip a hook for one command with --no-verify.
Contributing
Contributions are welcome. Please:
- Ensure tests pass
- Run ruff format and ruff check
- Keep type hints and update docs as needed
License
Licensed under the Apache License 2.0. See LICENSE for details.
Metadata
Release files for cuvis-ai-dataloader 0.6.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| cuvis_ai_dataloader-0.6.4.tar.gz | 205.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| cuvis_ai_dataloader-0.6.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 271.3 kB
Release files / cuvis_ai_dataloader-0.6.4.tar.gz
| Download URL | cuvis_ai_dataloader-0.6.4.tar.gz |
|---|---|
| Size | 205.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
87243fd70ccf7d2cbe46cd32ed08d991543c8cd3a80798c3a7f08851a4e46b8d
|
|
BLAKE2b-256 checksum How to use checksums |
7359437e452316ec66b49b1a3a6640e02e1963c13b01f857bf1918468024964f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.
Transparency logRelease files / cuvis_ai_dataloader-0.6.4-py3-none-any.whl
| Download URL | cuvis_ai_dataloader-0.6.4-py3-none-any.whl |
|---|---|
| Size | 65.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
93345c37766a9b52c0e2e9e9483fff383c85b10ee55cc053fe4cfee55b01f7e1
|
|
BLAKE2b-256 checksum How to use checksums |
4d7fc41c5d35b5f8837f8b3feb5ebe076a38fc0e6334d2c01b45783cc74d387f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.
Transparency log