Skip to main content

justdata

A TensorFlow-native data pipeline library with a modality-neutral core and first-class computer vision and acoustic recipes. justdata.core owns loading, adapter, preset, metadata, padding, and seeded execution machinery; justdata.vision owns image schemas, transforms, augmentations, corruptions, tasks, and vision presets. justdata.acoustic owns audio schemas, decoding, resampling, segmentation, frontends, augmentations, corruptions, DCASE helpers, stats, metadata/JAX helpers, and acoustic presets.


Table of Contents

  1. Installation
  2. Architecture
  3. Data Pipeline Stages
  4. Supervised Learning — Training Pipelines
  5. Automatic Augmentation Policies: RandAugment, TrivialAugment, and TrivialAugmentWide
  6. Supervised Learning — Validation Pipelines
  7. The timm ImageNet Recipes (ResNet Strikes Back: A1/A2/A3)
  8. Self-Supervised Learning Pipeline (DINOv2)
  9. Registry System
  10. Presets and Smart Merging
  11. Acoustic Pipelines
  12. Dataset Adapters
  13. Mini-C Corruption Benchmark
  14. Audio Corruption Benchmark
  15. Usage
  16. Development

Installation

pip install justdata

Requirements: Python 3.11–3.13, TensorFlow ≥ 2.18.1, TensorFlow Datasets ≥ 4.9.9. Install optional modality extras with justdata[vision] for Hugging Face vision datasets, justdata[wilds] for WILDS image classification datasets, or justdata[acoustic] for audio dataset/source dependencies.


Architecture

justdata is structured around a fixed, four-stage pipeline abstraction. The currently supported vision tasks, classification, semantic segmentation, and panoptic segmentation, each expose exactly four composable functions. Object detection and depth estimation are future scope and are not registered capabilities.

(preprocess_fn, augment_fn, late_augment_fn, postprocess_fn)

These are assembled by justdata.core.load_ds into the following execution graph:

fetch_ds -> adapter -> preprocess -> cache -> augment -> shuffle -> postprocess -> batch -> late_augment -> pad -> prefetch

The strict ordering reflects the execution domain requirements articulated throughout this document: format normalization occurs before the optional preprocessing cache; spatial and photometric distortions precede tensor conversion and normalization; and batch-level operations (Mixup, CutMix, random erasing) occur after batching on the GPU. An optional model-input cache can be inserted after deterministic postprocessing and before batching; for training it must be explicitly opted in and is placed before shuffle so epoch order is not frozen.

TensorFlow execution controls

load_ds uses tf.data.AUTOTUNE for loader-owned maps by default. Advanced callers may set map_parallel_calls, private_threadpool_size, and max_intra_op_parallelism to bound TensorFlow input-pipeline concurrency. Omitting these arguments preserves the default TensorFlow behavior. The map limit also applies to metadata sidecar writing when enabled.

Core Public API

Function Description
justdata.core.fetch_ds(dataset_names, splits_info, data_dir) Raw dataset loading through registered source loaders. Returns a tf.data.Dataset in canonical schema.
justdata.core.load_ds(...) Full pipeline for training or evaluation. Handles caching, augmentation, shuffling, batching, and prefetching.
justdata.core.admit_inventory(...) Strict offline validation of a resolved inventory into a disk snapshot with structured identity and failure reports.
justdata.core.open_inventory(snapshot_dir) Verify and reopen a completed inventory snapshot without original sources.
justdata.core.load_inventory(admitted, ...) Apply the shared pipeline to an admitted inventory, including raw epoch finalization.
justdata.core.ExecutedConfig Immutable, canonical snapshot returned by loaders when return_config=True.
justdata.vision.minic.create_minic_datasets(...) Constructs Mini-C corruption benchmark datasets from a shared preprocessed base dataset.
justdata.acoustic.corruptions.create_audio_corruption_datasets(...) Constructs acoustic corruption benchmark datasets from a shared preprocessed base dataset.
justdata.acoustic.dcase2025.make_source_dataset(...) Builds DCASE Task 1 source-domain datasets with split-safety checks.

Import justdata.vision before resolving built-in vision datasets or pipelines. Hugging Face vision datasets are referenced with the hf: prefix (e.g., hf:cifar10), and WILDS image classification datasets are referenced with the wilds: prefix (e.g., wilds:camelyon17).

Use the strict local inventory API when every requested source and record must be accounted for. It verifies expected content digests, preserves record IDs and order, and reports declared filtering. Existing fetch_ds and load_ds retain their legacy source-skipping behavior and do not provide the strict admission guarantee.


Data Pipeline Stages

Stage 1: Preprocessing (before cache)

Format normalization is performed before the optional decoded-dataset cache. Operations include rank fixing (2D -> 3D), CHW -> HWC transposition, grayscale -> RGB expansion, and RGBA -> RGB projection. The output shape is forced to [None, None, 3].

Stage 2: Augmentation (per-sample, training only)

Per-sample spatial and photometric augmentations are applied after the optional cache stage, ensuring that each training epoch receives independently sampled augmentations. All augmentation functions use tf.random.split for stateless randomization, enabling full reproducibility.

Stage 3: Postprocessing (every sample)

Resize, channel-wise normalization to zero mean and unit standard deviation, and optional HWC -> CHW transposition. This stage runs unconditionally for both training and evaluation.

Set cache_model_inputs=True with model_input_cache_path to cache these resized, normalized tensors on local SSD. For dataset_type="train", also set allow_train_model_input_cache=True and use it only when the training view is deterministic; stochastic crops or photometric augmentation will be materialized into the cache on first fill.

Stage 4: Late Augmentation (per-batch, training only)

Batch-level operations—Mixup, CutMix, and random erasing—are applied after batching. These operations are strictly training-only and require a formed batch to operate across the sample dimension.


Supervised Learning — Training Pipelines

To ensure mathematical correctness, the training pipeline is strictly segregated by execution domain. Spatial and photometric distortions are applied to [0, 255] image tensors (PIL-equivalent domain) before zero-mean tensor normalization.

1. Image Domain Augmentations

Geometric Augmentations

  • CIFAR-10/100: RandomCrop(32, padding=4, padding_mode='zeros') -> RandomHorizontalFlip(p=0.5).
  • ImageNet-1K: RandomResizedCrop(size=224) (or 256). Scale: (0.08, 1.0), Aspect ratio: (0.75, 1.33), Interpolation: Bicubic. -> RandomHorizontalFlip(p=0.5).

Photometric Augmentations (Mutually Exclusive Branches)

The photometric augmentation strategy is architecture-dependent and the two branches are mutually exclusive.

  • Modern Branch (ViT / ConvNeXt): Apply RandAugment(num_ops=2, magnitude=9), TrivialAugment, or TrivialAugmentWide (see the Automatic Augmentation Policies section for full specifications). Color Jitter is explicitly disabled to prevent redundant and destructive color space distortion.
  • Legacy Branch (ResNet): Apply ColorJitter(brightness=0.4, contrast=0.4, saturation=0.4, hue=0.1). RandAugment is disabled.

2. Tensor Domain

  • Conversion & Scaling: ToImage() -> ToDtype(float32, scale=True). Maps [0, 255] -> [0.0, 1.0].
  • Normalization:
    • CIFAR-10: mean=(0.4914, 0.4822, 0.4465), std=(0.2023, 0.1994, 0.2010)
    • CIFAR-100: mean=(0.5071, 0.4867, 0.4408), std=(0.2675, 0.2565, 0.2761)
    • ImageNet-1K: mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225)
  • Random Erasing (Cutout): Applied to the normalized tensor. p=0.25, scale=(0.02, 0.33), ratio=(0.3, 3.3).

3. Mini-Batch Domain

Applied across the batch dimension during training.

  • Repeated Augmentation (RA): Enabled for ViTs. Typically 3 repetitions per sample per mini-batch.
  • Mixup & CutMix: Controlled by mixup_prob=1.0 (probability of batch mixing) and switch_prob=0.5 (probability of selecting CutMix over Mixup).
    • Mixup: $\tilde{x} = \lambda x_i + (1 - \lambda) x_j$, where $\lambda \sim \mathrm{Beta}(0.8,, 0.8)$.
    • CutMix: Replaces a rectangular bounding box region; $\lambda \sim \mathrm{Beta}(1.0,, 1.0)$.
    • Note: The $\alpha$ values (0.8 and 1.0) reflect the DeiT baseline. Mixup/CutMix parameters are strictly recipe-dependent; see the A1/A2/A3 section for ResNet-specific variations.
  • Label Smoothing: Cross-entropy loss modification with $\varepsilon = 0.1$.

Automatic Augmentation Policies: RandAugment, TrivialAugment, and TrivialAugmentWide

justdata provides native TensorFlow implementations of three closely related automatic augmentation policies: RandAugment (Cubuk et al., 2020), TrivialAugment (Müller & Hutter, 2021), and TrivialAugmentWide (Müller & Hutter, 2021). All three are registered under the augment_strategy registry and share a common operation pool and magnitude discretization framework.

1. Pipeline Position and Domain Constraints

To guarantee deterministic reproduction of published results, all three policies are applied exclusively in the image domain (integer uint8 tensors, pixel values $\in [0, 255]$), before tensor conversion to floating-point and channel-wise normalization. The canonical execution order within the per-sample augmentation stage is:

  1. Random Resized Crop (or Random Pad Crop for CIFAR)
  2. RandAugment / TrivialAugment / TrivialAugmentWide
  3. Random Horizontal Flip (integrated into the crop strategy)
  4. ToTensor — scales to float $[0.0, 1.0]$ (postprocessing stage)
  5. Normalize — subtracts dataset mean, divides by standard deviation (postprocessing stage)

Applying photometric and geometric distortions prior to floating-point conversion ensures that operations such as Posterize and Solarize, which are defined on integer pixel arithmetic, remain numerically well-founded, and that fill values for geometric operations are expressed in the same integer domain as the source image.

2. The RA Operation Space ($K = 14$)

All three algorithms draw from the RA augmentation space, a fixed pool of $K = 14$ operations. The pool is partitioned into two subsets based on magnitude dependency.

Magnitude-Independent Operations (3 ops)

These operations ignore the sampled magnitude $m$ entirely.

Operation Description
Identity Returns the image unmodified.
AutoContrast Linearly scales the pixel intensity histogram so that the darkest pixel maps to 0 and the brightest to 255.
Equalize Equalizes the image histogram per channel using a cumulative distribution function.

Magnitude-Dependent Operations (11 ops)

These operations scale their physical intensity as a function of the sampled magnitude index $m$.

Sign randomization. For all geometric operations (Rotate, ShearX, ShearY, TranslateX, TranslateY) and all color enhancement operations (Brightness, Color, Contrast, Sharpness), the direction of the applied transformation is randomized bidirectionally. Given a raw physical magnitude $v$ computed from $m$, the applied value is:

$$v' = v \cdot s, \quad s \sim \mathcal{U}{-1, +1}$$

This sign randomization is sampled independently per operation per sample, using a stateless seed derived from the layer's random state.

Geometric fill. When pixels are shifted outside the image boundary by Rotate, ShearX, ShearY, TranslateX, or TranslateY, the vacated regions are filled with a constant value (default: 128). Geometric operations use the implementation's fixed nearest-neighbor interpolation for images.

For segmentation pipelines, these geometric operations reuse the same sampled parameters for the image and segmentation map. The image uses bilinear interpolation, while the map uses nearest-neighbor interpolation and preserves its integer dtype. Vacated map pixels use the configurable mask_fill_value (default 255), which must be treated as an ignore label by the training loss. Photometric operations continue to affect only the image.

Solarize semantics. Solarize inverts all pixel values that are greater than or equal to the computed threshold $\tau$. Formally, for each pixel $p$:

$$p' = \begin{cases} 255 - p & \text{if } p \geq \tau \ p & \text{otherwise} \end{cases}$$

where $\tau = 255 \cdot (1 - m / B)$. At $m = 0$, the threshold is 255 (no pixels inverted); at $m = B$, the threshold is 0 (all pixels inverted).

3. Magnitude Discretization and Physical Mappings

The magnitude scale is a 31-bin discrete framework with bin indices $m \in {0, 1, \ldots, 30}$ and $B = 30$ (the maximum bin index). This parameterization is the native scale of the torchvision (v0.13+) implementation of both RandAugment and TrivialAugmentWide.

The table below specifies the physical mapping formula and the operation-specific bound for each of the two spaces. $W$ and $H$ denote the input image width and height at the moment the augmentation is applied.

Operation Sign Rand. Physical Mapping Standard Bound (RA / TA) Wide Bound (TA-Wide)
Rotate Yes $(m / B) \cdot \text{MaxDeg} \cdot s$ 30.0° 135.0°
TranslateX Yes $(m / B) \cdot \text{MaxPx}_X \cdot s$ $(150 / 331) \times W$ 32.0 px
TranslateY Yes $(m / B) \cdot \text{MaxPx}_Y \cdot s$ $(150 / 331) \times H$ 32.0 px
ShearX Yes $(m / B) \cdot \text{MaxShear} \cdot s$ 0.3 0.99
ShearY Yes $(m / B) \cdot \text{MaxShear} \cdot s$ 0.3 0.99
Brightness Yes $1.0 + (m / B) \cdot \text{MaxDelta} \cdot s$ 0.9 0.99
Color Yes $1.0 + (m / B) \cdot \text{MaxDelta} \cdot s$ 0.9 0.99
Contrast Yes $1.0 + (m / B) \cdot \text{MaxDelta} \cdot s$ 0.9 0.99
Sharpness Yes $1.0 + (m / B) \cdot \text{MaxDelta} \cdot s$ 0.9 0.99
Posterize No $8 - \mathrm{round}!\left((m / B) \cdot \text{MaxBits}\right)$ MaxBits = 4 (min 4 bits retained) MaxBits = 6 (min 2 bits retained)
Solarize No $255.0 \cdot (1 - m / B)$ threshold down to 0 threshold down to 0
Identity — — — —
AutoContrast — — — —
Equalize — — — —

Note on TranslateX / TranslateY: Translation is a deliberate exception to the "wider bounds" pattern. The TA-Wide space specifies a fixed ceiling of 32 px for both axes. For any image dimension exceeding approximately 71 px, the Standard RA space—whose ceiling scales proportionally with the image dimension—admits a larger maximum translation than the Wide space.

4. Algorithm Specifications

4.1 RandAugment (RA)

RandAugment (Cubuk et al., 2020) reduces the search space of AutoAugment from $\mathcal{O}(K^N)$ policies to two scalar hyperparameters: the number of sequential operations $N$ and the global magnitude $M$.

Procedure. Given a training image $x$:

  1. Independently sample $N$ operation indices $k_1, \ldots, k_N$ uniformly at random with replacement from the pool of $K$ operations.
  2. For each selected operation $T_{k_i}$, map the fixed global magnitude $M$ to a physical parameter via the operation-specific formula in the table above, applying sign randomization where applicable.
  3. Apply the operations sequentially: $x \leftarrow T_{k_N}(\cdots T_{k_1}(x))$.

The magnitude $M$ is expressed on the 31-bin scale ($M \in {0, \ldots, 30}$) used throughout justdata. An optional magnitude perturbation $\sigma > 0$ samples the effective level per layer from $\mathcal{N}(M, \sigma^2)$, clipped to $[0, B]$.

API parameters (augment_strategy = "rand_augment" or direct call to rand_augment):

Parameter Type Default Description
num_layers int 2 Number of operations $N$ applied per image.
magnitude float 9.0 Global magnitude $M$ on the 31-bin scale.
magnitude_std float 0.0 Per-layer Gaussian magnitude noise $\sigma$. Set to 0.5 to replicate timm stochastic magnitude.
prob_to_apply float | None None If set, each layer is skipped with probability $1 - p$.
rotate_max float 30.0 Maximum rotation angle in degrees.
shear_max float 0.3 Maximum shear coefficient (Standard RA space).
enhance_max float 0.9 Maximum delta for brightness, color, contrast, sharpness.
posterize_max_bits int 4 Bits removed at maximum magnitude; min retained = $8 - \text{MaxBits}$.
translate_const float 100.0 Absolute translate ceiling in pixels. For Standard RA, set to $(150/331) \times \text{image_width}$.
exclude_ops list[str] | None None Operations to exclude from the sampling pool.

Published optimal hyperparameters and 31-bin scale conversion. The original paper (Cubuk et al., 2020) reports hyperparameters on an 11-level scale ($M_{\text{paper}} \in {0, \ldots, 10}$). To replicate those results using the 31-bin scale employed by justdata, apply the conversion $m = M_{\text{paper}} \times 3$.

Dataset Architecture $N$ $M_{\text{paper}}$ $m$ (31-bin)
CIFAR-10 WideResNet-28-2 3 4 12
CIFAR-10 WideResNet-28-10 3 5 15
CIFAR-10 PyramidNet + ShakeDrop 3 7 21
CIFAR-10 Shake-Shake 3 9 27
ImageNet-1K ResNet-50 2 9 27
ImageNet-1K EfficientNet-B7 2 — 28†

†The EfficientNet-B7 value was reported on an extended scale beyond 10; cross-verify against the target codebase before use. The torchvision default of magnitude=9 on the 31-bin scale corresponds to moderate augmentation and is not equivalent to the paper's $M_{\text{paper}} = 9$.

4.2 TrivialAugment (TA)

TrivialAugment (Müller & Hutter, 2021) eliminates hyperparameter search entirely by selecting a single operation and sampling its magnitude uniformly at random from the full discrete range on each forward pass.

Procedure. Given a training image $x$:

  1. Sample one operation index $k \sim \mathcal{U}{1, \ldots, K}$ uniformly from the 14-op RA pool.
  2. Sample a magnitude $m \sim \mathcal{U}{0, 1, \ldots, 30}$ uniformly from the full 31-bin range.
  3. Map $m$ to a physical parameter via the Standard bounds in the table above, applying sign randomization where applicable.
  4. Apply the single operation: $x \leftarrow T_k(x)$.

The proportional translate ceiling is computed dynamically as $(150 / 331) \times W$ (image width), matching the Standard RA space definition.

API: augment_strategy = "trivial_augment". Accepts translate_const (float, optional; computed from image width if not supplied) and exclude_ops (list of strings, optional).

4.3 TrivialAugmentWide (TA-Wide)

TrivialAugmentWide (Müller & Hutter, 2021) is the native torchvision variant of TrivialAugment. It uses the same zero-hyperparameter, single-operation protocol as baseline TA, but operates over the Wide magnitude bounds, substantially expanding the geometric and photometric search range.

Procedure. Identical to TrivialAugment, with two differences:

  1. The Wide physical bounds from the table above are applied in place of the Standard bounds.
  2. The translate ceiling is a fixed 32 px (not image-proportional).

Wide bounds summary:

Transformation Standard RA / TA TA-Wide
Rotation range ±30° ±135°
Shear range (X and Y) ±0.3 ±0.99
Enhancement delta (Brightness, Color, Contrast, Sharpness) ±0.9 ±0.99
Posterize (minimum bits retained) 4 2
Translation (fixed ceiling) $(150/331) \times \text{dim}$ 32 px

API: augment_strategy = "trivial_augment_wide". Accepts exclude_ops (list of strings, optional). The wide bounds are fixed by design and cannot be overridden via kwargs; to use custom bounds, call rand_augment directly with num_layers=1 and the desired parameters.

5. Algorithm Comparison

Property RandAugment TrivialAugment TrivialAugmentWide
Operations applied per image $N$ (tunable, with replacement) 1 (fixed) 1 (fixed)
Magnitude $m$ Fixed global $M$, constant across ops Sampled: $m \sim \mathcal{U}{0, \ldots, 30}$ Sampled: $m \sim \mathcal{U}{0, \ldots, 30}$
Hyperparameter search required Grid search over $(N, M)$ None None
Transformation bounds Standard RA space Standard RA space Wide space
Translate ceiling Configurable (default 100 px) $(150/331) \times W$ (image-proportional) 32 px (fixed)
Reference implementation torchvision automl/trivialaugment torchvision
justdata registry key rand_augment trivial_augment trivial_augment_wide

6. Non-RA Operations

Invert, Cutout, SolarizeAdd, and Grayscale are outside the fixed 14-op pool and cannot be selected through RandAugment, TrivialAugment, or TrivialAugmentWide. The public exclude_ops parameter only removes operations from the pool; listing these names has no effect. The cutout_const argument remains accepted for compatibility and has no effect. Random erasing is available through the random_erasing late-augmentation stage.


Supervised Learning — Validation Pipelines

Validation pipelines are strictly deterministic. The objective shifts from regularization to feature preservation and scale alignment.

CIFAR-10 / CIFAR-100 Standard Validation

Because CIFAR images are inherently 32×32 and contain minimal background, spatial cropping destroys the primary subject.

Pipeline: ToImage() -> ToDtype(float32, scale=True) -> Normalize (training-set statistics as above).

ImageNet-1K Standard Baseline (The 0.875 Rule)

For standard supervised models and baseline linear probing, the canonical 0.875 crop ratio discards peripheral background.

Pipeline: Resize(256, interpolation=Bicubic) -> CenterCrop(224) -> ToImage() -> ToDtype(float32, scale=True) -> Normalize.

ImageNet-1K Modern Recipes (FixRes Strategy)

Modern recipes correct train-test resolution discrepancies by manipulating the validation crop percentage (crop_pct).

  • A3 (Light) Validation: Train at 160×160, validate at 224×224. Resize shorter edge to ≈236, then CenterCrop(224).
  • A1 / A2 (Heavy/Moderate) Validation: Default test at 224×224 (crop_pct=1.0, so Resize(224) -> CenterCrop(224)). Accuracy improves further via FixRes evaluation at 288×288 (Resize(288) -> CenterCrop(288)) as demonstrated in the RSB paper.

The timm ImageNet Recipes (ResNet Strikes Back: A1/A2/A3)

The "ResNet Strikes Back" (RSB) recipes dynamically scale augmentation intensity and training schedules to match model capacity.

Crucial context: These recipes were designed specifically for ResNet-family architectures. ViTs typically use distinct recipes (e.g., DeiT, BEiT) with different Mixup $\alpha$, optimizers (AdamW), and loss functions. While the principle of scaling augmentation with model capacity generalizes, the specific hyperparameters below do not transfer directly to ViTs.

Parameter A1 (Heavy) A2 (Moderate) A3 (Light)
Target Architecture Large ResNets (e.g., ResNet-152/200) or high compute ResNet-50 (standard) ResNet-50 (fast) or smaller (e.g., ResNet-18)
Training Resolution 224 224 160 (FixRes strategy)
Test Resolution 224 (scales to 288 via FixRes) 224 224
Epochs 600 300 100
Optimizer LAMB LAMB LAMB
LR Schedule Cosine with warmup Cosine with warmup Cosine with warmup
Loss Function BCE (per-class binary) BCE CE (standard)
RandAugment $m \approx 7$, $n = 2$ $m \approx 6$, $n = 2$ $m \approx 6$, $n = 2$
Random Erasing p=0.35 p=0.25 Disabled (p=0.0)
Repeated Aug (RA) Enabled (3×) Enabled (3×) Disabled
Mixup $\alpha$ 0.2 0.2 0.1
CutMix $\alpha$ 1.0 1.0 1.0
Stochastic Depth Capacity-dependent (e.g., 0.05+) 0.0 0.0
EMA Yes Yes No (or lighter)
Weight Decay 0.02 0.02 0.02

Note: RandAugment parameter $n$ defaults to 2 per the original specification, though exact magnitude strings fluctuate across timm versions.


Self-Supervised Learning Pipeline (DINOv2)

DINOv2 employs a Teacher-Student knowledge distillation framework operating on multi-crop asymmetry to force the learning of semantic invariance over low-level frequency matching.

Multi-Crop Geometric Strategy

  • Global Crops (Context): 2 crops at 224×224. Scale: (0.32, 1.0). Passed to both Teacher and Student.
  • Local Crops (Detail): 8 crops at 96×96 (default). Use 98×98 ($14 \times 7$) to avoid positional embedding interpolation when using ViT-14 backbones. Scale: (0.05, 0.32). Passed to Student only.

Asymmetric Pipeline Implementation

The per-crop asymmetry is structurally enforced by three distinct Compose pipelines rather than conditional branching within a single block. Each source image passes through all three pipelines to produce the full 10-crop suite.

  • global_transform_1: Global Crop 1. Enforces strict blurring (p=1.0), disables solarization (p=0.0).
  • global_transform_2: Global Crop 2. Minimizes blurring (p=0.1), enables solarization (p=0.2).
  • local_transform: The 8 Local Crops. Moderate blurring (p=0.5), disables solarization (p=0.0).

Execution Order and Domain Separation (Per Crop)

Step Operation Parameters / Per-Crop Asymmetry Domain
1 RandomResizedCrop Aspect ratio (0.75, 1.33), Bicubic interpolation. Image
2 RandomHorizontalFlip p=0.5 (all crops) Image
3 ColorJitter p=0.8 (all crops). b=0.4, c=0.4, s=0.2, h=0.1 Image
4 RandomGrayscale p=0.2 (all crops) Image
5 GaussianBlur $\sigma \sim \mathrm{Uniform}(0.1, 2.0)$, dynamic kernel. Global 1: p=1.0 · Global 2: p=0.1 · Local: p=0.5 Image
6 Solarization Invert pixels $> 128$. Global 1: p=0.0 · Global 2: p=0.2 · Local: p=0.0 Image
7 Convert & Scale ToImage() -> ToDtype(float32, scale=True) Image -> Tensor
8 Normalize ImageNet mean and std. Tensor (zero-mean)

SSL Downstream Evaluation

Image-Level Classification (Linear Probing / $k$-NN): Standard supervised inference transforms: Resize(256, Bicubic) -> CenterCrop(224) -> ToImage() -> ToDtype(float32) -> Normalize. Features are extracted from the [CLS] token or a concatenation of [CLS] and average-pooled patch tokens.

Dense Tasks (Segmentation / Depth Validation — Patch Alignment): To avoid dropping boundary pixels or forcing complex interpolation during dense evaluation, images are resized to the target scale and padded (reflection or zero) on the bottom and right edges such that both height and width are exact multiples of the ViT patch size (e.g., 14). This is implemented in justdata.vision.transforms.pad_to_patch_multiple.

For semantic segmentation with replayable geometry, select vision/segmentation with apply_presets=False and explicit geometry_kwargs. It supports paired aspect-preserving training resize and crops, longer-side evaluation caps, independent RGB/ignore padding, per-pixel validity, and original-coordinate score restoration. See the dense geometry contract for configuration, record fields, replay, and batching requirements.

For instance-aware labels, select vision/panoptic_segmentation with a segment-ID map and category table. It preserves thing instances, merges stuff by category, and excludes crowd pixels from panoptic supervision. The panoptic contract and LaRS example show how to load either semantic or panoptic labels from local LaRS v1.0.0 archives and plot five images under the A1/A2/A3 presets with both label views.


Registry System

All extensible components in justdata use a decorator-based registry pattern with thread-safe lookups.

Registry Decorator Lookup
Crop strategies @register_crop_strategy(name) justdata.vision.augmentations.get_crop_strategy(name)
Augment strategies @register_augment_strategy(name) justdata.vision.augmentations.get_augment_strategy(name)
Corruptions @register_corruption(name, descriptor=...) justdata.vision.corruptions.apply_corruption(...)
Dataset adapters @register_adapter(dataset_name) justdata.core.get_adapter(dataset_name)
Source loaders @register_source_loader(prefix) justdata.core.get_source_loader(dataset_name)
Pipelines @register_pipeline("modality/task") justdata.core.get_pipeline(...)
Dataset metadata register_dataset(name, task_type, modality=...) justdata.core.get_dataset_info(name)
Audio frontends @register_audio_frontend(name) justdata.acoustic.get_audio_frontend(name)
Audio components @register_audio_* decorators justdata.acoustic.get_audio_*, list_audio_*, has_audio_*

get_pipeline is the high-level resolver: it infers the task type from the dataset name, merges preset defaults with user-supplied kwargs (via smart merge; see below), and invokes the appropriate pipeline factory.

Custom pipelines may register @register_pipeline(name, config_resolver=resolver) when they support strict overrides and executed-configuration export. The resolver receives a defensive copy of the merged configuration plus is_training and returns the resolved four-stage configuration and model-input contract. Legacy factories without a resolver remain buildable, but strict resolution and return_config=True reject them because their behavior cannot be described completely.

Built-in crop strategies: random_resized, random_resized_hvflip, random_pad, random_hflip, resize_random_hflip, random_rot90_hflip.

Built-in augment strategies: rand_augment, trivial_augment, trivial_augment_wide, color_jitter, none.

Crop-specific parameters are passed through aug_kwargs["crop_kwargs"]. size, seed, interpolation, padding, and pad_mode remain controlled by the existing top-level augmentation fields and are rejected inside crop_kwargs, avoiding ambiguous precedence. For example, the FMoW M1 crop is fully declarative:

pipeline = get_pipeline(
    dataset="wilds:fmow",
    aug_kwargs={
        "enable": True,
        "image_size": 224,
        "crop_type": "random_resized_hvflip",
        "interpolation": "bicubic",
        "crop_kwargs": {
            "scale": (0.85, 1.0),
            "ratio": (0.90, 1.10),
            "horizontal_flip_probability": 0.5,
            "vertical_flip_probability": 0.5,
        },
    },
)

The acoustic package registers audio decoders, resamplers, channel strategies, segment strategies, frontends, augmentations, normalizations, eval views, postprocessors, and corruptions. Use list_audio_* helpers to inspect registered acoustic components.


Presets and Smart Merging

justdata.vision.presets stores dataset-specific default kwargs for all four vision pipeline stages. The vision _default preset (ImageNet statistics, 224px, RandAugment) serves as the fallback for vision only.

Available presets:

Preset Description
_default ImageNet-1K statistics, 224px, RandAugment, modern branch
cifar 32px, TrivialAugmentWide, CIFAR-10 normalization, no resizing
cifar100 32px, TrivialAugmentWide, CIFAR-100 normalization
imagenet_resnet ImageNet statistics, legacy branch with ColorJitter
imagenet_a1 RSB A1: heavy augmentation, BCE loss, 600 epochs
imagenet_a2 RSB A2: moderate augmentation, BCE loss, 300 epochs
imagenet_a3 RSB A3: light augmentation, CE loss, 160px training
dinov2 Asymmetric multi-crop SSL pipeline
wilds:* WILDS benchmark defaults; matching _strong presets are opt-in

merge_with_presets(dataset, user_kwargs) implements a smart merge: user-supplied kwargs that are identical to the _default preset values are treated as "not explicitly overridden," allowing dataset-specific preset values to take precedence. Only kwargs that genuinely differ from the defaults are considered intentional user overrides.

For an authoritative recipe, pass overrides= to get_pipeline. Explicit values win even when they are False, zero, None (where supported), or equal to the modality default. Nested mappings preserve siblings from the selected preset:

pipeline = get_pipeline(
    dataset="cifar10",
    overrides={
        "aug_kwargs": {"enable": False, "image_size": 224},
        "laug_kwargs": {"enable": False},
        "postproc_kwargs": {
            "image_size": 224,
            "val_resize_size": None,
            "normalization_params": (
                (0.485, 0.456, 0.406),
                (0.229, 0.224, 0.225),
            ),
        },
    },
)

Supplying overrides={} selects strict resolution without changing values. Built-in resolvers reject unknown nested keys, missing required settings, contradictory configurations, and attempts to set runtime-owned fields before a loader opens its source. Existing keyword arguments retain smart-merge compatibility semantics.

Resolved presets are serializable and hashable across modalities. Use get_resolved_preset(name) from justdata.vision.presets, justdata.acoustic.presets, or justdata.core.presets to obtain an object with .to_json() and .hash(). Hashes use canonical JSON with sorted keys and a 16-character SHA-256 prefix.

Acoustic presets are typed with AudioPreset and nested frozen config dataclasses for preprocessing, segmentation, frontends, labels, normalization, train augment settings, eval views, layout, and metadata policy. The justdata.audio namespace is a compatibility alias for justdata.acoustic.

See docs/presets.md for acoustic preset contracts, including EfficientAT/DyMN, PaSST, and CED.

Executed configuration snapshots

Pass return_config=True to load_ds or load_inventory to receive (dataset, n_batches, config). The immutable ExecutedConfig records the selected pipeline and preset, resolved four-stage configuration, derived model-input contract, train/evaluation mode, randomness, shuffle, batching, metadata, caches, execution limits, prefetching, and NumPy conversion. Its canonical versioned encoding is available as config.to_bytes():

import hashlib
from pathlib import Path

dataset, n_batches, config = load_ds(
    dataset_names_arg="cifar10",
    splits_arg="test",
    dataset_type="validation",
    batch_size=128,
    seed=42,
    pipeline=pipeline,
    num_classes=10,
    return_config=True,
)
encoded = config.to_bytes()
Path("executed-justdata-config.json").write_bytes(encoded)
full_digest = hashlib.sha256(encoded).hexdigest()

For a JD-01 offline inventory, use the same return option after verified admission:

import hashlib
from pathlib import Path

from justdata.core import load_inventory, open_inventory

admitted = open_inventory("/data/lars-admitted")
dataset, n_batches, config = load_inventory(
    admitted,
    dataset_type="validation",
    batch_size=8,
    seed=42,
    pipeline=pipeline,
    return_config=True,
)
encoded = config.to_bytes()
Path("lars-justdata-execution.json").write_bytes(encoded)
full_digest = hashlib.sha256(encoded).hexdigest()

With return_raw_ds=True, the loader returns (prepared, tools, config) and marks the remaining stages as pending. Both tools["finalize_fn"] and tools["finalize_epoch"] accept return_config=True and append a snapshot of their actual finalization settings. Snapshot export rejects opaque callback replacements because their behavior cannot be represented by configuration data.

The snapshot describes JustData's configured execution. A consuming product must separately bind admitted source inventory identities, external transformations, exact library revisions, model/checkpoint identities, and other run artifacts.

Numeric identity and metadata sidecars

With metadata_mode="numeric_only", pass metadata_sidecar= for a complete source mapping prepared before iteration. For an admitted inventory, use MetadataSidecar.from_inventory(admitted); for another finite source, use MetadataSidecar.from_metadata(source_records). Each real batch row then has metadata.row_id, an int64 join key, and metadata.row_fingerprint, a numeric check against stale cache entries. Resolve keys through sidecar.records only where padding_mask is true. The sidecar preserves original string IDs and static source metadata while numeric geometry remains with each view.

For a view with additional string metadata, use the numeric metadata.view_id with row_id to read sidecar.view_records[(row_id, view_id)].

sidecar.write_jsonl(path, policy="create") saves a complete mapping and MetadataSidecar.read_jsonl(path) reopens it. The streaming sidecar_metadata_path= route resumes and validates an existing file by default; sidecar_metadata_policy="create" rejects reuse and "overwrite" explicitly replaces it. Streaming artifacts can be partial after interrupted iteration. Executed configuration records the sidecar policy or full mapping digest.

Automatic preset resolution

When a vision dataset has a registered preset (e.g., cifar10 -> cifar, cifar100 -> cifar100), get_pipeline automatically applies it. Vision datasets without a dedicated preset fall back to the vision _default.

import justdata.vision
from justdata.core.registry import get_pipeline

# cifar10 automatically gets the 'cifar' preset (32px, TrivialAugmentWide, CIFAR-10 stats)
pipeline = get_pipeline(dataset="cifar10")

Applying a named preset to any dataset

The dataset argument in get_pipeline drives preset lookup, not just task inference. To apply a specific named preset to a dataset that does not have its own preset (e.g., using imagenet_a3 for imagenette), pass it as preset:

import justdata.vision
from justdata.core.loader import load_ds
from justdata.core.registry import get_pipeline

# Resolve the A3 (RSB light) pipeline for imagenette
pipeline = get_pipeline(
    dataset="imagenette",
    preset="imagenet_a3",       # drives preset lookup: 160px train, 224px val, RandAugment m=6
)

train_ds, N = load_ds(
    dataset_names_arg=["imagenette"],  # actual dataset to load
    splits_arg={"imagenette": ["train"]},
    dataset_type="train",
    batch_size=128,
    seed=42,
    pipeline=pipeline,
    num_classes=10,
    cache_dataset=False,
)

The validation pipeline uses the same preset name; the A3 preset's FixRes strategy (train_image_size=160, val_resize_size=236) is applied automatically:

# The same pipeline can be used for validation; `load_ds` automatically sets `is_training=False` based on `dataset_type`
# and applies the deterministic center-crop resize (FixRes 236 → CenterCrop 224)
val_ds, N = load_ds(
    dataset_names_arg=["imagenette"],
    splits_arg={"imagenette": ["validation"]},
    dataset_type="validation",
    batch_size=256,
    seed=0,
    pipeline=pipeline,
    num_classes=10,
    cache_dataset=False,
)

Acoustic Pipelines

Import justdata.acoustic before resolving built-in acoustic datasets or pipelines. Acoustic samples use the canonical keys waveform, sample_rate, optional label, features, duration, and metadata.

import justdata.acoustic
from justdata.core.loader import load_ds
from justdata.core.registry import get_pipeline

pipeline = get_pipeline(
    dataset="dcase2025_task1",
    preset="dcase2025_task1_efficientat_32k_1s",
)

ds, n = load_ds(
    dataset_names_arg=["dcase2025:task1"],
    splits_arg={"dcase2025:task1": ["dev_train_25"]},
    dataset_type="train",
    batch_size=64,
    seed=0,
    pipeline=pipeline,
    num_classes=10,
    cache_dataset=False,
    data_dir="/path/to/dcase",
    metadata_mode="numeric_only",
    as_numpy=True,
)

Modality documentation:

  • docs/vision.md: vision schema, four stages, preset contracts, metadata modes, Mini-C, and parity matrix.
  • docs/acoustic.md: canonical schema, four stages, metadata modes, golden tests, corruption benchmark, and parity matrix.
  • docs/dcase2025.md: DCASE Task 1 source/target helpers and split safety.
  • docs/presets.md: vision and acoustic preset selection and hashable contracts.
  • docs/golden_tests.md: optional golden compatibility test workflow.

Examples are split by modality under examples/vision/ and examples/acoustic/. Each directory includes loading, Hugging Face source, statistics, and corruption benchmark examples.


Dataset Adapters

Core adapters normalize raw records into the modality-specific schema expected by a selected pipeline. justdata.core itself is schema-neutral, and the identity adapter is applied when no dataset-specific adapter is registered.

Vision adapter

Vision pipelines expect image, optional label, and task-specific fields such as mask.

from justdata.core.adapters import register_adapter

@register_adapter("my_dataset")
def my_adapter(sample):
    return {"image": sample["img"], "label": sample["class_id"]}

Acoustic adapter

Acoustic pipelines expect waveform, sample_rate, and an optional label.

from justdata.core.adapters import register_adapter

@register_adapter("my_audio_dataset")
def my_audio_adapter(sample):
    result = {
        "waveform": sample["audio"],
        "sample_rate": sample["sampling_rate"],
    }
    if "target" in sample:
        result["label"] = sample["target"]
    return result

Mini-C Corruption Benchmark

create_minic_datasets constructs finalized corruption evaluation datasets by forking from a preprocessed RGB dataset. Corruption runs after decoding and rank/channel normalization, but before resize, float conversion, normalization, batching, and padding. Pass cache_dataset=True with an explicit disk path when avoiding source re-reads across corruptions is worth the storage cost; otherwise the base remains streaming.

The original Mini-C family names remain pinned to their 1.0.0 behavior:

Name Version 1 implementation
blur Defocus blur with a disk kernel and alias Gaussian filter
noise Additive Gaussian noise using the legacy absolute standard deviations
digital Nearest-neighbor pixelation
weather Synthetic snow

Five exact-name TensorFlow-native operators are also available at version 1.0.0: gaussian_blur, gaussian_noise, jpeg_compression, contrast_reduction, and brightness_reduction. Their immutable descriptors contain the severity table, domains, dtype/shape contract, seed policy, clipping and rounding policy, backend identity, and codec/filter settings. Use get_corruption_descriptor(name, version) or list_corruption_descriptors() to serialize and fingerprint these semantics.

For a caller-owned scientific seed, use exact version dispatch directly:

import tensorflow as tf

from justdata.vision.corruptions import (
    apply_corruption,
    get_corruption_descriptor,
)

descriptor = get_corruption_descriptor("gaussian_noise", "1.0.0")
corrupted = apply_corruption(
    decoded_rgb_uint8,
    name=descriptor.name,
    version=descriptor.version,
    severity=3,
    seed=tf.constant([123, 456], dtype=tf.int32),
)

The complete seed is supplied by the caller; no global RNG is read. gaussian_noise@1.0.0 uses TensorFlow Philox. The JPEG descriptor fixes encode and decode options and is repeatable for a fixed TensorFlow/codec build, but does not claim byte identity across unrelated TensorFlow or codec versions.

import justdata.vision
from justdata.core.registry import get_pipeline
from justdata.vision.minic import create_minic_datasets

pipeline = get_pipeline(dataset="imagenet")

datasets, n = create_minic_datasets(
    corruption_types=["gaussian_noise", "jpeg_compression", "blur"],
    severity=3,
    dataset_names_arg=["imagenet2012"],
    splits_arg={"imagenet2012": ["validation"]},
    dataset_type="validation",
    batch_size=256,
    seed=0,
    pipeline=pipeline,
    num_classes=1000,
)

With metadata_mode="full", finalized samples include corruption, corruption_version, corruption_identity, corruption_identity_hash, severity, and corruption_domain. numeric_only retains the numeric identity hash and severity; none removes metadata according to the shared core contract. The dataset helper preserves legacy enumeration-position seed salting. Consumers with their own sample-level PRNG lineage should call apply_corruption directly.


Audio Corruption Benchmark

create_audio_corruption_datasets mirrors Mini-C for acoustic evaluation. It applies deterministic severity 1-5 corruptions in the waveform or spectrogram domain, attaches corruption metadata, runs the normal postprocessing stage, and preserves padding_mask.

import justdata.acoustic
from justdata.acoustic.corruptions.datasets import create_audio_corruption_datasets
from justdata.core.registry import get_pipeline

pipeline = get_pipeline(
    dataset="dcase2025_task1",
    preset="dcase2025_task1_efficientat_32k_1s",
)

datasets, n = create_audio_corruption_datasets(
    corruption_types=["additive_white_noise", "clipping"],
    severity=3,
    base_dataset="dcase2025:task1",
    preset="dcase2025_task1_efficientat_32k_1s",
    split="dev_test",
    pipeline=pipeline,
    batch_size=64,
    seed=0,
    num_classes=10,
    metadata_mode="numeric_only",
)

Usage

Basic Classification Pipeline

import justdata.vision
from justdata.core.loader import load_ds
from justdata.core.registry import get_pipeline

pipeline = get_pipeline(dataset="cifar10")

train_ds, N = load_ds(
    dataset_names_arg=["cifar10"],
    splits_arg={"cifar10": ["train"]},
    dataset_type="train",
    batch_size=128,
    seed=42,
    pipeline=pipeline,
    num_classes=10,
    # CIFAR-10 is small enough for an intentional memory cache.
    cache_dataset=True,
)

Reusable, Explicitly Seeded Epochs

Use return_raw_ds=True to prepare the source, pipeline, preprocessing, cache, and filters once, then materialize the complete remaining training pipeline for each epoch:

prepared_train, tools = load_ds(
    dataset_names_arg=["cifar10"],
    splits_arg={"cifar10": ["train"]},
    dataset_type="train",
    batch_size=128,
    seed=0,
    pipeline=pipeline,
    num_classes=10,
    cache_dataset=False,
    deterministic=True,
    return_raw_ds=True,
)

for epoch in range(num_epochs):
    train_iterator, n_batches = tools["finalize_epoch"](
        prepared_train,
        seed=derive_epoch_seed(epoch),
        as_numpy=True,
    )
    train_one_epoch(train_iterator)

finalize_epoch applies standard augmentation, shuffle, postprocessing, batching, late augmentation, padding, and prefetching. The epoch seed is split into standard- and late-augmentation seeds, with source and batch indices folded in. Shuffle uses the epoch seed with reshuffling disabled on the materialized dataset. With deterministic=True, repeating a seed therefore reproduces the epoch independently of previous iterator creation or consumption, including when the returned TensorFlow dataset is iterated more than once.

For an augmentation callback that depends only on its input sample and supplied seed, pass augment_is_stateless=True to finalize_epoch to run standard augmentation concurrently using map_parallel_calls (or AUTOTUNE). With deterministic=True, sample order and indexed RNG seeds remain unchanged:

train_iterator, n_batches = tools["finalize_epoch"](
    prepared_train,
    seed=derive_epoch_seed(epoch),
    augment_is_stateless=True,
    as_numpy=True,
)

This flag is the caller's assertion that the augmentation callback uses no stateful RNG, mutable state, or side effects. It defaults to False, preserving serial standard augmentation for deterministic epochs. It does not change the one-shot load_ds or finalize_fn seed scheduling.

With cache_dataset=False, source samples, preprocessing, and filters still run on every iteration; only loader and graph preparation are reused. A fresh NumPy iterator is created by each finalize_epoch(..., as_numpy=True) call.

finalize_epoch rejects cache_model_inputs=True when training or evaluation augmentation is enabled because that downstream cache intentionally freezes the first sampled augmented views. Existing one-shot load_ds and finalize_fn behavior is unchanged.

For a verified local inventory, load_replay_epoch binds an immutable snapshot and executed configuration to an explicit committed-batch cursor. It rebuilds the full epoch before skipping committed batches. See docs/inventory.md for checkpoint and resume usage. Count real rows with count_real_examples(batch) or the returned epoch's remaining_examples; padded rows have a false padding_mask. The epoch path also accepts prefetch=False when finalization must precede a caller-controlled prefetch stage.

Bounded input execution and protected caches

Pass prefetch=1 (or another positive batch count) to load_ds, load_inventory, or finalize_fn to bound final prefetch. finalize_epoch inherits the configured value unless overridden. False disables prefetch; True retains automatic tuning. A bounded input profile also sets positive integer map_parallel_calls, private_threadpool_size, and max_intra_op_parallelism, plus a finite shuffle buffer. The executed configuration records the effective values. Replay prefetches after skipping committed batches and defaults to one batch.

When another framework uses the accelerator, configure TensorFlow before creating tensors or datasets:

from justdata.core import configure_tensorflow_cpu

device_report = configure_tensorflow_cpu(intra_op_threads=1, inter_op_threads=1)
assert device_report["logical_gpus"] == []

The report includes physical and logical device names and process thread settings. The helper raises if the runtime has initialized too early for the requested change. Core, vision, acoustic, and audio imports permit this setup.

CachePolicy enables identity-checked file caches on the existing preprocessing and model-input cache stages. The cache path is a directory whose parent exists. The input identity must cover ordered content. load_inventory binds its verified manifest automatically; other sources require a content-derived input_identity. Callers must declare deterministic callbacks because arbitrary callbacks cannot be inspected for stateful random operations.

from justdata.core import CachePolicy, inspect_cache, load_inventory

policy = CachePolicy(
    max_bytes=1_073_741_824,
    max_examples=256,
    materialization="lazy",  # choose "eager" to finish before returning
    callbacks_are_deterministic=True,
)
batches, count, config = load_inventory(
    admitted, "validation", 2, 17,
    pipeline=pipeline,
    cache_dataset=True,
    cache_path="/existing/cache-parent/prepared",
    cache_policy=policy,
    return_config=True,
    map_parallel_calls=2,
    private_threadpool_size=2,
    max_intra_op_parallelism=1,
    prefetch=1,
)
for batch in batches:
    consume(batch)
assert policy.status("preprocess")["state"] == "complete"
assert inspect_cache("/existing/cache-parent/prepared")["state"] == "complete"

Lazy caches publish completion only after the source is fully exhausted. Eager caches finish before loading returns. Both modes record count, byte size, checksum, and a versioned stage/configuration identity. Changed input, normalization, geometry, or augmentation configuration rejects reuse at the same path. An incomplete or incompatible cache raises CacheError; choose a new path after inspecting or clearing the caller-owned artifact. Cache errors expose code and path. During lazy iteration TensorFlow may wrap the exception; policy.status(stage) retains its structured failure code. Quotas bound record count and serialized record bytes. Protected model-input caches require evaluation without active augmentation; preprocessing caches remain before per-epoch augmentation. Existing cache flags without a policy retain their behavior.

Loading a Hugging Face Dataset

train_ds, n_batches = load_ds(
    dataset_names_arg=["hf:cifar10"],
    splits_arg={"hf:cifar10": ["train"]},
    cache_dataset=False,
    ...
)

load_ds returns the dataset and its batch cardinality. The cardinality is a Python integer when TensorFlow can determine it and None for unknown or infinite pipelines.

For remote/downloaded sources, data_dir is a cache root. justdata namespaces source-owned caches under it: tfds/, hf/vision/, hf/acoustic/, wilds/, and zenodo/. If omitted, the root is ~/.cache/justdata.

Loader caches are separate from source-owned caches. cache_dataset/cache_path caches decoded, preprocessed samples before augmentation. Loader caching is off by default, so datasets remain streaming when cache_dataset is omitted. Set cache_dataset=True with an empty cache_path to opt into memory caching, or provide a nonempty path to cache on the filesystem. Large datasets should use an explicit disk path or remain uncached. cache_model_inputs and model_input_cache_path cache resized, normalized model inputs before batching. Use different paths for the two cache stages. This default changes performance, not output values.

Loading a WILDS Dataset

WILDS image classification datasets use the wilds: source prefix. Source options such as FMoW temporal split schemes live in the dataset string; when using a splits_arg dictionary, reuse that exact string as the key. Downloads are disabled by default (download=false). To download through WILDS under data_dir/wilds, add download=true to the dataset string. Use & between query options, not a second ?. The automatic presets are benchmark-faithful: Camelyon17 96 px, FMoW 224 px, iWildCam 448 px, and RxRx1 256 px with per-image channel standardization. Use preset="wilds:fmow_strong" or another _strong name to opt into stronger training augmentation.

local_ssd = "/local_ssd/justdata"
wilds_ds = "wilds:fmow?split_scheme=time_after_2016&download=true"
pipeline = get_pipeline(dataset="wilds:fmow")

train_ds, N = load_ds(
    dataset_names_arg=[wilds_ds],
    splits_arg={wilds_ds: ["train"]},
    dataset_type="train",
    batch_size=32,
    seed=42,
    pipeline=pipeline,
    num_classes=62,
    data_dir=f"{local_ssd}/sources",
    cache_dataset=True,
    cache_path=f"{local_ssd}/decoded/fmow-train",
    cache_model_inputs=True,
    model_input_cache_path=f"{local_ssd}/model-inputs/fmow-train-224",
    allow_train_model_input_cache=True,
    metadata_mode="numeric_only",
)

FMoW source inventories can opt into authoritative sequence identity and acquisition time without changing the default WILDS signature:

from justdata.core import fetch_ds

fmow = (
    "wilds:fmow?split_scheme=official&version=1.1&download=false"
    "&source_metadata=location_id,timestamp"
)
inventory = fetch_ds([fmow], {fmow: ["train", "id_val", "id_test", "val"]})

for sample in inventory.as_numpy_iterator():
    source = sample["metadata"]["wilds_source"]
    location_id = source["location_id"]  # UTF-8 bytes
    timestamp = source["timestamp"]      # UTF-8 bytes

The option accepts only location_id and timestamp. JustData maps every emitted wilds_index through WILDS' authoritative full_idxs mapping: location_id is the exact original FMoW sequence-directory basename, and timestamp preserves the source timezone-aware ISO-8601 text. Invalid fields, target-equivalent fields, missing source columns, or unreliable mappings fail closed. The wilds_source mapping is absent without the opt-in; it is retained by metadata_mode="full" and removed by numeric_only.

The base wilds:fmow preset has no stochastic per-sample training augmentation, so train model-input caching is safe when explicitly opted in. Do not use train model-input caching with wilds:fmow_strong unless freezing the first pass of stochastic augmentation is intentional.

Applying a Named Preset

To use a specific preset for any dataset, pass the preset name as preset:

# Load imagenette with the RSB A3 (light) recipe: 160px training, RandAugment m=6, Mixup α=0.1
pipeline = get_pipeline(
    dataset="imagenette",
    preset="imagenet_a3",  # apply the A3 preset to the imagenette task
)

train_ds, N = load_ds(
    dataset_names_arg=["imagenette"],
    splits_arg={"imagenette": ["train"]},
    dataset_type="train",
    batch_size=128,
    seed=42,
    pipeline=pipeline,
    num_classes=10,
    cache_dataset=False,
)

Overriding Preset Parameters

pipeline = get_pipeline(
    dataset="imagenet",
    aug_kwargs={"augment_type": "trivial_augment"},  # overrides preset default
)

DINOv2 SSL Pipeline

pipeline = get_pipeline(
    preset="dinov2",
    pipeline_name="vision/classification",
)

Development

This project uses devenv (Nix-based) with uv for Python dependency management.

devenv shell        # enter the development environment
direnv allow        # approve automatic devenv activation once
uv sync             # install/update dependencies
uv run pytest       # run all tests
uv run ruff check . # run lint checks
uv run ruff format --check .  # check formatting
uv run ty check src/          # type check the package
uv run mdformat --check --wrap 120 ./*.md benchmarks docs examples src tests  # check Markdown formatting
uv run pytest tests/path/to/test_file.py::test_name  # run a single test

The project targets Python 3.12 (see .python-version). LD_LIBRARY_PATH is configured by devenv for native libraries.

See Performance checks for reproducible CPU benchmarks and the performance audit results.

GitLab runs sh tests/doctor/test_lint.sh in a pinned Debian-based uv image. It checks Python with Ruff and Markdown with mdformat using the GFM, footnote, and MyST plugins from the lint dependency group alone, then type checks src/ with ty against the locked development environment so that third-party imports resolve. TensorFlow attaches Tensor operators at runtime, so ty treats tensorflow imports as Any (see [tool.ty.analysis] in pyproject.toml). The image provides the standard Linux dynamic loader required by the Ruff and ty PyPI executables, which the Nix image lacks.

Releases

GitLab builds a wheel and source distribution after lint, the default test suite, and all Python 3.11–3.13 package-support jobs pass. Build artifacts are retained for one week. Tag pipelines publish those artifacts to PyPI after the existing manual golden job also passes; start that job in GitLab to complete a release.

Before tagging, update pyproject.toml, src/justdata/__init__.py, and CHANGELOG.md, then run uv lock. Push a tag matching the package version, with an optional v prefix (for example, v1.2.1 or 1.2.1). Publishing rejects tags that disagree with either version declaration. Branch pipelines build artifacts without publishing.

Configure a PyPI GitLab trusted publisher for the justdata project with namespace rhizome-labs/public, project justdata, pipeline path .gitlab-ci.yml, and environment release. The publish job requests a GitLab ID token with audience pypi and requires trusted publishing; no stored PyPI API token is needed.

Metadata

Release files for justdata 1.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for justdata 1.2.1
File Size Uploaded
justdata-1.2.1.tar.gz 228.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for justdata 1.2.1
File Interpreter ABI Platform
justdata-1.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 492.5 kB

Release files / justdata-1.2.1.tar.gz

Download URL justdata-1.2.1.tar.gz
Size 228.7 kB
Tags Source
SHA-256 checksum
How to use checksums
056fdc8095add3a66c462e84f5bc2a9fbf3d624c2f939a61d147d019f48b09c0
BLAKE2b-256 checksum
How to use checksums
b020a96ba4bde6584a61f2703125174db765cb1005c2a437ab75416a991a9bb1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.16 {"installer":{"name":"uv","version":"0.11.16","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / justdata-1.2.1-py3-none-any.whl

Download URL justdata-1.2.1-py3-none-any.whl
Size 263.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5474d795e8d8452d0a08012b52b2a73c18825ffdf3f40c52a853e804be8a0235
BLAKE2b-256 checksum
How to use checksums
0f2c069e70c471d80982bba85099cd7fec997f1cb205b932ffa0d6fc7c37d8aa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.16 {"installer":{"name":"uv","version":"0.11.16","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

1.2.1 This release

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page