Skip to main content

Minimal Dependency Dataset Library for PyTorch

Bachelor Project — DFKI (German Research Center for Artificial Intelligence) | SS 2026

Project Description

Training large deep learning models on GPU clusters requires efficient data pipelines. When datasets reside on remote network storage, loading and preprocessing can become the primary bottleneck, leaving expensive GPUs underutilized. Existing solutions (WebDataset, MosaicML StreamingDataset, TFRecord) often require extensive dependencies and major changes to training code.

The goal of this project is to analyse performance bottlenecks in the data loading pipeline, then design and implement a simple library for storing and loading training data from a remote storage server. The library should be minimal — relying only on PyTorch and the Python standard library — while supporting multi-threaded loading, lock-free sampling, distributed training (DDP), and built-in instrumentation for performance profiling.

Project Goals & Status

# Goal Status
1 Train baseline model + profile data loading bottlenecks Done
2 Benchmark GPU compute delays across 10 GPU types (3,540 runs) Done
3 Build GPU simulator (Dummy Model) for GPU-free testing Done
4 Evaluate 6 storage formats for throughput, overhead, thread safety Done
5 Implement multi-threaded DataLoader with lock-free sampler Done
6 Add built-in instrumentation (MonitoredQueue + MetricsTracker) Done
7 Run benchmark sweep: 9 batch sizes × 6 worker counts, 100k images Done
8 Compare against PyTorch DataLoader on same hardware/data Done

Key Results

Storage Format Evaluation (100k images, tiny-imagenet, fscratch SSD)

Format Sequential (img/s) Random (img/s) Overhead Thread-Safe
Parquet 472.7 473.7 -0.3% Yes
MessagePack 431.4 312.9 0.0% Conditional
Tar 424.1 N/A 1.5% Sequential only
LMDB 415.8 304.7 1.8% Yes
Zip 341.5 233.2 0.1% No
Plain JPEG 266.6 255.9 0% Yes

DataLoader Performance (Parquet, A100 CPU node, 100k images)

Metric Our DataLoader PyTorch DataLoader
Peak Throughput 1,787 samples/s (BS=512, 16w) 2,865 samples/s (BS=512, 32w)
Optimal Workers 16 32
Scaling 1-16 workers 6.5x 9.8x
Scaling 16-32 workers Degrades (-10%) Continues (+8%)
Memory 14-17 GB 16 GB
Dependencies stdlib + torch PyTorch only (stdli + torch)

GPU Compute Delays (ResNet-50, batch_size=256, SGD)

GPU Avg Delay (ms) Models Tested
H200 522 15
H100 597 15
RTXB6000 537 15
L40S 663 11
A100-40GB 702 9
B200 711 16
RTXA6000 1,042 9

Repository Structure

bachelor-project/
├── minimal_dataset/ # Library module
│ ├── init.py # Public API
│ ├── dataset.py # BaseDataset (ImageFolder reader)
│ ├── parquet_dataset.py # ParquetDataset (Parquet reader)
│ ├── sampler.py # LockFreeSampler (lock-free index partitioning)
│ ├── dataloader.py # DataLoader (multi-threaded, staging + batch queues)
│ ├── monitored_queue.py # MonitoredQueue (queue instrumentation)
│ └── metrics.py # MetricsTracker + WorkerMetrics
├── benchmarks/
│ ├── gpu/ # GPU benchmarking & dummy model
│ │ ├── benchmark_gpu.py
│ │ ├── calibrate_dummy.py
│ │ ├── dummy_model.py
│ │ ├── build_delay_config.py
│ │ └── gpu_delays.json
│ └── storage/ # Storage format evaluation
│ ├── benchmark_plain.py
│ ├── benchmark_zip.py
│ ├── benchmark_tar.py
│ ├── benchmark_lmdb.py
│ ├── benchmark_parquet.py
│ ├── benchmark_msgpack.py
│ └── file_io.py
├── tests/ # DataLoader tests & benchmarks
│ ├── test_dataloader.py
│ ├── benchmark_dataloader.py
│ ├── benchmark_pytorch.py
│ ├── plot_comparison.py
│ └── launch_full_benchmark.sh
├── training/ # ResNet training scripts
│ ├── train_resnet50.py
│ ├── train_resnet50_ddp.py
│ └── analyze_dataloading.py
├── docs/
│ ├── images/ # Architecture diagrams & benchmark plots
│ │ ├── 00_MainOverview1.png
│ │ ├── 02_dataLoader1.png
│ │ ├── 04_Storage Format Evaluation.png
│ │ ├── 05_GPU Compute Delay.png
│ │ ├── plot_ours_throughput.png
│ │ ├── plot_pytorch_throughput.png
│ │ ├── plot_comparison_bs16.png
│ │ └── plot_speedup_comparison.png
│ └── architecture.md
└── README.md

Release files for minimal-dataset-pytorch 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for minimal-dataset-pytorch 0.1.0
File Size Uploaded
minimal_dataset_pytorch-0.1.0.tar.gz 16.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for minimal-dataset-pytorch 0.1.0
File Interpreter ABI Platform
minimal_dataset_pytorch-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 31.2 kB

Release files / minimal_dataset_pytorch-0.1.0.tar.gz

Download URL minimal_dataset_pytorch-0.1.0.tar.gz
Size 16.4 kB
Tags Source
SHA-256 checksum
How to use checksums
05c34b1e08f351a8e83bfff881f9398d6353183a3c3c0f469166e45d46124bc1
BLAKE2b-256 checksum
How to use checksums
47efad86411b40fb86d6a54a3e9ba8c7ef0f0dbb9c62249fa6205a697b67694e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.

Transparency log

Release files / minimal_dataset_pytorch-0.1.0-py3-none-any.whl

Download URL minimal_dataset_pytorch-0.1.0-py3-none-any.whl
Size 14.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9a874dc7821ad070bcd2c3ed7fd0c008210a3a4e4a820aa0c7abb1c665ba5687
BLAKE2b-256 checksum
How to use checksums
9d999f1bc35796607d8757922296cdeef07704d54fce90c8108054d393af1e41
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page