Skip to main content

hf-dataset-slow-disk

A Hugging Face-style Parquet dataset library for HDDs, network filesystems, and other slow storage.

Features

  • Sequential row-group reads with bounded memory
  • Lazy background download and audio decoding
  • Standard Hugging Face Hub Parquet cache
  • Resumable general metadata cache with visible progress logs
  • Deterministic shuffle, epochs, and exact batch resume
  • Sample-mixed concatenation with one shared worker pipeline

Install

pip install hf-dataset-slow-disk

Quick start

from hf_dataset_slow_disk import load_dataset

dataset = load_dataset(
    "capleaf/viVoice",
    split="train",
    batch_size=8,
    num_workers=2,
).shuffle(seed=42)

dataset.set_epoch(0)

try:
    for batch in dataset:
        train_step(batch)
finally:
    dataset.close()

Workers start lazily on the first get() or when dataset iteration begins. Defaults are one worker, an eight-batch decoded buffer, and a two-file Parquet buffer.

Core API

from hf_dataset_slow_disk import (
    Audio,
    ConcatDataset,
    Dataset,
    DatasetDict,
    concatenate_datasets,
    load_dataset,
)
  • load_dataset() loads Hub or local Parquet data.
  • shuffle(seed) and set_epoch(epoch) select deterministic order.
  • set_batch_idx(batch_idx) restores the next training batch.
  • for batch in dataset, get(), and iter_batches() consume batches.
  • iter_rows() performs synchronous sample-by-sample iteration.
  • concatenate_datasets() mixes compatible sources without oversampling.
  • cast_column() configures lazy audio decoding.

Documentation

Examples

uv run python examples/basic.py
uv run python examples/benchmark.py
uv run python examples/concat.py

The benchmark and concat examples use batch size 8 and stop after 10 batches.

Disclaimer

This project was developed with the assistance of AI coding tools. Its source code is publicly available at giangndm/hf-dataset-slow-disk. Please review and validate it for your own workloads before production use.

Development

uv sync
uv run pytest -q
uv build

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hf_dataset_slow_disk-0.1.1.tar.gz (85.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

hf_dataset_slow_disk-0.1.1-py3-none-any.whl (23.9 kB view details)

Uploaded Python 3

File details

Details for the file hf_dataset_slow_disk-0.1.1.tar.gz.

File metadata

File hashes

Hashes for hf_dataset_slow_disk-0.1.1.tar.gz
Algorithm Hash digest
SHA256 2a23ff7eef0bdad610970ad29cc7de0cd4de5428eb36e069402df2de7e3a1449
MD5 28ecee68ef43590684164cc8d41f8c86
BLAKE2b-256 b5b9fbdacdf07a6d96a724f1000210e8e50040374445d006e6a5d3379246570f

See more details on using hashes here.

File details

Details for the file hf_dataset_slow_disk-0.1.1-py3-none-any.whl.

File metadata

File hashes

Hashes for hf_dataset_slow_disk-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 fb97ec3ac967a13183ec9d3453bfd1b453c196b231f5d87c953ce66d64a74604
MD5 868f5e65f5685bdb722399b72808d081
BLAKE2b-256 f5b46f2a3af492d907145ccc7d27711de741a8aa25e048d491975a7d2e06852d

See more details on using hashes here.

Release history Release notifications | RSS feed

0.2.1

2 files

0.2.0

2 files

0.1.2

2 files

This release

0.1.1 This release

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page