Skip to main content

hf-dataset-slow-disk

A Hugging Face-style Parquet dataset library for HDDs, network filesystems, and other slow storage.

Features

  • Sequential row-group reads with bounded memory
  • Lazy background download and audio decoding
  • Standard Hugging Face Hub Parquet cache
  • Resumable general metadata cache with visible progress logs
  • Deterministic shuffle, epochs, and exact batch resume
  • Sample-mixed concatenation with one shared worker pipeline

Install

pip install hf-dataset-slow-disk

Quick start

from hf_dataset_slow_disk import load_dataset

dataset = load_dataset(
    "capleaf/viVoice",
    split="train",
    batch_size=8,
    num_workers=2,
).shuffle(seed=42)

dataset.set_epoch(0)

try:
    for batch in dataset:
        train_step(batch)
finally:
    dataset.close()

Workers start lazily on the first get() or when dataset iteration begins. Defaults are one worker, an eight-batch decoded buffer, and a two-file Parquet buffer.

Core API

from hf_dataset_slow_disk import (
    Audio,
    ConcatDataset,
    Dataset,
    DatasetDict,
    concatenate_datasets,
    load_dataset,
)
  • load_dataset() loads Hub or local Parquet data.
  • shuffle(seed) and set_epoch(epoch) select deterministic order.
  • set_batch_idx(batch_idx) restores the next training batch.
  • for batch in dataset, get(), and iter_batches() consume batches.
  • iter_rows() performs synchronous sample-by-sample iteration.
  • concatenate_datasets() mixes compatible sources without oversampling.
  • cast_column() configures lazy audio decoding.

Documentation

Examples

uv run python examples/basic.py
uv run python examples/benchmark.py
uv run python examples/concat.py

The benchmark and concat examples use batch size 8 and stop after 10 batches.

Disclaimer

This project was developed with the assistance of AI coding tools. Its source code is publicly available at giangndm/hf-dataset-slow-disk. Please review and validate it for your own workloads before production use.

Development

uv sync
uv run pytest -q
uv build

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hf_dataset_slow_disk-0.1.0.tar.gz (84.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

hf_dataset_slow_disk-0.1.0-py3-none-any.whl (23.7 kB view details)

Uploaded Python 3

File details

Details for the file hf_dataset_slow_disk-0.1.0.tar.gz.

File metadata

File hashes

Hashes for hf_dataset_slow_disk-0.1.0.tar.gz
Algorithm Hash digest
SHA256 eebc97b0c8fd2e4f6005de31a32aa1272ab078acd7221dd02c94ae5022366f20
MD5 439ee159e44092510e89ea5a9df93d28
BLAKE2b-256 5d5f3684eb08c0c5480dfaa901b96368df8005fc78858325785d0b009bfd8c0a

See more details on using hashes here.

File details

Details for the file hf_dataset_slow_disk-0.1.0-py3-none-any.whl.

File metadata

File hashes

Hashes for hf_dataset_slow_disk-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b6b507b073f3928e61b1b737e4f6a3b149d907090bde7cfd3d700a8c0aa5bfed
MD5 4b68bd53223acf55824537d9aef8fbfb
BLAKE2b-256 2b0c457889f358220bf3036b638bca37951b7f495cebbf1695eee07c4e78c4bd

See more details on using hashes here.

Release history Release notifications | RSS feed

0.2.1

2 files

0.2.0

2 files

0.1.2

2 files

0.1.1

2 files

This release

0.1.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page