Skip to main content
image

Fast dataloader and conversion utility for webdataset tar shards. Rust core with Python bindings.

Built for streaming large video and image datasets, but handles any byte data.

Install

pip install webshart

What is this?

Webshart is a fast reader for webdataset tar files with separate JSON index files. This format enables random access to any file in the dataset without downloading the entire archive.

The indexed format provides massive performance benefits:

  • Random access: Jump to any file instantly
  • Selective downloads: Only fetch the files you need
  • True parallelism: Read from multiple shards simultaneously
  • Cloud-optimized: Works efficiently with HTTP range requests
  • Aspect bucketing: Optionally include image geometry hints width, height and aspect for the ability to bucket images by shape
  • Logical sample APIs: Treat image.ext + image.json pairs as one sample while still allowing raw file access
  • Caption metadata: Store captions in shard metadata under the plural captions key as either a string or a list of strings
  • Custom DataLoader: Includes state dict methods on the DataLoader so that you can resume training deterministically
  • Rate-limit friendly: Local caching allows high-frequency random seeking without encountering storage provider rate limits
  • Instant start-up with pre-sorted aspect buckets

Growing ecosystem: While not all datasets use this format yet, you can easily create indices for any tar-based dataset (see below).

Quick Start

import webshart

# Find your dataset
dataset = webshart.discover_dataset(
    source="laion/conceptual-captions-12m-webdataset",
    # we're able to upload metadata separately so that we reduce load on huggingface infra.
    metadata="webshart/conceptual-captions-12m-webdataset-metadata",
)
print(f"Found {dataset.num_shards} shards")

loader = webshart.TarDataLoader(dataset)

# File-oriented access is still available.
files = dataset.list_files_in_shard(0)

# Sample-oriented access skips paired JSON sidecars.
samples = dataset.list_samples_in_shard(0)
entry = loader.load_sample(0, 0)
print(entry.path, entry.captions, entry.json_metadata)

Common Patterns

For real-world, working examples:

Creating Indices for / Converting Existing Datasets

Any tar-based webdataset can benefit from indexing! Webshart includes tools to generate indices:

A command-line tool that auto-discovers tars to process:

% webshart extract-metadata \
    --source laion/conceptual-captions-12m-webdataset \
    --destination laion_output/ \
    --checkpoint-dir ./laion_output/checkpoints \
    --max-workers 2 \
    --include-image-geometry

Or, if you prefer/require direct-integration to an existing Python application, use the API

Uploading Indices to HuggingFace

Once you've generated indices, share them with the community:

# Upload all JSON files to your dataset
huggingface-cli upload --repo-type=dataset \
    username/dataset-name \
    ./indices/ \
    --include "*.json" \
    --path-in-repo "indices/"

Or if you want to contribute to an existing dataset you don't own:

  1. Create a community dataset with indices: username/original-dataset-indices
  2. Upload the JSON files there
  3. Open a discussion on the original dataset suggesting they add the indices

Creating New Indexed Datasets

If you're creating a new dataset, generate indices during creation:

{
  "files": {
    "image_0001.webp": {"offset": 512, "length": 102400},
    "image_0002.webp": {"offset": 102912, "length": 98304},
    ...
  }
}

The JSON index should have the same name as the tar file (e.g., shard_0000.tarshard_0000.json).

Image + JSON Sidecar Samples

Webshart supports webdataset shards that store each sample as an image-like payload plus a paired JSON sidecar:

sample_0001.webp
sample_0001.json
sample_0002.webp
sample_0002.json

When metadata is extracted or loaded, sidecars are attached to their paired sample entries:

{
  "files": {
    "sample_0001.webp": {
      "offset": 512,
      "length": 102400,
      "width": 1024,
      "height": 1024,
      "aspect": 1.0,
      "json_path": "sample_0001.json",
      "json_offset": 103424,
      "json_length": 128,
      "captions": "a product photo on a white background",
      "json_metadata": {
        "caption": "a product photo on a white background"
      }
    },
    "sample_0001.json": {
      "offset": 103424,
      "length": 128
    }
  }
}

Use file-oriented APIs when you want every archive member, including sidecars:

dataset.list_files_in_shard(0)

reader = dataset.open_shard(0)
raw_file_bytes = reader.read_file(0)

Use sample-oriented APIs when you want training samples:

dataset.list_samples_in_shard(0)
dataset.get_shard_sample_count(0)

reader = dataset.open_shard(0)
image_bytes = reader.read_sample(0)
json_bytes = reader.read_sample_json(0)

entry = loader.load_sample(0, 0)
print(entry.path)
print(entry.captions)
print(entry.json_data)

Captions are canonicalized to the plural captions metadata key. The value may be a single string, a list of strings, or absent.

webshart.write_captions_to_metadata(
    "shard_0000.json",
    {
        "sample_0001.webp": "a short caption",
        "sample_0002": ["caption one", "caption two"],
    },
)

The writer updates existing webshart metadata JSON in place, removes old singular caption keys from updated samples, and leaves paired .json sidecar entries untouched.

Aspect Bucketing Samples

list_shard_aspect_buckets() is file-oriented and buckets any indexed file that has width and height.

For training pipelines, prefer list_shard_sample_aspect_buckets():

loader = webshart.TarDataLoader(dataset)
buckets = loader.list_shard_sample_aspect_buckets(
    [0],
    key="geometry-tuple",
    target_pixel_area=1024**2,
)[0]["buckets"]

for bucket_key, entries in buckets.items():
    for item in entries:
        virtual_id = f"webshart://0/{item['sample_idx']}/{item['filename']}"
        image = loader.load_sample(0, item["sample_idx"])

This uses logical samples from metadata.sample_range() / get_sample_by_index() and excludes paired JSON sidecars before bucketing. Each bucket entry includes sample_idx, so callers can build stable IDs and load images directly with loader.load_sample(shard_idx, sample_idx).

Why is it fast?

Problem: Standard tar files require sequential reading. To get file #10,000, you must read through files #1-9,999 first.

Solution: The indexed format stores byte offsets and sample metadata in a separate JSON file, enabling:

  • HTTP range requests for any file
  • True random access over network
  • Parallel reads from multiple shards
  • Large scale, aspect-bucketed datasets
  • No wasted bandwidth

The Rust implementation provides:

  • Real parallelism (no Python GIL)
  • Zero-copy operations where possible
  • Efficient HTTP connection pooling
  • Optimized tokio async runtime
  • Optional local caching for metadata and shards
  • Fast aspect bucketing for image data

Datasets Using This Format

I discovered after creating this library that cheesechaser is the origin of the indexed tar format, which webshart has formalised and extended to include aspect bucketing support.

Requirements

  • Python 3.12+
  • Linux/macOS/Windows

Roadmap

  • image decoding is currently not handled by this library, but it will be added with zero-copy.
  • more informative API for caching and other Rust implementation details
  • multi-gpu/multi-node friendly dataloader

Projects using webshart

  • CaptionFlow uses this library to solve memory use and seek performance issues typical to webdatasets

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

webshart-0.5.0.tar.gz (102.5 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

webshart-0.5.0-cp312-abi3-win_amd64.whl (4.0 MB view details)

Uploaded CPython 3.12+Windows x86-64

webshart-0.5.0-cp312-abi3-manylinux_2_39_x86_64.whl (4.4 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.39+ x86-64

webshart-0.5.0-cp312-abi3-manylinux_2_39_aarch64.whl (4.1 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.39+ ARM64

webshart-0.5.0-cp312-abi3-manylinux_2_35_x86_64.whl (4.4 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.35+ x86-64

webshart-0.5.0-cp312-abi3-macosx_11_0_arm64.whl (4.4 MB view details)

Uploaded CPython 3.12+macOS 11.0+ ARM64

File details

Details for the file webshart-0.5.0.tar.gz.

File metadata

  • Download URL: webshart-0.5.0.tar.gz
  • Upload date:
  • Size: 102.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for webshart-0.5.0.tar.gz
Algorithm Hash digest
SHA256 6085ca5b09895fc231047682db4eebe805abdbb5e2979ff33e6fb46ccb95dd20
MD5 9c7520caf3e396c9ce4cc296cdefed93
BLAKE2b-256 12d1bef819e86506cf93df8e5b664559ad81717d24921d8538ccfc297be91a32

See more details on using hashes here.

File details

Details for the file webshart-0.5.0-cp312-abi3-win_amd64.whl.

File metadata

  • Download URL: webshart-0.5.0-cp312-abi3-win_amd64.whl
  • Upload date:
  • Size: 4.0 MB
  • Tags: CPython 3.12+, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for webshart-0.5.0-cp312-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 cfe01f06e61504ba765fffe078c33205a460d73a3548ca9eaf3af452239ef0b7
MD5 1c9a58404e1ac891bbaf2bf212bf5401
BLAKE2b-256 91f05e3beae1c2c7a989775d890d05bffca0f0528aba2852580d4de03046403f

See more details on using hashes here.

File details

Details for the file webshart-0.5.0-cp312-abi3-manylinux_2_39_x86_64.whl.

File metadata

File hashes

Hashes for webshart-0.5.0-cp312-abi3-manylinux_2_39_x86_64.whl
Algorithm Hash digest
SHA256 8f70d04cce241fd869171671d2d1d751663edd88d3f95505a34dc08fe5a66c6b
MD5 53aa95299d3c9042dee86f7e2963be84
BLAKE2b-256 32f8e3891034e5480f0de099a8da1fac1c245704a433ac3a3ec97e3e8a435e41

See more details on using hashes here.

File details

Details for the file webshart-0.5.0-cp312-abi3-manylinux_2_39_aarch64.whl.

File metadata

File hashes

Hashes for webshart-0.5.0-cp312-abi3-manylinux_2_39_aarch64.whl
Algorithm Hash digest
SHA256 1781170fe66a183945fd2d078e41db26fbe097af7df90994fbdedd5098097a43
MD5 ba684df8103405c3d3afdabe0252696c
BLAKE2b-256 cb12bac6f283f8c15ffe9873fb85e067c100da118a663398023d76a5e816f523

See more details on using hashes here.

File details

Details for the file webshart-0.5.0-cp312-abi3-manylinux_2_35_x86_64.whl.

File metadata

File hashes

Hashes for webshart-0.5.0-cp312-abi3-manylinux_2_35_x86_64.whl
Algorithm Hash digest
SHA256 0f10b6bee9784aafdfa2bf1b0f27c8d9dd7b6cfea756bb934786eee522109e52
MD5 e7eb8da454e2a340c1d46647a81fe2e1
BLAKE2b-256 267206848383f6354e0d4fdff1dc1f092b19d66a53ebc533150761c45fdfbdbe

See more details on using hashes here.

File details

Details for the file webshart-0.5.0-cp312-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for webshart-0.5.0-cp312-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 aabf7e46feb7282b91f9c7ae2b63ce1a6f1d0403683db2a6e089aeaddf45b530
MD5 0a3a14b4be2e74345ecff04b2cfdafaa
BLAKE2b-256 76fb1e443b1c90bde9a585e65155ccff84e369d1a471082274aaecc7737f9dd5

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page