Skip to main content
image

Fast dataloader and conversion utility for webdataset tar shards. Rust core with Python bindings.

Built for streaming large video and image datasets, but handles any byte data.

Install

pip install webshart

What is this?

Webshart is a fast reader for webdataset tar files with separate JSON index files. This format enables random access to any file in the dataset without downloading the entire archive.

The indexed format provides massive performance benefits:

  • Random access: Jump to any file instantly
  • Selective downloads: Only fetch the files you need
  • True parallelism: Read from multiple shards simultaneously
  • Cloud-optimized: Works efficiently with HTTP range requests
  • Aspect bucketing: Optionally include image geometry hints width, height and aspect for the ability to bucket images by shape
  • Logical sample APIs: Treat image.ext + image.json pairs as one sample while still allowing raw file access
  • Caption metadata: Store captions in shard metadata under the plural captions key as either a string or a list of strings
  • Custom DataLoader: Includes state dict methods on the DataLoader so that you can resume training deterministically
  • Rate-limit friendly: Local caching allows high-frequency random seeking without encountering storage provider rate limits
  • Instant start-up with pre-sorted aspect buckets

Growing ecosystem: While not all datasets use this format yet, you can easily create indices for any tar-based dataset (see below).

Quick Start

import webshart

# Find your dataset
dataset = webshart.discover_dataset(
    source="laion/conceptual-captions-12m-webdataset",
    # we're able to upload metadata separately so that we reduce load on huggingface infra.
    metadata="webshart/conceptual-captions-12m-webdataset-metadata",
)
print(f"Found {dataset.num_shards} shards")

loader = webshart.TarDataLoader(dataset)

# File-oriented access is still available.
files = dataset.list_files_in_shard(0)

# Sample-oriented access skips paired JSON sidecars.
samples = dataset.list_samples_in_shard(0)
entry = loader.load_sample(0, 0)
print(entry.path, entry.captions, entry.json_metadata)

max_file_size is a visibility limit for loader APIs. Files larger than the configured limit are omitted from iteration, batches, direct sample loading, and aspect buckets instead of being returned with empty data. Direct load_sample() calls return None for an oversized sample. The loader's list_samples_in_shard() returns dictionaries containing sample_idx and filename, so filtered listings retain the stable index required by load_sample().

Common Patterns

For real-world, working examples:

Creating Indices for / Converting Existing Datasets

Any tar-based webdataset can benefit from indexing! Webshart includes tools to generate indices:

A command-line tool that auto-discovers tars to process:

% webshart extract-metadata \
    --source laion/conceptual-captions-12m-webdataset \
    --destination laion_output/ \
    --checkpoint-dir ./laion_output/checkpoints \
    --max-workers 2 \
    --include-image-geometry

Or, if you prefer/require direct-integration to an existing Python application, use the API

Uploading Indices to HuggingFace

Once you've generated indices, share them with the community:

# Upload all JSON files to your dataset
huggingface-cli upload --repo-type=dataset \
    username/dataset-name \
    ./indices/ \
    --include "*.json" \
    --path-in-repo "indices/"

Or if you want to contribute to an existing dataset you don't own:

  1. Create a community dataset with indices: username/original-dataset-indices
  2. Upload the JSON files there
  3. Open a discussion on the original dataset suggesting they add the indices

Creating New Indexed Datasets

If you're creating a new dataset, generate indices during creation:

{
  "files": {
    "image_0001.webp": {"offset": 512, "length": 102400},
    "image_0002.webp": {"offset": 102912, "length": 98304},
    ...
  }
}

The JSON index should have the same name as the tar file (e.g., shard_0000.tarshard_0000.json).

Image + JSON Sidecar Samples

Webshart supports webdataset shards that store each sample as an image-like payload plus a paired JSON sidecar:

sample_0001.webp
sample_0001.json
sample_0002.webp
sample_0002.json

When metadata is extracted or loaded, sidecars are attached to their paired sample entries:

{
  "files": {
    "sample_0001.webp": {
      "offset": 512,
      "length": 102400,
      "width": 1024,
      "height": 1024,
      "aspect": 1.0,
      "json_path": "sample_0001.json",
      "json_offset": 103424,
      "json_length": 128,
      "captions": "a product photo on a white background",
      "json_metadata": {
        "caption": "a product photo on a white background"
      }
    },
    "sample_0001.json": {
      "offset": 103424,
      "length": 128
    }
  }
}

Use file-oriented APIs when you want every archive member, including sidecars:

dataset.list_files_in_shard(0)

reader = dataset.open_shard(0)
raw_file_bytes = reader.read_file(0)

Use sample-oriented APIs when you want training samples:

dataset.list_samples_in_shard(0)
dataset.get_shard_sample_count(0)

reader = dataset.open_shard(0)
image_bytes = reader.read_sample(0)
json_bytes = reader.read_sample_json(0)

entry = loader.load_sample(0, 0)
print(entry.path)
print(entry.captions)
print(entry.json_data)

Captions are canonicalized to the plural captions metadata key. The value may be a single string, a list of strings, or absent.

webshart.write_captions_to_metadata(
    "shard_0000.json",
    {
        "sample_0001.webp": "a short caption",
        "sample_0002": ["caption one", "caption two"],
    },
)

The writer updates existing webshart metadata JSON in place, removes old singular caption keys from updated samples, and leaves paired .json sidecar entries untouched.

Aspect Bucketing Samples

list_shard_aspect_buckets() is file-oriented and buckets any indexed file that has width and height.

For training pipelines, prefer list_shard_sample_aspect_buckets():

loader = webshart.TarDataLoader(dataset)
buckets = loader.list_shard_sample_aspect_buckets(
    [0],
    key="geometry-tuple",
    target_pixel_area=1024**2,
)[0]["buckets"]

for bucket_key, entries in buckets.items():
    for item in entries:
        virtual_id = f"webshart://0/{item['sample_idx']}/{item['filename']}"
        image = loader.load_sample(0, item["sample_idx"])

This uses logical samples from metadata.sample_range() / get_sample_by_index() and excludes paired JSON sidecars before bucketing. Each bucket entry includes sample_idx, so callers can build stable IDs and load images directly with loader.load_sample(shard_idx, sample_idx).

Why is it fast?

Problem: Standard tar files require sequential reading. To get file #10,000, you must read through files #1-9,999 first.

Solution: The indexed format stores byte offsets and sample metadata in a separate JSON file, enabling:

  • HTTP range requests for any file
  • True random access over network
  • Parallel reads from multiple shards
  • Large scale, aspect-bucketed datasets
  • No wasted bandwidth

The Rust implementation provides:

  • Real parallelism (no Python GIL)
  • Zero-copy operations where possible
  • Efficient HTTP connection pooling
  • Optimized tokio async runtime
  • Optional local caching for metadata and shards
  • Fast aspect bucketing for image data

Datasets Using This Format

I discovered after creating this library that cheesechaser is the origin of the indexed tar format, which webshart has formalised and extended to include aspect bucketing support.

Requirements

  • Python 3.12+
  • Linux/macOS/Windows

Roadmap

  • image decoding is currently not handled by this library, but it will be added with zero-copy.
  • more informative API for caching and other Rust implementation details
  • multi-gpu/multi-node friendly dataloader

Projects using webshart

  • CaptionFlow uses this library to solve memory use and seek performance issues typical to webdatasets

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

webshart-0.5.1.tar.gz (104.3 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

webshart-0.5.1-cp312-abi3-win_amd64.whl (4.0 MB view details)

Uploaded CPython 3.12+Windows x86-64

webshart-0.5.1-cp312-abi3-manylinux_2_39_x86_64.whl (4.4 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.39+ x86-64

webshart-0.5.1-cp312-abi3-manylinux_2_39_aarch64.whl (4.2 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.39+ ARM64

webshart-0.5.1-cp312-abi3-manylinux_2_35_x86_64.whl (4.4 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.35+ x86-64

webshart-0.5.1-cp312-abi3-macosx_11_0_arm64.whl (4.4 MB view details)

Uploaded CPython 3.12+macOS 11.0+ ARM64

File details

Details for the file webshart-0.5.1.tar.gz.

File metadata

  • Download URL: webshart-0.5.1.tar.gz
  • Upload date:
  • Size: 104.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for webshart-0.5.1.tar.gz
Algorithm Hash digest
SHA256 2645ce3acf0779a69ac4cd18916af95247f3b60c0b180b937adb394e3bac144d
MD5 b2ffd638b6feb5f2d0bc92ba861159c4
BLAKE2b-256 ea053bb45590b78f97355c76ff46072930a4d568bcae5c4faa3ef443e0a845a9

See more details on using hashes here.

File details

Details for the file webshart-0.5.1-cp312-abi3-win_amd64.whl.

File metadata

  • Download URL: webshart-0.5.1-cp312-abi3-win_amd64.whl
  • Upload date:
  • Size: 4.0 MB
  • Tags: CPython 3.12+, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for webshart-0.5.1-cp312-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 6610f1c8815f5cb948d84eaa9fb9bfa71bcf6ffb77972ec4e006bd0691c60776
MD5 ae426b1ab18ed6445f022b8513eb1302
BLAKE2b-256 3424d1606dde6f3c10ec69beedc97516d7db8f4520f98eb81309cf8535772359

See more details on using hashes here.

File details

Details for the file webshart-0.5.1-cp312-abi3-manylinux_2_39_x86_64.whl.

File metadata

File hashes

Hashes for webshart-0.5.1-cp312-abi3-manylinux_2_39_x86_64.whl
Algorithm Hash digest
SHA256 0ee97731b9b52531625061f07d104c1627e3cf2aa454e535ecf8bc037942d733
MD5 2a5ff51763f379a7d1a653333035f218
BLAKE2b-256 3616df009bbca03b4a795343a00188e2eeef1d04987d9f2b771efd6487b98508

See more details on using hashes here.

File details

Details for the file webshart-0.5.1-cp312-abi3-manylinux_2_39_aarch64.whl.

File metadata

File hashes

Hashes for webshart-0.5.1-cp312-abi3-manylinux_2_39_aarch64.whl
Algorithm Hash digest
SHA256 20f2b064baa0555493e111ec593c51f26cce96816908e2482120ee68c6fc7940
MD5 f7dab32947147ce199b040d061c91b4c
BLAKE2b-256 689c3f60921bd3ab839ffdc57d16024475212d9a90bdd95dc07e97c5c16a4b77

See more details on using hashes here.

File details

Details for the file webshart-0.5.1-cp312-abi3-manylinux_2_35_x86_64.whl.

File metadata

File hashes

Hashes for webshart-0.5.1-cp312-abi3-manylinux_2_35_x86_64.whl
Algorithm Hash digest
SHA256 f52beb7dd114a5de4ccab04d4800787665a9dc63c41150fa13b871a1cc9ddb50
MD5 a7c64fea93bbb7d8ec805df01d51c22a
BLAKE2b-256 09e3700b96b5a1a24791f2208ea9c99b4b764147870ff0808a786f8646870ba4

See more details on using hashes here.

File details

Details for the file webshart-0.5.1-cp312-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for webshart-0.5.1-cp312-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 cee0d6840d2579d94fbfc002a262f7e176f32c0820c165e921676d51e8aa9820
MD5 1be764b0752ac8173424811106497f6f
BLAKE2b-256 f695a7fe80597e76faade10ef6f1bef6cca99563f55dd60c1bef038a93e418ca

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page