Skip to main content
image

Fast dataloader and conversion utility for webdataset tar shards. Rust core with Python bindings.

Built for streaming large video and image datasets, but handles any byte data.

Install

pip install webshart

What is this?

Webshart is a fast reader for webdataset tar files with separate JSON index files. This format enables random access to any file in the dataset without downloading the entire archive.

The indexed format provides massive performance benefits:

  • Random access: Jump to any file instantly
  • Selective downloads: Only fetch the files you need
  • True parallelism: Read from multiple shards simultaneously
  • Cloud-optimized: Works efficiently with HTTP range requests
  • Aspect bucketing: Optionally include image geometry hints width, height and aspect for the ability to bucket images by shape
  • Logical sample APIs: Treat image.ext + image.json pairs as one sample while still allowing raw file access
  • Caption metadata: Store captions in shard metadata under the plural captions key as either a string or a list of strings
  • Custom DataLoader: Includes state dict methods on the DataLoader so that you can resume training deterministically
  • Rate-limit friendly: Local caching allows high-frequency random seeking without encountering storage provider rate limits
  • Instant start-up with pre-sorted aspect buckets

Growing ecosystem: While not all datasets use this format yet, you can easily create indices for any tar-based dataset (see below).

Quick Start

import webshart

# Find your dataset
dataset = webshart.discover_dataset(
    source="laion/conceptual-captions-12m-webdataset",
    # we're able to upload metadata separately so that we reduce load on huggingface infra.
    metadata="webshart/conceptual-captions-12m-webdataset-metadata",
)
print(f"Found {dataset.num_shards} shards")

loader = webshart.TarDataLoader(dataset)

# File-oriented access is still available.
files = dataset.list_files_in_shard(0)

# Sample-oriented access skips paired JSON sidecars.
samples = dataset.list_samples_in_shard(0)
entry = loader.load_sample(0, 0)
print(entry.path, entry.captions, entry.json_metadata)

max_file_size is a visibility limit for loader APIs. Files larger than the configured limit are omitted from iteration, batches, direct sample loading, and aspect buckets instead of being returned with empty data. Direct load_sample() calls return None for an oversized sample. The loader's list_samples_in_shard() returns dictionaries containing sample_idx and filename, so filtered listings retain the stable index required by load_sample().

Common Patterns

For real-world, working examples:

Creating Indices for / Converting Existing Datasets

Any tar-based webdataset can benefit from indexing! Webshart includes tools to generate indices:

A command-line tool that auto-discovers tars to process:

% webshart extract-metadata \
    --source laion/conceptual-captions-12m-webdataset \
    --destination laion_output/ \
    --checkpoint-dir ./laion_output/checkpoints \
    --max-workers 2 \
    --include-image-geometry

Or, if you prefer/require direct-integration to an existing Python application, use the API

Uploading Indices to HuggingFace

Once you've generated indices, share them with the community:

# Upload all JSON files to your dataset
huggingface-cli upload --repo-type=dataset \
    username/dataset-name \
    ./indices/ \
    --include "*.json" \
    --path-in-repo "indices/"

Or if you want to contribute to an existing dataset you don't own:

  1. Create a community dataset with indices: username/original-dataset-indices
  2. Upload the JSON files there
  3. Open a discussion on the original dataset suggesting they add the indices

Creating New Indexed Datasets

If you're creating a new dataset, generate indices during creation:

{
  "files": {
    "image_0001.webp": {"offset": 512, "length": 102400},
    "image_0002.webp": {"offset": 102912, "length": 98304},
    ...
  }
}

The JSON index should have the same name as the tar file (e.g., shard_0000.tarshard_0000.json).

Caption layouts and sidecars

Webshart recognizes both JSON metadata sidecars and plain-text caption sidecars:

sample_0001.webp
sample_0001.json
sample_0002.webp
sample_0002.txt

Paired .json and .txt members are excluded from logical sample indexes. You can inspect the layout from shard metadata without downloading tar members:

layout = dataset.probe_caption_layout(max_shards=16)
print(layout["layout"])  # embedded, json_sidecar, txt_sidecar, mixed, or none

When metadata is extracted or loaded, sidecars are attached to their paired sample entries:

{
  "files": {
    "sample_0001.webp": {
      "offset": 512,
      "length": 102400,
      "width": 1024,
      "height": 1024,
      "aspect": 1.0,
      "json_path": "sample_0001.json",
      "json_offset": 103424,
      "json_length": 128,
      "captions": "a product photo on a white background",
      "json_metadata": {
        "caption": "a product photo on a white background"
      }
    },
    "sample_0001.json": {
      "offset": 103424,
      "length": 128
    }
  }
}

Use file-oriented APIs when you want every archive member, including sidecars:

dataset.list_files_in_shard(0)

reader = dataset.open_shard(0)
raw_file_bytes = reader.read_file(0)

Use sample-oriented APIs when you want training samples:

dataset.list_samples_in_shard(0)
dataset.get_shard_sample_count(0)

reader = dataset.open_shard(0)
image_bytes = reader.read_sample(0)
json_bytes = reader.read_sample_json(0)

entry = loader.load_sample(0, 0)
print(entry.path)
print(entry.captions)
print(entry.json_data)

# Direct caption lookup also handles paired .txt sidecars.
caption = loader.load_caption(0, 0)

Captions are canonicalized to the plural captions metadata key. The value may be a single string, a list of strings, or absent.

webshart.write_captions_to_metadata(
    "shard_0000.json",
    {
        "sample_0001.webp": "a short caption",
        "sample_0002": ["caption one", "caption two"],
    },
)

The writer updates existing webshart metadata JSON in place, removes old singular caption keys from updated samples, and leaves paired .json sidecar entries untouched.

To avoid repeated .txt range reads, fold all sidecar captions into standard webshart metadata files. If metadata caching is enabled, omitting the destination persists the enriched indexes in webshart's cache:

dataset.enable_metadata_cache("cache/metadata", init_shard_count=0)
loader = webshart.TarDataLoader(dataset, load_file_data=False)
loader.coalesce_caption_metadata()

# Or create a portable export tree for copying or upload.
loader.coalesce_caption_metadata("caption-metadata")
webshart.upload_caption_metadata(
    "caption-metadata",
    "organization/dataset-metadata",
    hf_token="hf_...",
)

The CLI provides the same operation. --shard-cache-dir lets coalescing reuse full cached shards instead of issuing one range read per sidecar:

webshart optimize-captions \
  --source organization/dataset \
  --metadata organization/dataset-metadata \
  --destination caption-metadata \
  --shard-cache-dir cache/shards \
  --push-to-hub organization/dataset-metadata

Hub reads accept hf_token= and also honor HF_TOKEN. This includes gated datasets and separately hosted metadata. Local discovery recursively pairs tar and JSON indexes, preserving their relative subdirectories.

Aspect Bucketing Samples

list_shard_aspect_buckets() is file-oriented and buckets any indexed file that has width and height.

For training pipelines, prefer list_shard_sample_aspect_buckets():

loader = webshart.TarDataLoader(dataset)
buckets = loader.list_shard_sample_aspect_buckets(
    [0],
    key="geometry-tuple",
    target_pixel_area=1024**2,
)[0]["buckets"]

for bucket_key, entries in buckets.items():
    for item in entries:
        virtual_id = f"webshart://0/{item['sample_idx']}/{item['filename']}"
        image = loader.load_sample(0, item["sample_idx"])

This uses logical samples from metadata.sample_range() / get_sample_by_index() and excludes paired JSON sidecars before bucketing. Each bucket entry includes sample_idx, so callers can build stable IDs and load images directly with loader.load_sample(shard_idx, sample_idx).

Why is it fast?

Problem: Standard tar files require sequential reading. To get file #10,000, you must read through files #1-9,999 first.

Solution: The indexed format stores byte offsets and sample metadata in a separate JSON file, enabling:

  • HTTP range requests for any file
  • True random access over network
  • Parallel reads from multiple shards
  • Large scale, aspect-bucketed datasets
  • No wasted bandwidth

The Rust implementation provides:

  • Real parallelism (no Python GIL)
  • Zero-copy operations where possible
  • Efficient HTTP connection pooling
  • Optimized tokio async runtime
  • Optional local caching for metadata and shards
  • Fast aspect bucketing for image data

Datasets Using This Format

I discovered after creating this library that cheesechaser is the origin of the indexed tar format, which webshart has formalised and extended to include aspect bucketing support.

Requirements

  • Python 3.12+
  • Linux/macOS/Windows

Roadmap

  • image decoding is currently not handled by this library, but it will be added with zero-copy.
  • more informative API for caching and other Rust implementation details
  • multi-gpu/multi-node friendly dataloader

Projects using webshart

  • CaptionFlow uses this library to solve memory use and seek performance issues typical to webdatasets

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

webshart-0.5.2.tar.gz (111.6 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

webshart-0.5.2-cp312-abi3-win_amd64.whl (4.1 MB view details)

Uploaded CPython 3.12+Windows x86-64

webshart-0.5.2-cp312-abi3-manylinux_2_39_x86_64.whl (4.4 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.39+ x86-64

webshart-0.5.2-cp312-abi3-manylinux_2_39_aarch64.whl (4.2 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.39+ ARM64

webshart-0.5.2-cp312-abi3-manylinux_2_35_x86_64.whl (4.4 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.35+ x86-64

webshart-0.5.2-cp312-abi3-macosx_11_0_arm64.whl (4.5 MB view details)

Uploaded CPython 3.12+macOS 11.0+ ARM64

File details

Details for the file webshart-0.5.2.tar.gz.

File metadata

  • Download URL: webshart-0.5.2.tar.gz
  • Upload date:
  • Size: 111.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for webshart-0.5.2.tar.gz
Algorithm Hash digest
SHA256 fc8ed8f8080ad1fa98a49752d74082ee4b3ee0e2af4b336d97d99a4c7eeb56d5
MD5 831351f5e096bb5fa8b5913f3a3aefd1
BLAKE2b-256 024f76f6e393bf7cec4ff0d3a866a3261f8c793a16936ca4725ad49244e86528

See more details on using hashes here.

File details

Details for the file webshart-0.5.2-cp312-abi3-win_amd64.whl.

File metadata

  • Download URL: webshart-0.5.2-cp312-abi3-win_amd64.whl
  • Upload date:
  • Size: 4.1 MB
  • Tags: CPython 3.12+, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for webshart-0.5.2-cp312-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 dd87b72ceb16c3a0add56fe823bb22ce1e9b579233d6c099be120ca4e91a7281
MD5 3510c123f61ac03e70ad052286ef26e0
BLAKE2b-256 1737088efce0135294a49f967e07525bea8e7cdfa3d09fe6b53ed5f6fc5f209a

See more details on using hashes here.

File details

Details for the file webshart-0.5.2-cp312-abi3-manylinux_2_39_x86_64.whl.

File metadata

File hashes

Hashes for webshart-0.5.2-cp312-abi3-manylinux_2_39_x86_64.whl
Algorithm Hash digest
SHA256 53e22848533fe2e340e9c3dc0a8aff480dc95c0a222e9e450dac774da50b5dca
MD5 07bba3f50cfb790b10f96c3a9adf4714
BLAKE2b-256 c946b385be8039e962099d09c54477f7ef58731b2e69c7ddad5143e22761198c

See more details on using hashes here.

File details

Details for the file webshart-0.5.2-cp312-abi3-manylinux_2_39_aarch64.whl.

File metadata

File hashes

Hashes for webshart-0.5.2-cp312-abi3-manylinux_2_39_aarch64.whl
Algorithm Hash digest
SHA256 2f06c92f111dd04aa48c2b40e235256560950d36bc1fe9ce2f58edbc9e01f4e8
MD5 9bff73286a04491fed0a16ff89c1387b
BLAKE2b-256 80f47c321d05ecc4ad281448aa7855d0b16740f2846f37fa3c2dfd216fd0f494

See more details on using hashes here.

File details

Details for the file webshart-0.5.2-cp312-abi3-manylinux_2_35_x86_64.whl.

File metadata

File hashes

Hashes for webshart-0.5.2-cp312-abi3-manylinux_2_35_x86_64.whl
Algorithm Hash digest
SHA256 e6ccc7a2de6e33bb308b83e3d8bbf7910978cc0a22173c7f5f309b57cf66c7e9
MD5 e767484ec45cea1ed254ef346613685f
BLAKE2b-256 679c5bb98c4f50fe5aa1f5e9cd576ff46852ca25dc5ca01c10e16cf464276fe0

See more details on using hashes here.

File details

Details for the file webshart-0.5.2-cp312-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for webshart-0.5.2-cp312-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 eeaca4d7993fc65b468a0a6d8713f3f7b401f09779f7e2c8e6a0f0e9aa06a2c2
MD5 a4c40fc662773d025d9f8a6575f5b22c
BLAKE2b-256 33f860f80c9af6597b2c96893036972f313cad22e002de86d5d47edc653cacbb

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page