Skip to main content

H5Record

Large dataset ( > 100G, <= 1T) storage format for Pytorch (wip)

Why?

  • Writing large dataset is still a wild west in pytorch. Approaches seen in the wild include:

    • large directory with lots of small files : slow IO when complex file is fetched, deserialized frequently
    • database approach : depend on what kind of database engine used, usually multi-process read is not supported
    • the above method scale non linear in terms of data - storage size
  • TFRecord solved the above problems well ( multiprocess fetch, (de)compression ), fast serialization ( protobuf )

  • However TFRecord port does not support data size evaluation (used frequently by Dataloader ), no index level access available ( important for data evaluation or verification )

H5Record aim to tackle TFRecord problems by compressing the dataset into HDF5 file with an easy to use interface through predefined interfaces ( String, Image, Sequences, Integer).

Simple usage

  1. Sentence Similarity
from h5record import H5Dataset, Float, String

schema = (
    String(name='sentence1'),
    String(name='sentence2'),
    Float(name='label')
)
data = [
    ['Sent 1.', 'Sent 2', 0.1],
    ['Sent 3', 'Sent 4', 0.2],
]

def pair_iter():
    for row in data:
        yield {
            'sentence1': row[0],
            'sentence2': row[1],
            'label': row[2]
        }

dataset = H5Dataset(schema, './question_pair.h5', pair_iter())
for idx in range(len(dataset)):
    print(dataset[idx])

Note

Due to in progress development, this package should be use in care in storage with FAT, FAT-32 format

Comparison between different compression algorithm

No chunking is used

Compression Type File size Read speed row/second
no compression 2.0G 2084.55 it/s
lzf 1.7G 1496.14 it/s
gzip 1.1G 843.78 it/s

benchmarked in i7-9700, 1TB NVMe SSD

TODO

  • Test combinations of different data modalities

  • Do more tuning and experiments on different driver settings

  • Performance benchmark

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

h5record-1.0.2.tar.gz (7.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

h5record-1.0.2-py3-none-any.whl (6.0 kB view details)

Uploaded Python 3

File details

Details for the file h5record-1.0.2.tar.gz.

File metadata

  • Download URL: h5record-1.0.2.tar.gz
  • Upload date:
  • Size: 7.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.4.1 importlib_metadata/3.7.0 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.60.0 CPython/3.6.5

File hashes

Hashes for h5record-1.0.2.tar.gz
Algorithm Hash digest
SHA256 b6227826fc5246440228278ea078879a8ec111917ce5d6c16f4c8bfa9816f264
MD5 8a11d16118546faa272287a9dbe93def
BLAKE2b-256 603818420f91d04910a2586514c3c77f30bdf439451fd5e8ddb8ea6fbf6059c7

See more details on using hashes here.

File details

Details for the file h5record-1.0.2-py3-none-any.whl.

File metadata

  • Download URL: h5record-1.0.2-py3-none-any.whl
  • Upload date:
  • Size: 6.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.4.1 importlib_metadata/3.7.0 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.60.0 CPython/3.6.5

File hashes

Hashes for h5record-1.0.2-py3-none-any.whl
Algorithm Hash digest
SHA256 83d9d8fd447b103e1263423c829227b9cc40d63db490747d5e88ab18faafa532
MD5 c1bec48947eafaa7235cd538a70c211a
BLAKE2b-256 6e35ea67af81d8df6a47b4340e3a2b0540b9d32caa9eb05f56014fb6687c00d6

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page