Skip to main content

H5Record

Large dataset ( > 100G, <= 1T) storage format for Pytorch (wip)

Why?

  • Writing large dataset is still a wild west in pytorch. Approaches seen in the wild include:

    • large directory with lots of small files : slow IO when complex file is fetched, deserialized frequently
    • database approach : depend on what kind of database engine used, usually multi-process read is not supported
    • the above method scale non linear in terms of data - storage size
  • TFRecord solved the above problems well ( multiprocess fetch, (de)compression ), fast serialization ( protobuf )

  • However TFRecord port does not support data size evaluation (used frequently by Dataloader ), no index level access available ( important for data evaluation or verification )

H5Record aim to tackle TFRecord problems by compressing the dataset into HDF5 file with an easy to use interface through predefined interfaces ( String, Image, Sequences, Integer).

Simple usage

  1. Sentence Similarity
from h5record import H5Record, Float, Sentence

schema = (
    String(name='sentence1'),
    String(name='sentence2'),
    Float(name='label')
)
data = [
    ['Sent 1.', 'Sent 2', 0.1],
    ['Sent 3', 'Sent 4', 0.2],
]

def pair_iter():
    for row in data:
        yield {
            'sentence1': row[0],
            'sentence2': row[1]
        }

dataset = H5Dataset(schema, './question_pair.h5', pair_iter())
for idx in range(len(dataset)):
    print(dataset[idx])

Note

Due to in progress development, this package should be use in care in storage with FAT, FAT-32 format

Comparison between different compression algorithm

No chunking is used

Compression Type File size Read speed row/second
no compression 2.0G 2084.55 it/s
lzf 1.7G 1496.14 it/s
gzip 1.1G 843.78 it/s

benchmarked in i7-9700, 1TB NVMe SSD

TODO

  • Test combinations of different data modalities

  • Do more tuning and experiments on different driver settings

  • Performance benchmark

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

h5record-1.0.1.tar.gz (6.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

h5record-1.0.1-py3-none-any.whl (6.0 kB view details)

Uploaded Python 3

File details

Details for the file h5record-1.0.1.tar.gz.

File metadata

  • Download URL: h5record-1.0.1.tar.gz
  • Upload date:
  • Size: 6.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.4.1 importlib_metadata/3.7.0 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.60.0 CPython/3.6.5

File hashes

Hashes for h5record-1.0.1.tar.gz
Algorithm Hash digest
SHA256 0cf9d53ea13c18896e4cf23d553018e3b49331cfbcd5eaf32999d8bbc5c0e95e
MD5 d9186a969f9b8d69fa31a977827dcd4b
BLAKE2b-256 880850bd684c2ab3b47f5b8104d34805513c7447247e4b637adf2ad6340a017b

See more details on using hashes here.

File details

Details for the file h5record-1.0.1-py3-none-any.whl.

File metadata

  • Download URL: h5record-1.0.1-py3-none-any.whl
  • Upload date:
  • Size: 6.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.4.1 importlib_metadata/3.7.0 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.60.0 CPython/3.6.5

File hashes

Hashes for h5record-1.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 ed06c668c5cf1b1baaf582304238aa4a55070dd101784c39b0cb291c7c61e08d
MD5 34b9307c67cab4158d4e6a64f0142ce2
BLAKE2b-256 2327c7132d69370373eca9685c1ba2ab0dee57e732ac7867fa7231082a66ef0c

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page