Skip to main content

numpy2tfrecord

Simple helper library to convert numpy data to tfrecord and build a tensorflow dataset.

Installation

$ git clone git@github.com:yonetaniryo/numpy2tfrecord.git
$ cd numpy2tfrecord
$ pip install .

or simply using pip:

$ pip install numpy2tfrecord

How to use

Convert a collection of numpy data to tfrecord

You can convert samples represented in the form of a dict to tf.train.Example and save them as a tfrecord.

import numpy as np
from numpy2tfrecord import Numpy2TFRecordConverter

with Numpy2TFRecordConverter("test.tfrecord") as converter:
    x = np.arange(100).reshape(10, 10).astype(np.float32)  # float array
    y = np.arange(100).reshape(10, 10).astype(np.int64)  # int array
    a = 5  # int
    b = 0.3  # float
    sample = {"x": x, "y": y, "a": a, "b": b}
    converter.convert_sample(sample)  # convert data sample

You can also convert a list of samples at once using convert_list.

with Numpy2TFRecordConverter("test.tfrecord") as converter:
    samples = [
        {
            "x": np.random.rand(64).astype(np.float32),
            "y": np.random.randint(0, 10),
        }
        for _ in range(32)
    ]  # list of 32 samples

    converter.convert_list(samples)

Or a batch of samples at once using convert_batch.

with Numpy2TFRecordConverter("test.tfrecord") as converter:
    samples = {
        "x": np.random.rand(32, 64).astype(np.float32),
        "y": np.random.randint(0, 10, size=32).astype(np.int64),
    }  # batch of 32 samples

    converter.convert_batch(samples)

So what are the advantages of Numpy2TFRecordConverter compared to tf.data.datset.from_tensor_slices? Simply put, when using tf.data.dataset.from_tensor_slices, all the samples that will be converted to a dataset must be in memory. On the other hand, you can use Numpy2TFRecordConverter to sequentially add samples to the tfrecord without having to read all of them into memory beforehand..

Build a tensorflow dataset from tfrecord

Samples once stored in the tfrecord can be streamed using tf.data.TFRecordDataset.

from numpy2tfrecord import build_dataset_from_tfrecord

dataset = build_dataset_from_tfrecord("test.tfrecord")

The dataset can then be used directly in the for-loop of machine learning.

for batch in dataset.as_numpy_iterator():
    x, y = batch.values()
    ...

Speeding up PyTorch data loading with numpy2tfrecord!

https://gist.github.com/yonetaniryo/c1780e58b841f30150c45233d3fe6d01

import os
import time

import numpy as np
from numpy2tfrecord import Numpy2TfrecordConverter, build_dataset_from_tfrecord
import torch
from torchvision import datasets, transforms

dataset = datasets.MNIST(".", download=True, transform=transforms.ToTensor())

# convert to tfrecord
with Numpy2TfrecordConverter("mnist.tfrecord") as converter:
    converter.convert_batch({"x": dataset.data.numpy().astype(np.int64), 
                        "y": dataset.targets.numpy().astype(np.int64)})

torch_loader = torch.utils.data.DataLoader(dataset, batch_size=32, pin_memory=True, num_workers=os.cpu_count())
tic = time.time()
for e in range(5):
    for batch in torch_loader:
        x, y = batch
elapsed = time.time() - tic
print(f"elapsed time with pytorch dataloader: {elapsed:0.2f} sec for 5 epochs")

tf_loader = build_dataset_from_tfrecord("mnist.tfrecord").batch(32).prefetch(1)
tic = time.time()
for e in range(5):
    for batch in tf_loader.as_numpy_iterator():
        x, y = batch.values()
elapsed = time.time() - tic
print(f"elapsed time with tf dataloader: {elapsed:0.2f} sec for 5 epochs")

⬇️

elapsed time with pytorch dataloader: 41.10 sec for 5 epochs
elapsed time with tf dataloader: 17.34 sec for 5 epochs

Release files for numpy2tfrecord 0.0.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for numpy2tfrecord 0.0.3
File Size Uploaded
numpy2tfrecord-0.0.3.tar.gz 5.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for numpy2tfrecord 0.0.3
File Interpreter ABI Platform
numpy2tfrecord-0.0.3-py3-none-any.whl Python 3 none any Details

Total release size: 10.5 kB

Release files / numpy2tfrecord-0.0.3.tar.gz

Download URL numpy2tfrecord-0.0.3.tar.gz
Size 5.1 kB
Tags Source
SHA-256 checksum
How to use checksums
fa44db6cc26677f3886ef1c5dc0bda13f3cf390247907388ee62acf12035f111
BLAKE2b-256 checksum
How to use checksums
4e0b919950e84385fa697966ef54683b3b4d981a206f4063c478517477263c67
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.10.8

Release files / numpy2tfrecord-0.0.3-py3-none-any.whl

Download URL numpy2tfrecord-0.0.3-py3-none-any.whl
Size 5.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e21e3507f92c3c5e90633fe8483d466543c16b754f062269575aa460f2090ed7
BLAKE2b-256 checksum
How to use checksums
8465265df14bfda999f279f34070b58a0f38df56cf2079206193082f29baf32d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.10.8

Release history Release notifications | RSS feed

This release

0.0.3 This release

2 release files

0.0.2

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page