Omni-IO
Efficient Python library for reading and writing multimedia data (audio, video, text) from binary archive blobs with support for both local and remote HTTP range requests.
Features
-
Multi-format support: Audio (FLAC, WAV, WebM/Opus), Video (MP4), Text (zstandard compressed)
-
Local and remote access: Seamlessly read from local files or remote URLs using HTTP range requests
-
Efficient storage: Binary blob archives with PyArrow/Parquet metadata indexing
-
Frame-level slicing: Extract specific time ranges from audio/video without loading entire files
-
Parallel processing: Multi-process append operations for fast archive creation
-
Streaming operations: Memory-efficient handling of large multimedia files
-
Kaldi ark/scp compatibility via
omniio.kaldi, an MIT-licensed drop-in forkaldiio
Why Omni-IO?
Most multimedia datasets outgrow naive storage approaches quickly. Omni-IO is designed for the scale and access patterns that matter in practice.
- Raw files on disk create serious filesystem overhead at scale — inode exhaustion, slow directory scans, and poor I/O throughput. Omni-IO packs everything into large .bin files, enabling fast sequential I/O and efficient bulk transfers.
- WebDataset eliminates the small-files problem but sacrifices random access. Omni-IO stores byte offsets in Parquet, so any item can be fetched in O(1) with a single range read — filter by any metadata column and shuffle freely.
- HuggingFace Datasets / Parquet blobs force audio and video into columnar formats they weren't designed for, inflating storage and defeating compression. Omni-IO keeps data in its native format (FLAC, WebM, zstd) and reserves Parquet for lightweight metadata only.
- HDF5 binary blobs do not expose the byte-range access needed for frame-level seeking, making it inefficient for partial reads and remote access.
- Numpy dumps store uncompressed PCM, ballooning storage 10–15×. Omni-IO decodes on demand from compressed formats, keeping archives compact while retaining full metadata.
- Lhotse manages where files are, but doesn't consolidate how they are stored — you still end up with individual files or WebDataset.
- Remote support: the same Parquet metadata file works for local and remote access. Swap a local .bin path for an HTTPS URL and the API is identical — HTTP range requests fetch only the bytes needed per sample.
Installation
git clone https://github.com/wavlab-speech/omniio.git
cd omniio
pip install -e .
Quick Start
Reading from Archives
Audio
from omniio.interface import audio_read
# Read audio from local or remote archive
result = audio_read(
archive_path="/path/to/archive.bin", # or "https://example.com/archive.bin"
start_offset=1024,
file_size=50000,
start_time=5.0, # optional: start at 5 seconds
end_time=10.0 # optional: end at 10 seconds
)
print(f"Sample rate: {result.sample_rate}")
print(f"Audio shape: {result.array.shape}") # (frames, channels)
Video
from omniio.video.read import video_read_local
# Read video with frame-based slicing
result = video_read_local(
archive_path="/path/to/archive.bin",
start_offset=2048,
file_size=1000000,
start_frame=100,
end_frame=200
)
print(f"FPS: {result.fps}")
print(f"Video shape: {result.video_array.shape}") # (frames, height, width, 3)
print(f"Audio shape: {result.audio_array.shape}") # (samples, channels)
Text
from omniio.text.read import text_read_local
# Read compressed text
result = text_read_local(
archive_path="/path/to/archive.bin",
start_offset=512,
file_size=2048
)
print(result.text)
Writing to Archives
Creating an Archive
from omniio.blob.blob import Blob
# Initialize archive
blob = Blob(
archive_dir="./my_archive",
modality="audio",
max_bin_size=320 * 1024 * 1024 # 320MB per bin file
)
# Append audio files in parallel
blob.append(
items=["audio1.wav", "audio2.flac", "audio3.mp3"],
ids=["sample_001", "sample_002", "sample_003"],
num_workers=4,
target_format="flac",
target_bit_depth=16
)
# View archive statistics
blob.summary()
Audio Format Conversion
from omniio.audio.write import audio_write
# Convert audio to different format
raw_bytes, metadata = audio_write(
audio_path="input.wav",
item_id="converted_audio",
target_format="flac", # 'flac', 'wav', 'webm'
target_bit_depth=24
)
print(f"Channels: {metadata['channels']}")
print(f"Sample rate: {metadata['sample_rate']}")
print(f"Compressed size: {len(raw_bytes)} bytes")
Text Compression
from omniio.text.write import text_write
# Compress text data
raw_bytes, metadata = text_write(
path_or_string="document.txt",
item_id="doc_001",
is_path=True,
compression_level=3
)
print(f"Original size: {metadata['original_size']} bytes")
print(f"Compressed size: {metadata['compressed_size']} bytes")
Kaldi ark/scp Compatibility
omniio.kaldi reads and writes Kaldi ark/scp archives, so a project that
only needs Kaldi I/O can drop its kaldiio dependency:
from omniio import kaldi as kaldiio # same names, same signatures
with kaldiio.ReadHelper("scp:feats.scp") as reader:
for utt_id, feats in reader:
...
with kaldiio.WriteHelper("ark,scp:feats.ark,feats.scp") as writer:
writer["utt1"] = feats # float32/float64 matrix or vector
kaldiio.save_ark("wav.ark", {"utt1": (16000, wave)}, scp="wav.scp")
array = kaldiio.load_mat("feats.ark:1234") # random access via an scp entry
Exported: ReadHelper, WriteHelper, load_ark, load_scp,
load_scp_sequential, load_wav_scp, load_mat, load_segments, save_ark,
save_mat, open_like_kaldi, parse_specifier, parse_rspecifier,
parse_wspecifier, LazyLoader, SegmentedLoader, ReadError.
load_scp and load_scp_sequential take separator= for an scp with its own
delimiter and segments= to key the result by a segments file instead of by
recording; load_mat takes fd_dict=, a caller-owned handle cache worth using
when reading many entries out of a few archives.
The code lives in omniio/tools/kaldi/, but omniio.kaldi is the supported
import path and import omniio.kaldi, from omniio import kaldi and
from omniio.kaldi import ReadHelper all work.
Supported on-disk formats:
| Form | Notes |
|---|---|
FM/DM, FV/DV |
float32/float64 matrices and vectors |
CM/CM2/CM3 |
compressed matrices, all seven compression_method values |
std::vector<int32> |
alignments |
WaveHolder (bare RIFF) |
decoded at its native PCM width |
AUDIO-framed blobs |
the extended archive layout, any container soundfile can decode |
text (ark,t:) |
matrices and vectors |
Also supported: segments files, scp random access with a bounded file
descriptor cache (max_cache_fd), and Kaldi extended filenames including pipes
(sox ... |, | gzip -c > x.gz), - for stdin/stdout, and .gz.
Every public name kaldiio exports is present. The remaining differences are
additive optional arguments (endian= on the readers, write_kwargs= on
WriteHelper, return_position= on load_ark) and the name of the first
parameter on four functions, which matters only if you pass it by keyword:
kaldiio |
omniio.kaldi |
|
|---|---|---|
ReadHelper |
wspecifier |
rspecifier — it reads, so that is what it takes |
load_ark |
fname |
file_or_fd — an open file is accepted too |
load_mat |
ark_name |
name |
save_mat |
fname |
path |
Archives written here are byte-identical to Kaldi's own. Two deliberate
divergences from kaldiio are documented in omniio/kaldi/compression.py:
kSpeechFeature compression of matrices with fewer than five rows follows
Kaldi rather than kaldiio's wrapping arithmetic, and the fixed-range methods
kOneByteUnsignedInteger/kOneByteZeroOne clip out-of-range input instead of
letting it wrap.
This code is written against the format description in Kaldi itself
(Apache-2.0) and carries omniio's MIT license; kaldiio is not a dependency
and none of its code is used.
Archive Structure
Archives are organized as follows:
archive_dir/
├── blob_0.bin # Binary data (first chunk)
├── blob_1.bin # Binary data (second chunk, if > max_bin_size)
└── metadata.parquet # PyArrow table with byte offsets and metadata
The metadata table contains:
id: Unique identifier for each entrystart_byte: Byte offset where entry beginsend_byte: Byte offset where entry endsbin_index: Which bin file contains the entry- Format-specific metadata (sample_rate, channels, dimensions, etc.)
Data Formats
Audio
- Input formats: FLAC, WAV, OGG, WebM/Opus
- Output shape:
(frames, channels)asfloat32 - Supported bit depths: 8, 16, 24, 32 (PCM formats only)
Video
- Input formats: MP4 with H.264/H.265 video and AAC/Opus audio
- Video output shape:
(frames, height, width, 3)asuint8RGB24 - Audio output shape:
(samples, channels)asfloat32
Text
- Compression: Zstandard (levels 1-22)
- Encoding: UTF-8
Requirements
- Python >= 3.8
- numpy
- av (PyAV)
- soundfile
- requests
- zstandard
- pyarrow
License
MIT License
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file omniio-0.1.1.tar.gz.
File metadata
- Download URL: omniio-0.1.1.tar.gz
- Upload date:
- Size: 78.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cb796e720300f93e507b9d7b7cf4d880d870d9ec985cb4450aa255e550a9a2cd
|
|
| MD5 |
42715f19c986fd3e14bc2dc80723b183
|
|
| BLAKE2b-256 |
c8336e3b7baabb88bbf8766d650ff26c3f89374213a3e92e2d424146afeda83f
|
Provenance
The following attestation bundles were made for omniio-0.1.1.tar.gz:
Publisher:
publish.yml on wavlab-speech/omniio
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
omniio-0.1.1.tar.gz -
Subject digest:
cb796e720300f93e507b9d7b7cf4d880d870d9ec985cb4450aa255e550a9a2cd - Sigstore transparency entry: 2715025186
- Sigstore integration time:
-
Permalink:
wavlab-speech/omniio@00924a70b7cb6379d654bdced169791553811fb4 -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/wavlab-speech
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@00924a70b7cb6379d654bdced169791553811fb4 -
Trigger Event:
release
-
Statement type:
File details
Details for the file omniio-0.1.1-py3-none-any.whl.
File metadata
- Download URL: omniio-0.1.1-py3-none-any.whl
- Upload date:
- Size: 50.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f451ef0b61df5288ac9b646fac8dc282de41c487854d608e88d19ef25e738b87
|
|
| MD5 |
b8e24932ae2f8e6031ac2065cf23404b
|
|
| BLAKE2b-256 |
a2d574ea70e27b4df96c373f22e0c809aa5f36cb0ef0368de7f7dfef383838b4
|
Provenance
The following attestation bundles were made for omniio-0.1.1-py3-none-any.whl:
Publisher:
publish.yml on wavlab-speech/omniio
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
omniio-0.1.1-py3-none-any.whl -
Subject digest:
f451ef0b61df5288ac9b646fac8dc282de41c487854d608e88d19ef25e738b87 - Sigstore transparency entry: 2715025218
- Sigstore integration time:
-
Permalink:
wavlab-speech/omniio@00924a70b7cb6379d654bdced169791553811fb4 -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/wavlab-speech
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@00924a70b7cb6379d654bdced169791553811fb4 -
Trigger Event:
release
-
Statement type: