Asynchronous gzip file reading and writing for Python asyncio — aiofiles-style API, streaming text/binary modes, decompression-bomb guards, optional zlib-ng acceleration.
Project description
aiogzip ⚡️
An asynchronous API modeled after Python's gzip module for reading and
writing gzip-compressed files.
It can substantially outperform sequential gzip when async work overlaps or
optional zlib-ng accelerates bulk decompression; direct single-file line
iteration remains faster with synchronous gzip. See
Performance and optional acceleration.
Installation
pip install aiogzip
Python 3.8 through 3.14 are supported by the 1.x release line.
Text-mode quickstart
File methods are asynchronous, and line iteration uses async for:
import aiogzip
async with aiogzip.open("events.jsonl.gz", "rt") as f:
async for line in f:
print(line)
Binary-mode quickstart
import aiogzip
async with aiogzip.open("payload.bin.gz", "wb") as f:
await f.write(b"Hello, async world!")
async with aiogzip.open("payload.bin.gz", "rb") as f:
payload = await f.read()
For small files that comfortably fit in memory, use the binary whole-file helpers:
import aiogzip
data = await aiogzip.read("payload.bin.gz")
await aiogzip.write("copy.bin.gz", data)
read() and write() load the entire decompressed or uncompressed payload
into memory. Use open() to stream large files. The existing
AsyncGzipFile() factory remains fully supported for compatibility.
For arbitrary asynchronous byte sources, compress or decompress without adapting the source to a file object:
async for data in aiogzip.decompress_chunks(compressed_source()):
await consume(data)
async for data in aiogzip.compress_chunks(raw_source(), mtime=0):
await send(data)
Complete decompression integrity validation occurs only if the iterator is consumed to the end. Compression output is incomplete if its source fails or the consumer exits early. See the async-iterable streaming guide for backpressure, limits, cancellation, metadata, and lifecycle behavior.
See the recipes for JSON Lines, untrusted input, reproducible output, append mode, seeking, cancellation recovery, and external async streams.
Why use aiogzip?
- Async file I/O built on
asyncioandaiofiles, so independent streams can overlap I/O waits. - Binary and text modes with distinct, typed concrete classes.
- Async
read,write,readline,seek,tell,peek,readinto, and line iteration. - Interoperable gzip output, concatenated-member reads, and append support.
- Configurable gzip metadata for reproducible archives.
- Bounded decompression and rewind-cache controls for untrusted or non-seekable input.
- Pull-driven compression and decompression for
AsyncIterable[bytes]sources. - Optional zlib-ng acceleration without a required runtime dependency.
- Verified tarfile-style access patterns and
aiocsvworkflows.
Migrating from gzip
The API follows gzip, but file operations must be awaited:
| Standard library | aiogzip |
|---|---|
gzip.open(path, "rt") |
aiogzip.open(path, "rt") |
f.read() |
await f.read() |
f.readline() |
await f.readline() |
for line in f |
async for line in f |
f.seek(offset) |
await f.seek(offset) |
f.close() |
await f.close() |
Synchronous code:
import gzip
with gzip.open("events.jsonl.gz", "rt") as f:
for line in f:
process(line)
becomes asynchronous code:
import aiogzip
async with aiogzip.open("events.jsonl.gz", "rt") as f:
async for line in f:
await process(line)
Important differences and caveats:
aiogzipdefaults tocompresslevel=6;gzip.open()defaults to9. Passcompresslevel=9when that parity matters.- Paths (including
pathlib.Path) are accepted directly. Supported external asynchronous sources and destinations are passed withfilename=Noneandfileobj=...; theirread()orwrite()methods must be async. - Append modes (
"ab"and"at") create a new gzip member. Both libraries transparently read concatenated members as one decompressed stream. - An open file object is stateful and is not safe for simultaneous use by multiple tasks. Use one handle per task or serialize access.
- Gzip has no random-access index. Backward seeks rewind and replay decompression, so mixed-direction access can be O(n).
- Cancelling an executor-backed decompression can leave that reader unusable; close it and open a new handle before continuing.
Compatibility and operational behavior
aiogzip reads and writes standard gzip streams and supports text and binary
modes, tarfile-style reads, aiocsv, append mode, and concatenated members.
It is an asynchronous API modeled after gzip, not a synchronous drop-in
replacement.
- Lifecycle: Prefer
async with. When that is impractical, callawait f.open()and pair it withawait f.close()infinally. - Seeking: Backward seeks replay decompression from the start. Forward
access is fastest. Text
tell()may return a handle-bound opaque cookie when decoder state is buffered; do not persist that cookie across reopens. - Non-seekable sources: Up to 128 MiB of compressed input is cached by
default for replay. Tune
max_rewind_cache_size, or passNonefor an unbounded cache. - Untrusted input:
max_decompressed_sizecaps cumulative decompressed output for a read pass. Overflow raisesOSErrorwithout first materializing the complete expansion. - Task safety: Do not operate on one open handle concurrently from several tasks; internal codec and buffer state is mutable and intentionally unlocked.
- Cancellation: If cancellation occurs during executor-backed
decompression, later reads and seeks raise
OSError. Close and reopen the reader. A similarly cancelled compression makes that output member unusable; discard it and start a new writer. - Append mode: Each append creates another member instead of extending the existing deflate stream. Standards-compliant readers concatenate the members.
- Large writes: Gzip's 32-bit
ISIZEwraps after 4 GiB, as it does ingzip.open(). Passstrict_size=Trueto reject a member that would cross that limit. - Compression metadata:
mtimeand the embedded original filename affect output bytes. Set both explicitly when reproducibility across paths matters.
Performance and optional acceleration
aiogzip's performance advantage comes from async concurrency and optional codec acceleration. Comparisons use identical compressed fixtures for reads, compression level 6 for both writers, and median timings from repeated runs.
On a representative Python 3.12 Linux run, the direct I/O cases used 8 MiB inputs and the concurrency case used ten 1 MiB files:
- overlapping ten files with simulated 10 ms latency was about 6.6-7.3x
faster than processing them sequentially with
gzip; - optional zlib-ng made a highly compressible bulk
read(-1)about 5.3x faster thangzip, and stdlib zlib finished slightly ahead (~1.08x); - an LF-only universal-newline fast path made the representative zlib-ng bulk
text read about 1.8x faster than
gzip(the stdlib engine narrowed to about 1.2x slower); - bounded
readlines()batches brought an 8 MiB JSONL read-and-parse workload slightly ahead ofgzipon both engines (~1.03-1.05x faster), about 9-14% faster than direct iteration; - equal-level bulk text writes were at parity; and
- direct single-file JSONL iteration remained about 1.4-1.8x slower than
gzipbecause each line crosses an async-iterator boundary — batch withreadlines(hint)when that path is hot.
The concurrency result measures overlapped waiting, not a faster deflate
codec, and benchmark ratios vary by hardware, storage, Python version, and
data. Large codec calls are offloaded to the default executor so independent
tasks can keep making progress. Line splitting, readlines(), and
writelines() use bounded batching to reduce aiogzip's own coroutine overhead.
For large UTF-8 JSON Lines files with \n terminators, the measured fast path
uses newline="\n" and chunk_size=512 * 1024. Tune memory and throughput for
your workload rather than assuming one chunk size fits every application.
When CPU-bound per-line processing permits it, repeated
await f.readlines(hint) calls can process bounded groups of complete lines
with fewer async transitions than async for; the hint is an approximate
decoded-character target, not a hard memory limit.
Install the optional codec with:
pip install "aiogzip[fast]"
When installed, zlib-ng is selected automatically for decompression. Its gain
depends on the input and access pattern: it helps decompression-heavy bulk
reads far more than per-line Python iteration. Compression remains on stdlib
zlib so installation alone does not change gzip bytes; pass
fast_compress=True per writer to opt in. Set
AIOGZIP_ENGINE=stdlib to force stdlib behavior. Inspect the default selections
for a diagnostic report:
import aiogzip
print(aiogzip.engine_info())
The engine names are informational, not a stable machine-readable interface. See the performance guide for benchmarks and tuning guidance.
Development and contributing
The 1.x line is the last to support Python 3.8 and 3.9; aiogzip 2.0 will require Python 3.11+. Older interpreters will continue to resolve the latest compatible 1.x release from PyPI.
See the contributing guide for setup, tests, linting, typing, documentation, and benchmark workflows.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file aiogzip-1.10.2.tar.gz.
File metadata
- Download URL: aiogzip-1.10.2.tar.gz
- Upload date:
- Size: 182.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
881abbd59bc7208161118cf16cf54af61ec16fe544ebe8018b9a15c4bd5655bf
|
|
| MD5 |
5a89dfdb1337b546ad999f27bd41bc0e
|
|
| BLAKE2b-256 |
327bb49c14da7f7c99ec5fd2a2093a3fc53077257f547e25e0c91d844f34ea52
|
Provenance
The following attestation bundles were made for aiogzip-1.10.2.tar.gz:
Publisher:
publish.yml on geoff-davis/aiogzip
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
aiogzip-1.10.2.tar.gz -
Subject digest:
881abbd59bc7208161118cf16cf54af61ec16fe544ebe8018b9a15c4bd5655bf - Sigstore transparency entry: 2187271369
- Sigstore integration time:
-
Permalink:
geoff-davis/aiogzip@ab7a23e15bb7835bbef90bdefc80ee191485928d -
Branch / Tag:
refs/tags/v1.10.2 - Owner: https://github.com/geoff-davis
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@ab7a23e15bb7835bbef90bdefc80ee191485928d -
Trigger Event:
push
-
Statement type:
File details
Details for the file aiogzip-1.10.2-py3-none-any.whl.
File metadata
- Download URL: aiogzip-1.10.2-py3-none-any.whl
- Upload date:
- Size: 51.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
07c6bce6bf2e89786dd3067ab041df5da58af6a9d7e6743b2b2b49cbc52e7c4c
|
|
| MD5 |
12b44e9a89184cde8edd62a8bde4b6fa
|
|
| BLAKE2b-256 |
8004c2403043a2c6a1c8e198b897ea9213d8a32441d720ea770e3c319e0c50e9
|
Provenance
The following attestation bundles were made for aiogzip-1.10.2-py3-none-any.whl:
Publisher:
publish.yml on geoff-davis/aiogzip
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
aiogzip-1.10.2-py3-none-any.whl -
Subject digest:
07c6bce6bf2e89786dd3067ab041df5da58af6a9d7e6743b2b2b49cbc52e7c4c - Sigstore transparency entry: 2187271494
- Sigstore integration time:
-
Permalink:
geoff-davis/aiogzip@ab7a23e15bb7835bbef90bdefc80ee191485928d -
Branch / Tag:
refs/tags/v1.10.2 - Owner: https://github.com/geoff-davis
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@ab7a23e15bb7835bbef90bdefc80ee191485928d -
Trigger Event:
push
-
Statement type: