Skip to main content

jzpack

JSON-record compression using columnar encoding, MessagePack, and Zstandard.

Status: beta. The public API is intentionally small, while the JZPK binary format is being stabilized for long-term and cross-language use.

Installation

pip install jzpack

Quick Start

from jzpack import compress, decompress

data = [{"service": "api", "status": "ok", "latency": 42} for _ in range(10000)]

compressed = compress(data)
original = decompress(compressed)

Workloads and performance

jzpack stores record-shaped JSON and NDJSON data in a columnar binary archive. Size and speed depend on the data, runtime, and hardware; benchmark results do not predict every workload. See the benchmark guide for tested workloads, reproduction steps, and limitations.

API

from jzpack import (
    JZPackCompressor,
    StreamingCompressor,
    compress,
    decompress,
    iter_decompress,
    iter_decompress_recover,
    write_records,
)

# Simple API
compressed = compress(data, level=3, fast=False)
original = decompress(compressed)

# Class-based API
compressor = JZPackCompressor(compression_level=3, fast=False)
compressed = compressor.compress(data)
compressor.decompress(compressed)
compressor.compress_to_file(data, "out.jzpk")
compressor.decompress_from_file("out.jzpk")

# Incremental in-memory ingestion
stream = StreamingCompressor(compression_level=3, fast=False)
stream.add_record(record)
stream.add_batch(records)
stream.finalize()
stream.clear()

Chunked v3 reads

JZPK version 3 is the sole public wire format. compress emits a v3 container, and iter_decompress reads one chunk at a time from paths and binary streams without an unbounded read() call. Standalone v1/v2 payloads are unsupported; the v2 payload embedded inside a v3 chunk is an internal format detail.

from jzpack import ChunkError, ChunkRecords, iter_decompress, iter_decompress_recover

for record in iter_decompress("archive-v3.jzpk", max_chunks=1000):
    consume(record)

# Recovery is explicit: ChunkError means the output is incomplete, and no
# records from its sequence are yielded.
for event in iter_decompress_recover("archive-v3.jzpk"):
    if isinstance(event, ChunkRecords):
        consume_many(event.records)
    else:
        assert isinstance(event, ChunkError)
        report_corrupt_chunk(event.sequence, event.error)

Both iterator functions accept bytes-like input, paths, and binary file-like objects. Their optional limits are max_output_size, max_records, max_chunks, max_chunk_uncompressed_bytes, and max_chunk_payload_bytes. max_output_size is the aggregate uncompressed MessagePack body size. decompress also accepts valid v3 bytes as a list-returning adapter; its returned list is naturally not bounded-memory.

Bounded v3 writing

write_records accepts a built-in dictionary or iterable of them, emitting independent v3 chunks to a binary sink or filesystem path without retaining the complete input. It returns the number of archive bytes written.

import json

from jzpack import write_records

with open("events.ndjson", "r", encoding="utf-8") as source:
    records = (json.loads(line) for line in source)
    written = write_records(records, "events.jzpk")

The writer applies finite row, byte, schema, path, node, and depth limits. It retries positive short writes to a caller-owned binary sink, which remains open; a failure may leave a partial archive. For a path, it writes to a same-directory temporary file and atomically replaces the destination after success. Exact defaults, validation rules, and failure behavior are in the writer resource model.

File helpers

compress_to_file and decompress_from_file accept paths and binary file-like objects. Path writes use a temporary file and atomic replacement; streams use their current position and remain open. These helpers buffer the complete JZPK payload and are not bounded-memory APIs. compress_to_file returns the compressed byte count.

import io

buffer = io.BytesIO()
compressor.compress_to_file(data, buffer)
buffer.seek(0)
assert compressor.decompress_from_file(buffer) == data

Parameters:

  • level: zstd compression level 1-22 (default: 3)
  • fast: skip column encoding analysis for speed (default: False)
  • max_output_size: optional decompression limit in bytes
  • max_records: optional decompression limit in records

compress accepts a mapping, a list of mappings, or an iterable of mappings; mapping keys must be strings. Supported values include None, booleans, strings, bytes, MessagePack-range integers, binary64 floats, lists, and nested mappings. It preserves missing versus explicit null and exact float bits; mapping key order is not guaranteed. Tuples may be returned as lists, and jzpack does not provide a custom conversion hook. See the format specification for details.

Older v3 archives remain readable, including historical numeric DELTA payloads. Some earlier writers could lose float or mixed numeric distinctions before the archive was stored, and those original values cannot be recovered by a newer reader.

StreamingCompressor accumulates column data until finalize() and emits the same v3 format as compress; it is not a bounded-memory archive writer. Use write_records above for bounded-memory output to a binary sink or path. See the roadmap. add_batch is not atomic: after a validation failure it may retain a prefix of valid records. Call clear() or discard that compressor before restarting a failed ingestion.

How It Works

  1. Schema grouping — records grouped by field structure
  2. Columnar storage — fields stored as columns
  3. Smart encoding — RLE, Delta, Dictionary per column type
  4. MessagePack + Zstandard — binary serialization + compression

Format and compatibility

JZPK version 3 is the sole reader and writer format. It includes deterministic schema identifiers, explicit row counts, collision-safe nested paths, chunk integrity checks, and bounded sequential iteration. See the format specification.

Malformed payloads raise typed exceptions exported from the package, including InvalidFormatError, UnsupportedVersionError, and ResourceLimitError.

Development

python -m pip install -e ".[dev]"
python -m pytest
python -m build

License

MIT

Metadata

Release files for jzpack 0.5.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jzpack 0.5.5
File Size Uploaded
jzpack-0.5.5.tar.gz 252.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jzpack 0.5.5
File Interpreter ABI Platform
jzpack-0.5.5-py3-none-any.whl Python 3 none any Details

Total release size: 279.8 kB

Release files / jzpack-0.5.5.tar.gz

Download URL jzpack-0.5.5.tar.gz
Size 252.5 kB
Tags Source
SHA-256 checksum
How to use checksums
2f53139bdf0faa83848254ca0a9049ea2265e2014995df75b1c611d949112830
BLAKE2b-256 checksum
How to use checksums
ad80a2d5c82e0322f43404cdf62846a3aa87a5249b40e690109ec9614b6161bf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / jzpack-0.5.5-py3-none-any.whl

Download URL jzpack-0.5.5-py3-none-any.whl
Size 27.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0ec8e8d822d5283f1aaf39a2951d57722ccc7729aca687f6939912717523781d
BLAKE2b-256 checksum
How to use checksums
5f58e8bbd3b4ca426c213868b10c94802f265d1596dbf1af1f0037c88b0b04f4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.7

2 release files

0.5.6

2 release files

This release

0.5.5 This release

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page