Skip to main content

oparq — smaller Parquet files for S3 and GCS

oparq is a Python library for reducing Parquet storage bytes by choosing compression-friendly row orders. Learn sort keys from samples once, save the plan, and reuse it across a local, Amazon S3, or Google Cloud Storage dataset. PyArrow writes the files; optional DuckDB sorting can spill to disk.

It changes physical row order, not the logical row values. Compression gains depend on the data and existing clustering; smaller files and faster queries are not guaranteed.

Install

The library is MIT-licensed and requires Python 3.12+ and PyArrow 16+.

pip install oparq
pip install "oparq[duckdb]"  # optional larger-than-memory DuckDB sorting

Write an Arrow table

import pyarrow as pa
import oparq

table = pa.table({
    "account": ["b", "a", "b", "a"],
    "payload": ["repeat-b", "repeat-a", "repeat-b", "repeat-a"],
})
result = oparq.write(table, "optimized.parquet", algorithm="auto")
print(result.plan.sort_keys, result.file_size)
print(result.planning_seconds, result.permutation_seconds,
      result.gathering_seconds, result.write_seconds)

New-table writes use Zstandard's Arrow codec default, currently level 1, with dictionaries and statistics enabled. auto uses bounded codec_fast key selection. Use algorithm="none" with the same writer settings for a controlled baseline.

Learn once, rewrite a bucket

import json
from pathlib import Path
import oparq

# An exporter manifest must record the actual source codec/level.
source_settings = json.loads(Path("source-compression.json").read_text())
plan = oparq.fit_dataset(
    "s3://source-bucket/events",
    algorithms=("codec_fast", "portfolio"),
    prefix=("event_date",),  # Optional physical leading key; discover its suffix.
    compression_manifest=source_settings,
)
plan.save("events-order.json")

result = oparq.rewrite_dataset(
    "s3://source-bucket/events",
    "gs://destination-bucket/events",
    plan=oparq.RewritePlan.load("events-order.json"),
    compression_manifest=source_settings,
    engine="duckdb",  # Install the extra; "arrow" is the default.
    memory_limit="6GB",
)
print(result.rewritten_files, result.copied_files, result.skipped_files)

The saved plan skips future profiling and algorithm selection. Rewrites retain Parquet relative paths, Hive partition directories, and file boundaries, sorting within each physical file. They do not coalesce partitions, copy arbitrary sidecars, or provide a bucket-wide transaction. Prefix keys must be physical columns; virtual Hive columns are not injected.

Defaults that protect your data

  • File rewrites preserve source codecs and known encoder levels. Parquet footers do not store compression levels: unknown ZSTD levels require trusted provenance or an explicit output level, never a silent jump to 9.
  • With preserved settings, no mandatory prefix, and skip_unchanged=True, no-key plans copy/skip the original. A sorted output that is not smaller is discarded before publication. The attempted work still costs time.
  • Explicit codec/level/layout changes and required prefixes are honored even if files grow. Existing destinations are protected unless overwrite=True.
  • A source page index and bloom filters (PyArrow 25+) are rebuilt for the new row order, so rewritten files keep the same pruning structures.
  • Remote reads avoid a staging copy of source files, but data still travels through the worker and DuckDB may spill locally. Run near your buckets.
  • Nested and detected JSON columns are not automatic sort candidates but are carried through sorting. Parquet already omits null-value payloads.
  • Sorting can improve or weaken query clustering. Statistics and advertised natural sort metadata remain truthful; no universal pruning win is promised.

Measured evidence

The September 30, 2026 full-corpus comparison processed every row of 21 datasets: 967,938,981 rows in 986 source files, sorting each dataset globally across its files. Under controlled identical writer settings, portfolio saved 16.28% in aggregate; the best measured method per dataset saved 18.28%. That best mixture was selected after full results, not predicted by auto. Some sample-chosen orders regressed sharply.

rewrite_dataset sorts within each file, so its savings differ from that global sort. The separate Arrow/DuckDB comparison applied the same portfolio keys file by file to all 26,015,956 rows in 40 stock, Hits, and Solana files, with exact per-file value/order checks: stock saved 20.61% and Solana 29.64% (33.54% and 43.00% when sorted globally). Arrow was faster in that run; DuckDB is retained for bounded-memory execution, not a demonstrated universal speed advantage.

The measurements above are from historical development runs, not a claim that the original inputs are public. Methodology, regressions, and caveats are in Benchmarks and limitations; runners and machine-readable results are in benchmarks/.

Documentation

Metadata

Release files for oparq 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for oparq 0.3.0
File Size Uploaded
oparq-0.3.0.tar.gz 47.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for oparq 0.3.0
File Interpreter ABI Platform
oparq-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 104.2 kB

Release files / oparq-0.3.0.tar.gz

Download URL oparq-0.3.0.tar.gz
Size 47.3 kB
Tags Source
SHA-256 checksum
How to use checksums
367559614df1dded514f1fd906b64b6e586cb34ab48c73adc9ef09e77c43e92d
BLAKE2b-256 checksum
How to use checksums
38896258e3d54259c2c93898695b2f9613910e8e42518c7b0e7c7b8aff06ba8f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.7

Release files / oparq-0.3.0-py3-none-any.whl

Download URL oparq-0.3.0-py3-none-any.whl
Size 56.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f08e07b8332995938cba948038651ce31de9c15b16ac7e306d36c243b1b4c702
BLAKE2b-256 checksum
How to use checksums
b707b008b269c0ef477c1161a3cbaf414043a6f54dde15e218ede60ba3171448
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.7

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page