Fast Parquet/CSV for pandas: column pruning, PyArrow pushdown, unified multi-file scans, tuple filters, optional parallel I/O, footer-only schema.
Project description
snapparquet
Small, pandas-first helpers for fast Parquet and CSV workflows: column pruning, tuple-based filters with PyArrow dataset predicate pushdown, optional legacy integer date/partition ranges, unified multi-file scans, optional parallel per-file I/O, and footer-only schema utilities for discovery and safe column lists.
Non-goals: snapparquet is not a query engine. For SQL over huge Parquet lakes, lazy plans, or TB-scale analytics, use DuckDB, Polars, or similar; keep snapparquet for dashboards, ETL scripts, and pipelines that want a thin layer over pandas + PyArrow.
Requirements
- Python 3.9+
- pandas ≥ 1.5
- PyArrow ≥ 12
Install
Current release: 0.3.1
pip install "snapparquet==0.3.1"
Compatible future releases:
pip install "snapparquet>=0.3.1"
From a source checkout (editable):
cd snapparquet
pip install -e ".[dev]"
Quick start
import snapparquet as sp
cols = ["user_id", "load_date", "revenue"]
# Parquet with column pruning + legacy inclusive integer range on load_date
df = sp.read_tabular(
"metrics.parquet",
columns=cols,
filter_column="load_date",
range_start=20250301,
range_end=20250331,
)
# Many daily Parquet files: one PyArrow dataset scan when possible
df = sp.read_tabular_many(
["day1.parquet", "day2.parquet", "day3.parquet"],
columns=cols,
filter_column="load_date",
range_start=20250301,
range_end=20250331,
)
# Tuple filters: AND-combined; pushed to Parquet via pyarrow.dataset when reading Parquet
df = sp.read_parquet(
"events.parquet",
columns=["status", "revenue"],
filters=[("status", "in", ["active", "trial"]), ("revenue", ">", 0)],
)
# Pandas 2+ optional Arrow-backed dtypes (numpy fallback + warning on older pandas)
df = sp.read_parquet("wide.parquet", backend="pyarrow")
Reading API
read_parquet(path, ...)
Reads a single Parquet file into a pandas DataFrame.
| Argument | Description |
|---|---|
columns |
Optional list of column names to decode (projection / pruning). Filter columns are auto-included when needed. |
filter_column, range_start, range_end |
Legacy inclusive integer range on filter_column. If that column is integer-typed in the Parquet footer and use_pushdown=True, the range is merged into the PyArrow dataset filter so row groups may be skipped. |
use_pushdown |
If False, the legacy range is not pushed (still applied in pandas when the column is not pushed). |
filters |
Sequence of (column, operator, value) tuples (see Tuple filters). AND-combined; invalid specs raise ValueError. |
backend |
"numpy" (default) or "pyarrow". Invalid values raise ValueError immediately (before I/O). |
dictionary_as_categorical |
On the numpy path, map Arrow dictionary-encoded columns to pandas.Categorical to save memory. When True, self_destruct during Arrow→pandas is disabled so post-processing can still read dictionary buffers safely. |
Read path (implementation summary):
- Try PyArrow
datasetwith a Parquet format that enables row-group pre-buffering when supported. - Call
Dataset.to_table(columns=..., filter=..., use_threads=True)so projection, predicate pushdown, and threaded decode run together. - Convert with
Table.to_pandasusingsplit_blocks/use_threads, and optionaltypes_mapper=pd.ArrowDtypeforbackend="pyarrow". - If the legacy range could not be expressed in the dataset filter (e.g. column not integer in file, or
use_pushdown=False), apply the same inclusive range in pandas after the Arrow read—even when tuplefilterswere pushed down—so results match the pandas fallback path. - On any failure opening or scanning via dataset, fall back to
pandas.read_parquetand apply range + filters in pandas.
read_csv(path, ...)
CSV → pandas with optional usecols (from columns=), legacy range, and tuple filters. No Arrow pushdown (filters and range are applied in pandas).
read_tabular(path, ...)
Dispatches on the path suffix (.parquet / .csv, case-insensitive).
- Parquet: forwards to
read_parquet(includingfilters,backend,dictionary_as_categorical). - CSV: forwards to
read_csv.
If columns is set, Parquet uses schema intersection with the file (intersect_columns); CSV reads the header and keeps overlapping names. If the intersection is empty and full_read_if_no_column_match=True (default), all columns are read then filters apply; if False, an empty DataFrame is returned when intersection is empty.
read_tabular_many(paths, ...) and read_parquet_many(paths, ...)
Read many paths and concatenate (row-wise), index reset.
| Argument | Description |
|---|---|
unified_parquet |
Default True. If every path ends with .parquet, try one pyarrow.dataset over the full list (single scan, shared pushdown across fragments). On failure (schemas, I/O quirks), fall back to per-file reads. |
parallel_workers |
None = auto-scaled thread pool for the per-file fallback (I/O-bound oversubscription). 0 or 1 = strict sequential. Otherwise caps at file count. |
ignore_errors |
If True, failed files become empty frames instead of raising (use with care). |
| Other kwargs | Same as read_tabular / read_parquet where applicable (columns, filter_column, range_start, range_end, use_pushdown, filters, backend, dictionary_as_categorical, full_read_if_no_column_match). |
read_parquet_many is an alias of read_tabular_many for naming clarity in Parquet-only code.
Tuple filters
Each filter is a 3-tuple: (column_name, operator, value).
Supported operators: ==, !=, >, >=, <, <=, in, not_in.
- For
in/not_in,valuemust be a non-string sequence (e.g. list, tuple); strings are rejected to avoid foot-guns with character iteration.
Filters are AND-merged. For Parquet, build_filter() produces a pyarrow.compute expression consumed by the dataset API.
Programmatic reuse: apply_filters_pandas(df, filters) applies the same semantics in memory (CSV path, or debugging). It raises if a referenced column is missing.
Column names are normalized with norm_col (strip whitespace). Overly long names or control characters in filter columns raise ValueError (defensive guard for untrusted specs).
Schema helpers (footer / metadata only)
These avoid decoding column data where possible.
| Function | Purpose |
|---|---|
schema_column_names(path) |
frozenset of column names from Parquet metadata, or None if unreadable. |
intersect_columns(path, wanted) |
Ordered intersection of wanted with file columns; None if schema unreadable; [] if no overlap. |
column_is_integer_type(path, name) |
Whether name exists and is an Arrow integer type (used to decide legacy range pushdown). |
schema_has_all(names, required) |
All required columns present in a name set. |
schema_matches_any_group(names, groups) |
True if at least one group’s columns are all present (e.g. alternate mart layouts). |
norm_col(name) |
Normalize a column identifier string. |
backend and dictionary columns
| Option | Behavior |
|---|---|
backend="numpy" (default) |
Classic numpy-backed pandas dtypes (after Arrow conversion). |
backend="pyarrow" |
Uses types_mapper=pd.ArrowDtype when pandas ≥ 2; otherwise emits UserWarning and uses numpy. |
dictionary_as_categorical=True |
After conversion, dictionary-encoded Arrow columns are turned into pandas.Categorical on the numpy path where possible. |
Performance notes
- Arrow-first single-file reads: Parquet reads prefer
pyarrow.dataset+ threadedto_tablefor both filtered and unfiltered cases, then fall back to pandas only on failure. - Pre-buffering: Parquet fragment scan options use pre-buffered row-group reads when the installed PyArrow supports it (throughput vs memory; safe fallback if options are unavailable).
- Unified multi-file Parquet: One dataset over many files reduces open/scan overhead and improves fragment-level pruning vs naive per-file loops.
- Parallel fallback: Default worker count scales with CPU and file count (capped) to hide I/O latency; use
parallel_workers=1on very slow disks or tiny file counts if needed. - Heavier workloads: For lazy query planning or massive scans, prefer Polars/DuckDB; snapparquet stays small and pandas-centric.
Development and tests
pip install -e ".[dev]"
pytest tests/ -q # full suite
pytest tests/ -q -m "not stress" # skip @pytest.mark.stress cases
Stress-marked tests are larger synthetic scans (still seconds-scale, not microbenchmarks).
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file snapparquet-0.3.1.tar.gz.
File metadata
- Download URL: snapparquet-0.3.1.tar.gz
- Upload date:
- Size: 21.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e5bd7235014ec4a3e4ce193df10ba81c574e73ad3f56682bc6b3fd51e244c16d
|
|
| MD5 |
53ea71495a9aa1f5d43a4984d406c15e
|
|
| BLAKE2b-256 |
78db4707bd922c61368dfa35609e906a361f4ab3ea1d4973fa0f5ab4fc14a153
|
File details
Details for the file snapparquet-0.3.1-py3-none-any.whl.
File metadata
- Download URL: snapparquet-0.3.1-py3-none-any.whl
- Upload date:
- Size: 16.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8376298a865e3cf6eb4e6f105c0b8bdc427bac05323066532de82c574e064710
|
|
| MD5 |
87fe2753aa87cdff7f4c3011cb2a346b
|
|
| BLAKE2b-256 |
88efa5053ac03c1a69919bfee1c5751b6901b226cfab1912eb977afac1e1849b
|