A utility library for generating Polars expressions to work with nested data structures
Project description
Polars Nexpresso ☕
Polars Nexpresso is a utility library for working with nested and hierarchical data in Polars. It provides two main capabilities:
- Nested Expression Builder - Clean, intuitive syntax for transforming deeply nested structs and lists
- Hierarchical Packer - Pack/unpack operations for hierarchical data, similar to pandas MultiIndex but using Polars' native nested types
Nexpresso = Nested Expression + ☕ (espresso)
Installation
pip install polars-nexpresso
Or using uv:
uv add polars-nexpresso
Quick Start
Nested Expression Builder
Transform deeply nested data with intuitive dictionary syntax:
import polars as pl
from nexpresso import generate_nested_exprs
df = pl.DataFrame({
"order": [
{"customer": "Alice", "items": [{"name": "Laptop", "price": 999}, {"name": "Mouse", "price": 25}]},
{"customer": "Bob", "items": [{"name": "Keyboard", "price": 75}]},
]
})
# Define transformations declaratively
fields = {
"order": {
"items": {
"price": lambda x: x * 1.1, # 10% price increase
"discounted": pl.field("price") * 0.9, # New field
}
}
}
exprs = generate_nested_exprs(fields, df.schema, struct_mode="with_fields")
result = df.select(exprs)
Hierarchical Packer
Build and navigate hierarchical data from normalized tables:
from nexpresso import HierarchicalPacker, HierarchySpec, LevelSpec
# Define hierarchy
spec = HierarchySpec.from_levels(
LevelSpec(name="region", id_fields=["id"]),
LevelSpec(name="store", id_fields=["id"], parent_keys=["region_id"]),
)
packer = HierarchicalPacker(spec)
# Build from separate tables (like database tables)
nested = packer.build_from_tables({
"region": regions_df,
"store": stores_df,
})
# Navigate between granularities
flat = packer.unpack(nested, "store") # Explode to store level
packed = packer.pack(flat, "region") # Aggregate back to region level
Features
Nested Expression Builder
| Feature | Description |
|---|---|
| Field Selection | Keep fields as-is with None |
| Transformations | Apply lambdas: lambda x: x * 2 |
| New Fields | Create with pl.Expr: pl.field("a") + pl.field("b") |
| Deep Nesting | Works with any depth of structs/lists |
| Two Modes | select (keep specified) or with_fields (keep all) |
Hierarchical Packer
| Feature | Description |
|---|---|
| Build from Tables | Join normalized tables into nested hierarchy |
| Pack/Unpack | Navigate between granularity levels |
| Streaming Pack/Unpack | Memory-bounded pack_streaming / unpack_streaming for large data |
| Normalize/Denormalize | Split into per-level tables and reconstruct |
| Validation | Check for null keys and data integrity |
| Custom Separators | Use any separator (default: .) |
| Type Preservation | DataFrame in = DataFrame out |
Core Concepts
Field Value Types
When defining transformations:
None: Keep the field unchangeddict: Recursively process nested structuresCallable: Apply function to field (e.g.,lambda x: x * 2)pl.Expr: Create/modify field with full expression
Struct Modes
"select": Only keep fields specified in the dictionary"with_fields": Keep all fields, add/modify specified ones
Hierarchy Levels
Define your data hierarchy with LevelSpec:
LevelSpec(
name="store", # Level identifier
id_fields=["id"], # Unique key columns
parent_keys=["region_id"], # Foreign key to parent (for build_from_tables)
)
Examples
Lists of Structs
df = pl.DataFrame({
"items": [[{"name": "A", "qty": 5}, {"name": "B", "qty": 3}]]
})
fields = {
"items": {
"qty": lambda x: x * 2,
"total": pl.field("qty") * 10,
}
}
result = apply_nested_operations(df, fields, struct_mode="with_fields")
Conditional Logic
fields = {
"customer": {
"discount": pl.when(pl.field("tier") == "Gold")
.then(0.15)
.when(pl.field("tier") == "Silver")
.then(0.10)
.otherwise(0.05),
}
}
Building Hierarchies from Database Tables
# Tables with foreign key relationships
regions = pl.DataFrame({"id": ["west", "east"], "name": ["West", "East"]})
stores = pl.DataFrame({
"id": ["s1", "s2"],
"name": ["Store 1", "Store 2"],
"region_id": ["west", "east"]
})
spec = HierarchySpec.from_levels(
LevelSpec(name="region", id_fields=["id"]),
LevelSpec(name="store", id_fields=["id"], parent_keys=["region_id"]),
)
packer = HierarchicalPacker(spec)
nested = packer.build_from_tables({"region": regions, "store": stores})
# Result: Stores nested within their regions
Normalize and Denormalize
# Split nested data into separate tables
tables = packer.normalize(nested_df)
# {"region": region_df, "store": store_df, ...}
# Reconstruct from separate tables
rebuilt = packer.denormalize(tables)
Memory-bounded packing for large data
pack builds the nested result with a group_by whose state holds every group
in memory, so peak memory scales with the whole dataset even under the streaming
engine. For datasets that don't fit comfortably in RAM, pack_streaming buckets
the input by the root-level key (keeping each entity's rows together), packs
each bucket independently while sinking to Parquet, and returns a chainable
LazyFrame — bounding peak memory to a single bucket.
# Bound peak memory by processing the data in root-key buckets.
# Accepts a DataFrame, LazyFrame, or a Parquet path/glob (scanned lazily).
packed = packer.pack_streaming(flat_df, "region", partitions=32)
# It returns a LazyFrame, so you can keep composing lazily:
top = (
packer.pack_streaming("s3_dump/*.parquet", "region", partitions=64)
.filter(pl.col("region.id").is_in(active_regions))
.collect(engine="streaming")
)
# defer=False sinks eagerly and returns a scan handle so downstream work streams
# straight from disk — safest when even the packed result is too big for RAM.
handle = packer.pack_streaming(flat_df, "region", partitions=64, defer=False)
# unpack already streams; unpack_streaming keeps it lazy / disk-to-disk.
leaves = packer.unpack_streaming("packed.parquet", "store", sink_path="leaves.parquet")
More buckets means lower peak memory (and more temporary files). See
benchmarks/ for a peak-RSS comparison of pack vs
pack_streaming.
Heavy root attributes: parent_strategy="split_join"
When the root level carries heavy attributes that repeat across every leaf row
(e.g. a per-entity blob, thumbnail, or embedding), carrying them through the pack
aggregation is wasteful. pack(..., parent_strategy="split_join") instead pulls
those attributes into a small dimension table (unique per root key) and reattaches
them after packing — identical results, but the heavy column is touched once per
entity instead of once per leaf row.
# Equivalent to the default pack, but far cheaper when root attributes dominate.
packed = packer.pack(flat_df, "region", parent_strategy="split_join")
This wins big when root attributes are heavy relative to the child data (measured
up to ~9x faster, half the memory); for child-dominated data it adds join overhead
for no gain, so it is opt-in. See benchmarks/ for the trade-offs.
Note on ordering: packing no longer performs a global sort, so top-level row order is not guaranteed. Child-list order is still preserved when
preserve_child_order=True(the default) or via a level'sorder_by. De-duplication and null handling of parent attributes are unaffected.
API Reference
Nested Expressions
generate_nested_exprs(fields, schema, struct_mode="select")
Generate Polars expressions for nested data operations.
Parameters:
fields: Dictionary defining operations on columns/fieldsschema: DataFrame schema (or DataFrame/LazyFrame to extract schema)struct_mode:"select"or"with_fields"
Returns: list[pl.Expr]
apply_nested_operations(df, fields, struct_mode="select", use_with_columns=False)
Apply nested operations directly to a DataFrame.
Hierarchical Packer
HierarchicalPacker(spec, *, granularity_separator=".", escape_char="\\", preserve_child_order=True, validate_on_pack=True)
Main class for hierarchical operations.
Key Methods:
pack(frame, to_level, *, extra_columns="preserve", parent_strategy="aggregate")- Pack to coarser granularity (parent_strategy="split_join"reattaches heavy root attributes via a join)pack_streaming(source, to_level, *, partitions=16, ...)- Memory-bounded pack for large dataunpack(frame, to_level)- Unpack to finer granularitynormalize(frame)- Split into per-level tablesdenormalize(tables)- Reconstruct from per-level tablesbuild_from_tables(tables)- Build hierarchy from normalized tablesvalidate(frame)- Check data integrity
HierarchySpec.from_levels(*levels, key_aliases=None)
Create a hierarchy specification from level definitions.
LevelSpec(name, id_fields, required_fields=None, order_by=None, parent_keys=None)
Define a single level in the hierarchy.
Running Examples
# Run comprehensive examples
python examples.py
# Or run specific module examples
python -m nexpresso.hierarchical_packer
Performance
Both components generate native Polars expressions, so performance is equivalent to hand-written code. All operations are lazy-compatible and benefit from Polars' query optimization.
License
MIT License - see LICENSE file for details.
Contributing
Contributions welcome! Please feel free to submit issues and pull requests.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file polars_nexpresso-0.4.0.tar.gz.
File metadata
- Download URL: polars_nexpresso-0.4.0.tar.gz
- Upload date:
- Size: 197.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
abe1a3b4e4eb2cfedacd1f535f125d2e52feb09496f4a152d4df7524f3586e82
|
|
| MD5 |
97b9f3d5392c7732f494427e5c0a24ea
|
|
| BLAKE2b-256 |
9d666a366b2c643b4cac2f3fdc6216902a832dd1a88d011d4f5088775dc28464
|
Provenance
The following attestation bundles were made for polars_nexpresso-0.4.0.tar.gz:
Publisher:
publish.yml on heshamdar/polars-nexpresso
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
polars_nexpresso-0.4.0.tar.gz -
Subject digest:
abe1a3b4e4eb2cfedacd1f535f125d2e52feb09496f4a152d4df7524f3586e82 - Sigstore transparency entry: 1866997322
- Sigstore integration time:
-
Permalink:
heshamdar/polars-nexpresso@bf558a947d30b6eb40eb584334d6733a1f545439 -
Branch / Tag:
refs/tags/v0.4.1 - Owner: https://github.com/heshamdar
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@bf558a947d30b6eb40eb584334d6733a1f545439 -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file polars_nexpresso-0.4.0-py3-none-any.whl.
File metadata
- Download URL: polars_nexpresso-0.4.0-py3-none-any.whl
- Upload date:
- Size: 43.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1d1f9af73969650d4b7d2388d47d6b78537f9f3cc7155ca342f5cf1474afb084
|
|
| MD5 |
c20cccfd8c22d7845077ca78d4ca4b08
|
|
| BLAKE2b-256 |
b48a7d4b8e60277876228d9dcf4ad998d6d5d0affeba8288bfe3670b080a28b7
|
Provenance
The following attestation bundles were made for polars_nexpresso-0.4.0-py3-none-any.whl:
Publisher:
publish.yml on heshamdar/polars-nexpresso
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
polars_nexpresso-0.4.0-py3-none-any.whl -
Subject digest:
1d1f9af73969650d4b7d2388d47d6b78537f9f3cc7155ca342f5cf1474afb084 - Sigstore transparency entry: 1866997337
- Sigstore integration time:
-
Permalink:
heshamdar/polars-nexpresso@bf558a947d30b6eb40eb584334d6733a1f545439 -
Branch / Tag:
refs/tags/v0.4.1 - Owner: https://github.com/heshamdar
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@bf558a947d30b6eb40eb584334d6733a1f545439 -
Trigger Event:
workflow_dispatch
-
Statement type: