Write a dataframe pipeline once, run it on Dask or pandas
Project description
BetterFrame
Write a dataframe pipeline once, run it on Dask or pandas.
Not a DataFrame implementation. BetterFrame adapts the execution semantics
that differ between the two engines — per-partition application, meta schemas,
aggregation keywords, persistence — so a pipeline that has to work on both does
not fill up with if is_dask: branches.
Why
Dask is the right tool when data outgrows memory, and a needless tax when it does not. A pipeline that wants both ends up either duplicated or littered with engine checks. The differences are few but awkward:
| Dask | pandas | |
|---|---|---|
| apply a function per partition | df.map_partitions(fn, meta=…) |
fn(df) |
| index level names | df.index._meta.names |
df.index.names |
groupby().agg() extras |
split_out=… |
— |
| make the result concrete | df.persist() |
already is |
| column assignment | builds a graph | mutates in place |
Install
pip install betterframe # pandas only
pip install "betterframe[dask]" # to drive Dask frames as well
Use
from betterframe import BetterFrame
def compute(records):
frame = BetterFrame(records).mutable() # records may be Dask or pandas
frame = frame.apply(normalise, meta=lambda: build_meta(frame.meta_source()))
result = frame.pipe(
lambda df, kw=frame.agg_kwargs(): df.groupby("key").agg({"scaled": "sum"}, **kw)
)
return result.finalize()
BetterFrame binds a frame to the operations its engine needs, so the frame
stops being an argument to every call. Methods that produce a frame return a
BetterFrame, so a pipeline chains; finalize() ends the chain and hands back
a native frame, and native reaches the underlying one at any point.
It is deliberately thin — it does not proxy the dataframe API. Real work
still happens on the frame itself, through pipe() or native. Wrapping exists
to answer engine questions, not to replace pandas.
The same function now runs on either engine, and there is exactly one place — this package — that knows the difference.
The part worth knowing about
Most of the surface is mechanical. The subtle part is Dask's meta handling,
because getting it wrong produces frames that differ only in dtype, only for
empty inputs, and only sometimes:
- Dask coerces every partition to a meta schema. A helper that short-circuits on an empty frame and returns it unchanged still comes out with the populated schema, because Dask repairs it afterwards.
- Where no explicit meta is given, Dask infers one by running the function against a synthetic non-empty frame. So even without a meta, an empty partition ends up with the schema the function would have produced had there been rows.
pandas does neither. PandasOps.apply reproduces both, so an empty frame is
not silently a different dtype depending on which engine produced it. This is
the failure mode BetterFrame exists to prevent: it is invisible in every test
whose fixtures happen to be fully populated.
Set aggregations
Neither engine ships one — pandas has no set aggregation at all, and Dask needs
a three-stage Aggregation whose chunk and combine steps have to agree. Both
are provided, and both return frozenset, so a value aggregated one way
compares equal to the same value aggregated the other:
bf = BetterFrame(records)
grouped = records.groupby("key").agg(
{
"name": bf.ops.set_union(), # distinct values -> frozenset
"tags": bf.ops.set_union_flatten(),
}, # union of set-valued cells
**bf.agg_kwargs(),
)
Flattening is built on betterset, which
handles what a naive set().union(*values) gets wrong:
| input | set().union(*values) |
set_union_flatten |
|---|---|---|
["abc"] |
{"a", "b", "c"} — string shredded |
{"abc"} |
[42, {"d"}] |
TypeError |
{42, "d"} |
[None, {"d"}] |
TypeError |
{"d"} |
The string case is the one that bites: it silently turns a set of names into a set of letters, and nothing raises.
Dask stringifies set-valued columns unless you stop it
Dask's dataframe.convert-string is on by default, and it rewrites object
columns to its own string dtype when a frame is constructed. A column holding
sets is turned into a column holding their reprs, before any aggregation runs:
frame = pd.DataFrame({"g": ["x", "x"], "tags": ["abc", frozenset({"d"})]})
dd.from_pandas(frame, npartitions=1) # tags dtype -> string
# set_union_flatten now yields {"abc", "frozenset({'d'})"}
with dask.config.set({"dataframe.convert-string": False}):
dd.from_pandas(frame, npartitions=1) # tags dtype -> object
# set_union_flatten yields {"abc", "d"}
No library can recover the values once that has happened -- the conversion occurs at construction, upstream of anything BetterFrame sees. If you hold sets in a column, turn the conversion off.
Extending
Subclass DaskOps / PandasOps for engine-specific behaviour of your own —
custom aggregations, say — and hand the subclasses to BetterFrame:
from betterframe import BetterFrame, DaskOps, PandasOps
class MyDaskOps(DaskOps):
def set_union(self):
return some_dask_aggregation()
class MyPandasOps(PandasOps):
def set_union(self):
return some_pandas_callable
frame = BetterFrame(records, dask_ops=MyDaskOps, pandas_ops=MyPandasOps)
frame.ops.set_union() # `ops` reaches engine questions that take no frame
Domain-specific aggregations are deliberately left out of the core so the package stays about engine semantics.
Development
uv sync --dev
uv run pytest
uv run ruff check .
uv run ruff format --check .
License
MIT
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file betterframe-0.2.1.tar.gz.
File metadata
- Download URL: betterframe-0.2.1.tar.gz
- Upload date:
- Size: 78.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7a480f4a074cf1d53fed7a693085b19ab62ba87747c6c385c764a5dfce817b3b
|
|
| MD5 |
3eabad8be823e71b70ce65f979315a7d
|
|
| BLAKE2b-256 |
97229068e4a8a90919274992e873bcc02fe73a6807413686e2d1f4ba784d7539
|
Provenance
The following attestation bundles were made for betterframe-0.2.1.tar.gz:
Publisher:
release.yaml on izzet/betterframe
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
betterframe-0.2.1.tar.gz -
Subject digest:
7a480f4a074cf1d53fed7a693085b19ab62ba87747c6c385c764a5dfce817b3b - Sigstore transparency entry: 2263516285
- Sigstore integration time:
-
Permalink:
izzet/betterframe@674bff095c77280eeba9842069c1ab89b4a8d228 -
Branch / Tag:
refs/tags/v0.2.1 - Owner: https://github.com/izzet
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yaml@674bff095c77280eeba9842069c1ab89b4a8d228 -
Trigger Event:
release
-
Statement type:
File details
Details for the file betterframe-0.2.1-py3-none-any.whl.
File metadata
- Download URL: betterframe-0.2.1-py3-none-any.whl
- Upload date:
- Size: 11.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2fe6ab5ad1ac06374b30d0577f18dedc5ff838f7314227426dc156cee272f229
|
|
| MD5 |
30f0b1b87aa13c232baf78d136ef9327
|
|
| BLAKE2b-256 |
fbb3617bbaf3fcfb85d50773b776ca882f59087a4465f193ea8d7c01be21fea8
|
Provenance
The following attestation bundles were made for betterframe-0.2.1-py3-none-any.whl:
Publisher:
release.yaml on izzet/betterframe
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
betterframe-0.2.1-py3-none-any.whl -
Subject digest:
2fe6ab5ad1ac06374b30d0577f18dedc5ff838f7314227426dc156cee272f229 - Sigstore transparency entry: 2263516377
- Sigstore integration time:
-
Permalink:
izzet/betterframe@674bff095c77280eeba9842069c1ab89b4a8d228 -
Branch / Tag:
refs/tags/v0.2.1 - Owner: https://github.com/izzet
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yaml@674bff095c77280eeba9842069c1ab89b4a8d228 -
Trigger Event:
release
-
Statement type: