duckpd
DuckPD is an experimental lazy DataFrame library with a pandas-shaped frontend and DuckDB as its execution engine.
[!WARNING] DuckPD is a work in progress and is not yet recommended for production-critical workloads. The API and supported pandas semantics may change between
0.xreleases, and many pandas operations are intentionally unsupported. Validate results and resource behavior for each intended workload before adopting it.
DuckPD intentionally supports a small, explicit subset of pandas rather than
silently falling back to materializing a complete pandas DataFrame. See the
release policy for the pre-1.0 stability policy.
Current capabilities
- Lazy pandas, Arrow, Parquet, DuckDB table, and read-only SQL sources.
- Column selection, boolean filtering, arithmetic expressions,
assign,sort_values, andlimit. - Eager DataFrame and Series
count,size,sum,mean,min, andmaxreductions over numeric and boolean data, includingskipna,min_count, and DataFramenumeric_onlysupport. - Explicit lazy indexes with
set_index()/reset_index()and sourceindex=/order_by=declarations. - Explicit pandas collection, bounded
head, Arrow tables and record batches, physical plan inspection, and direct Parquet writes. - Session-level memory, spill-directory, temporary-size, and thread settings.
- Rejection of ambiguous cross-frame alignment and mutating SQL.
Example
import duckpd as pd
orders = pd.read_parquet("orders/*.parquet")
result = (
orders[orders["status"] == "paid"]
.assign(net=lambda frame: frame["amount"] - frame["refund_amount"])
.sort_values("net", ascending=False)[["order_id", "net"]]
.limit(100)
)
print(result.explain())
preview = result.head(10)
result.write_parquet("largest-paid-orders.parquet")
pandas_result = result.collect()
Transformations above are lazy. explain(), head(), collect(), Arrow output,
and file output are explicit execution boundaries. limit() stays lazy while
head() returns a bounded pandas preview.
Demos
Small runnable programs are available in demo/:
uv run python demo/basic_pipeline.py
uv run python demo/parquet_pipeline.py
uv run python demo/reduction_pipeline.py
uv run python demo/generate_market_data.py
uv run python demo/market_data_demo.py
See the benchmark results for performance and memory comparisons between DuckPD and pandas across 100 MB, 1 GB, and 5 GB datasets.
Development
uv sync --frozen --group dev
make check
make build
GNU Make is optional. The equivalent commands are:
uv run pytest
uv run ruff check .
uv run ruff format --check .
uv run pyright
uv build
See the documentation index for the implementation roadmap, architecture decisions, benchmarks, research, and changelog.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file duckpd-0.0.2.tar.gz.
File metadata
- Download URL: duckpd-0.0.2.tar.gz
- Upload date:
- Size: 145.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.21 {"installer":{"name":"uv","version":"0.11.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b2284aadb1c29a181441ab261b61432c79aa9d2547e8b5ad549b11fa39f9950a
|
|
| MD5 |
4d698fb4422387c4c869381eda6046ef
|
|
| BLAKE2b-256 |
9748f25742e55a45dfbc8b956a00d45a56d7344caa7f2a880655b7a781a77c32
|
File details
Details for the file duckpd-0.0.2-py3-none-any.whl.
File metadata
- Download URL: duckpd-0.0.2-py3-none-any.whl
- Upload date:
- Size: 32.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.21 {"installer":{"name":"uv","version":"0.11.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a305a72716d52d75ba25f9561595f89bc17b71e98e7da67d9b4fd309d9f4ce72
|
|
| MD5 |
aa7cbfcd226f9d6b09b0c9eed5c5e67c
|
|
| BLAKE2b-256 |
688b0e3056ab03a60f042b9fadb1e4659a6ca68b1639556498fdedcc53dbfe26
|