[!WARNING] DuckPD is a work in progress and is not yet recommended for production-critical workloads. The API and supported pandas semantics may change between
0.xreleases, and many pandas operations are intentionally unsupported. Validate results and resource behavior for each intended workload before adopting it.
DuckPD 🦆❤️🐼
DuckPD is DuckDB dressed as a pandas DataFrame.
DuckPD is a lazy DataFrame library with a pandas-shaped API and DuckDB as its execution engine. The goal is to make working with DuckDB feel familiar to pandas users while preserving the performance, scalability, and query-optimization advantages of DuckDB.
Where practical, DuckPD aims to match pandas APIs and semantics closely enough that existing pandas knowledge — and eventually a large amount of pandas-oriented code — transfers naturally. It does not, however, aim to reproduce pandas by sacrificing the properties that make DuckDB valuable.
Project directives
These principles define the direction of DuckPD and should guide API and implementation decisions:
-
Pandas-shaped, DuckDB-native. The public API should feel like pandas, but operations should map naturally onto DuckDB's relational and vectorized execution model.
-
Stay lazy by default. Transformations should build a query plan rather than execute immediately. Execution should happen only at clear and intentional boundaries such as
collect(),head(), Arrow conversion, or file output. -
Never silently fall back to pandas. Unsupported operations should fail explicitly rather than unexpectedly materializing an entire dataset into memory. Users should always be able to reason about where computation happens.
-
Push work into DuckDB. Filtering, projection, joins, aggregation, sorting, expressions, and other supported operations should be translated into DuckDB operations whenever possible so DuckDB can optimize the complete query.
-
Preserve pandas semantics where we claim compatibility. API similarity alone is not enough. Supported operations should match pandas behavior as closely as practical, including edge cases around nulls, indexes, dtypes, grouping, and column behavior.
-
Correctness before coverage. It is better to support a smaller pandas surface correctly than to advertise broad compatibility backed by incomplete semantics, hidden fallbacks, or surprising execution behavior.
-
Make execution visible and predictable. Users should be able to understand when data is scanned, materialized, transferred, or written. Laziness must be a useful property, not hidden magic.
-
Exploit the ecosystem boundaries. DuckPD should interoperate cleanly with pandas, Arrow, Parquet, SQL, and DuckDB itself. Crossing those boundaries should be explicit and inexpensive wherever the underlying systems allow it.
The long-term ambition is broad pandas API coverage where those APIs can be implemented without violating these directives. Compatibility is the interface; DuckDB-native execution is the foundation.
Current capabilities
- Lazy pandas, Arrow, Parquet, DuckDB table, and read-only SQL sources.
- Column selection, boolean filtering, arithmetic expressions,
assign,sort_values,limit, and distinct/drop_duplicates deduplication. - Relational DataFrame joins (
merge) supportinginner,left,right,outer, andcrosswith column collision suffix management. - Multi-DataFrame row-wise concatenation (
duckpd.concat) with schema alignment and null-padding. - Vectorized
.str(e.g.upper,lower,strip,len,contains,replace) and.dt(e.g.year,month,day,hour,minute,second,strftime,to_period) accessor pipelines. - Multi-column
groupby()supporting eager and lazyagg(),sum(),mean(),min(),max(),std(),var(), andcount(). - Eager DataFrame and Series reductions:
count,size,sum,mean,min,max,std,var,median,quantile,any, andallover numeric and boolean data, includingskipna,min_count, and DataFramenumeric_onlysupport. - Explicit lazy indexes with
set_index()/reset_index()and sourceindex=/order_by=declarations, including exact and partial MultiIndex.locselection. - Stable snapshot order for pandas and Arrow inputs, used for deterministic positional operations, duplicate retention, ranking, and top-N ties.
- Context-local implicit sessions, allowing frames created by separate module-level helpers to participate in the same lazy plan.
- Explicit pandas collection, bounded
head, Arrow tables and record batches, physical plan inspection (explain), and direct zero-copy Parquet writes.
Supported pandas API Coverage
DuckPD maps pandas semantics directly to DuckDB's vectorized analytical engine:
| API Category | Supported Methods & Operations | Execution Model |
|---|---|---|
| I/O & Data Loading | read_parquet(), read_sql(), from_pandas(), from_arrow(), sql(), connect() |
Lazy (scans metadata / registers source) |
| Transformations & Projections | df[cols], df[bool_filter], assign(), sort_values(), limit(), drop_duplicates(), set_index(), reset_index() |
Lazy (appends to logical query graph) |
| Joins & Merges | merge() (inner, left, right, outer, cross, custom suffixes) |
Lazy (relational hash join) |
| Concatenation | duckpd.concat() (multi-frame row union, schema alignment, null padding) |
Lazy (union with projection padding) |
String Accessor (.str) |
upper(), lower(), strip(), len(), startswith(), endswith(), contains(), replace() |
Lazy (DuckDB SQL functions) |
Datetime Accessor (.dt) |
year, month, day, hour, minute, second, strftime(), to_period() |
Lazy (DuckDB timestamp extractors) |
| GroupBy Aggregations | groupby().agg(), .sum(), .mean(), .min(), .max(), .std(), .var(), .count() (as_index=True/False) |
Lazy for .agg(), Eager for reductions |
| Statistical Reductions | sum(), mean(), min(), max(), count(), size, std(), var(), median(), quantile(), any(), all() |
Eager (single aggregate SQL pushdown) |
| Collection & Output | collect(), head(n), explain(), write_parquet(), to_arrow_table(), to_arrow_batches() |
Explicit Execution Boundary |
Example
import duckpd as pd
orders = pd.read_parquet("orders/*.parquet")
result = (
orders[orders["status"] == "paid"]
.assign(net=lambda frame: frame["amount"] - frame["refund_amount"])
.sort_values("net", ascending=False)[["order_id", "net"]]
.limit(100)
)
print(result.explain())
preview = result.head(10)
result.write_parquet("largest-paid-orders.parquet")
pandas_result = result.collect()
Transformations above are lazy. explain(), head(), collect(), Arrow output,
and file output are explicit execution boundaries. limit() stays lazy while
head() returns a bounded pandas preview.
Ordering, indexing, and sessions
Pandas and Arrow inputs are snapshots with a stable source row order. DuckPD
tracks that order with hidden relational metadata so operations such as
.iloc, drop_duplicates(keep=...), rank(method="first"), and top-N tie
selection remain deterministic without exposing a synthetic pandas index.
Parquet, CSV, SQL, and DuckDB table scans remain unordered unless order_by=
is provided. Ordering-sensitive operations fail with
UnorderedOperationError rather than relying on accidental scan order.
Label selections remain lazy and therefore return DuckPD DataFrame or
Series handles. Exact pandas return-type switching for df.loc[label]
depends on runtime index uniqueness and is intentionally deferred to a bounded
eager scalar/row API. MultiIndex exact and prefix keys are supported; ordered
label-list reindexing and cross-frame assignment alignment remain unsupported.
Module-level readers reuse a context-local implicit session, so independently
created helper frames can be combined. Explicit Session context managers are
still recommended when resource limits, database lifetime, or deterministic
cleanup matter.
Demos
Interactive notebooks and small runnable programs are available in demo/:
demo/DuckPD_Quickstart.ipynb— 5-minute quickstart on the Goodreads Books dataset.demo/DuckPD_Features_Walkthrough.ipynb— Deep dive into recent additions (remote cloud parquet, multi-table joins,.str/.dtaccessors,duckpd.concat, statistical reductions, and multi-column groupbys) using the AlphaDojo stock news dataset (~3.9M rows).
uv run python demo/basic_pipeline.py
uv run python demo/parquet_pipeline.py
uv run python demo/reduction_pipeline.py
uv run python demo/generate_market_data.py
uv run python demo/market_data_demo.py
See the benchmark results for performance and memory comparisons between DuckPD and pandas across 100 MB, 1 GB, and 5 GB datasets.
Development
uv sync --frozen --group dev
make check
make build
GNU Make is optional. The equivalent commands are:
uv run pytest
uv run ruff check .
uv run ruff format --check .
uv run pyright
uv build
See the documentation index for the implementation roadmap, architecture decisions, benchmarks, research, and changelog.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file duckpd-0.0.7.tar.gz.
File metadata
- Download URL: duckpd-0.0.7.tar.gz
- Upload date:
- Size: 1.5 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.21 {"installer":{"name":"uv","version":"0.11.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fcb8c97844e16c20e5446b2e5a0dc217eb421c78afbde9d2f65a14e8dfc790ab
|
|
| MD5 |
4b612a9e3d7265e12105047ae795e3e3
|
|
| BLAKE2b-256 |
f01c61d86922717b41b9beb3081da6118a12bb492d08a76b1ae3d14044c2d7d6
|
File details
Details for the file duckpd-0.0.7-py3-none-any.whl.
File metadata
- Download URL: duckpd-0.0.7-py3-none-any.whl
- Upload date:
- Size: 61.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.21 {"installer":{"name":"uv","version":"0.11.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b3831a3ce1834fdc6d9848c4d10efa1c169537b4cefac5618dc1fc33cd40cfc6
|
|
| MD5 |
b59cb4004fb289a24d84181508a1d512
|
|
| BLAKE2b-256 |
9cd13a39f065f598afefe88402432d449c66159f0135dc2100ef9e303122ed72
|