Skip to main content

[!WARNING] DuckPD is a work in progress and is not yet recommended for production-critical workloads. The API and supported pandas semantics may change between 0.x releases, and many pandas operations are intentionally unsupported. Validate results and resource behavior for each intended workload before adopting it.

DuckPD mascot - a duck dressed as a panda

DuckPD 🦆❤️🐼

DuckPD is DuckDB dressed as a pandas DataFrame.

DuckPD is a lazy DataFrame library with a pandas-shaped API and DuckDB as its execution engine. The goal is to make working with DuckDB feel familiar to pandas users while preserving the performance, scalability, and query-optimization advantages of DuckDB.

Where practical, DuckPD aims to match pandas APIs and semantics closely enough that existing pandas knowledge — and eventually a large amount of pandas-oriented code — transfers naturally. It does not, however, aim to reproduce pandas by sacrificing the properties that make DuckDB valuable.

Project directives

These principles define the direction of DuckPD and should guide API and implementation decisions:

  1. Pandas-shaped, DuckDB-native. The public API should feel like pandas, but operations should map naturally onto DuckDB's relational and vectorized execution model.

  2. Stay lazy by default. Transformations should build a query plan rather than execute immediately. Execution should happen only at clear and intentional boundaries such as collect(), head(), Arrow conversion, or file output.

  3. Never silently fall back to pandas. Unsupported operations should fail explicitly rather than unexpectedly materializing an entire dataset into memory. Users should always be able to reason about where computation happens.

  4. Push work into DuckDB. Filtering, projection, joins, aggregation, sorting, expressions, and other supported operations should be translated into DuckDB operations whenever possible so DuckDB can optimize the complete query.

  5. Preserve pandas semantics where we claim compatibility. API similarity alone is not enough. Supported operations should match pandas behavior as closely as practical, including edge cases around nulls, indexes, dtypes, grouping, and column behavior.

  6. Correctness before coverage. It is better to support a smaller pandas surface correctly than to advertise broad compatibility backed by incomplete semantics, hidden fallbacks, or surprising execution behavior.

  7. Make execution visible and predictable. Users should be able to understand when data is scanned, materialized, transferred, or written. Laziness must be a useful property, not hidden magic.

  8. Exploit the ecosystem boundaries. DuckPD should interoperate cleanly with pandas, Arrow, Parquet, SQL, and DuckDB itself. Crossing those boundaries should be explicit and inexpensive wherever the underlying systems allow it.

The long-term ambition is broad pandas API coverage where those APIs can be implemented without violating these directives. Compatibility is the interface; DuckDB-native execution is the foundation.

Current capabilities

  • Lazy pandas, Arrow, Parquet, DuckDB table, and read-only SQL sources.
  • Column selection, boolean filtering, arithmetic expressions, assign, sort_values, limit, and distinct/drop_duplicates deduplication.
  • Relational DataFrame joins (merge) supporting inner, left, right, outer, and cross with column collision suffix management.
  • Multi-DataFrame row-wise concatenation (duckpd.concat) with schema alignment and null-padding.
  • Vectorized .str (e.g. upper, lower, strip, len, contains, replace) and .dt (e.g. year, month, day, hour, minute, second, strftime, to_period) accessor pipelines.
  • Multi-column groupby() supporting eager and lazy agg(), sum(), mean(), min(), max(), std(), var(), and count().
  • Eager DataFrame and Series reductions: count, size, sum, mean, min, max, std, var, median, quantile, any, and all over numeric and boolean data, including skipna, min_count, and DataFrame numeric_only support.
  • Explicit lazy indexes with set_index()/reset_index() and source index=/order_by= declarations, including exact and partial MultiIndex .loc selection.
  • Stable snapshot order for pandas and Arrow inputs, used for deterministic positional operations, duplicate retention, ranking, and top-N ties.
  • Context-local implicit sessions, allowing frames created by separate module-level helpers to participate in the same lazy plan.
  • Explicit pandas collection, bounded head, Arrow tables and record batches, physical plan inspection (explain), and direct zero-copy Parquet writes.

Supported pandas API Coverage

DuckPD maps pandas semantics directly to DuckDB's vectorized analytical engine:

API Category Supported Methods & Operations Execution Model
I/O & Data Loading read_parquet(), read_sql(), from_pandas(), from_arrow(), sql(), connect() Lazy (scans metadata / registers source)
Transformations & Projections df[cols], df[bool_filter], assign(), sort_values(), limit(), drop_duplicates(), set_index(), reset_index() Lazy (appends to logical query graph)
Joins & Merges merge() (inner, left, right, outer, cross, custom suffixes) Lazy (relational hash join)
Concatenation duckpd.concat() (multi-frame row union, schema alignment, null padding) Lazy (union with projection padding)
String Accessor (.str) upper(), lower(), strip(), len(), startswith(), endswith(), contains(), replace() Lazy (DuckDB SQL functions)
Datetime Accessor (.dt) year, month, day, hour, minute, second, strftime(), to_period() Lazy (DuckDB timestamp extractors)
GroupBy Aggregations groupby().agg(), .sum(), .mean(), .min(), .max(), .std(), .var(), .count() (as_index=True/False) Lazy for .agg(), Eager for reductions
Statistical Reductions sum(), mean(), min(), max(), count(), size, std(), var(), median(), quantile(), any(), all() Eager (single aggregate SQL pushdown)
Collection & Output collect(), head(n), explain(), write_parquet(), to_arrow_table(), to_arrow_batches() Explicit Execution Boundary

Example

import duckpd as pd

orders = pd.read_parquet("orders/*.parquet")

result = (
    orders[orders["status"] == "paid"]
    .assign(net=lambda frame: frame["amount"] - frame["refund_amount"])
    .sort_values("net", ascending=False)[["order_id", "net"]]
    .limit(100)
)

print(result.explain())
preview = result.head(10)
result.write_parquet("largest-paid-orders.parquet")
pandas_result = result.collect()

Transformations above are lazy. explain(), head(), collect(), Arrow output, and file output are explicit execution boundaries. limit() stays lazy while head() returns a bounded pandas preview.

Ordering, indexing, and sessions

Pandas and Arrow inputs are snapshots with a stable source row order. DuckPD tracks that order with hidden relational metadata so operations such as .iloc, drop_duplicates(keep=...), rank(method="first"), and top-N tie selection remain deterministic without exposing a synthetic pandas index.

Parquet, CSV, SQL, and DuckDB table scans remain unordered unless order_by= is provided. Ordering-sensitive operations fail with UnorderedOperationError rather than relying on accidental scan order.

Label selections remain lazy and therefore return DuckPD DataFrame or Series handles. Exact pandas return-type switching for df.loc[label] depends on runtime index uniqueness and is intentionally deferred to a bounded eager scalar/row API. MultiIndex exact and prefix keys are supported; ordered label-list reindexing and cross-frame assignment alignment remain unsupported.

Module-level readers reuse a context-local implicit session, so independently created helper frames can be combined. Explicit Session context managers are still recommended when resource limits, database lifetime, or deterministic cleanup matter.

Demos

Interactive notebooks and small runnable programs are available in demo/:

  • demo/DuckPD_Quickstart.ipynb — 5-minute quickstart on the Goodreads Books dataset.
  • demo/DuckPD_Features_Walkthrough.ipynb — Deep dive into recent additions (remote cloud parquet, multi-table joins, .str/.dt accessors, duckpd.concat, statistical reductions, and multi-column groupbys) using the AlphaDojo stock news dataset (~3.9M rows).
uv run python demo/basic_pipeline.py
uv run python demo/parquet_pipeline.py
uv run python demo/reduction_pipeline.py
uv run python demo/generate_market_data.py
uv run python demo/market_data_demo.py

See the benchmark results for performance and memory comparisons between DuckPD and pandas across 100 MB, 1 GB, and 5 GB datasets.

Development

uv sync --frozen --group dev
make check
make build

GNU Make is optional. The equivalent commands are:

uv run pytest
uv run ruff check .
uv run ruff format --check .
uv run pyright
uv build

See the documentation index for the implementation roadmap, architecture decisions, benchmarks, research, and changelog.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

duckpd-0.0.7.tar.gz (1.5 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

duckpd-0.0.7-py3-none-any.whl (61.1 kB view details)

Uploaded Python 3

File details

Details for the file duckpd-0.0.7.tar.gz.

File metadata

  • Download URL: duckpd-0.0.7.tar.gz
  • Upload date:
  • Size: 1.5 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.21 {"installer":{"name":"uv","version":"0.11.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for duckpd-0.0.7.tar.gz
Algorithm Hash digest
SHA256 fcb8c97844e16c20e5446b2e5a0dc217eb421c78afbde9d2f65a14e8dfc790ab
MD5 4b612a9e3d7265e12105047ae795e3e3
BLAKE2b-256 f01c61d86922717b41b9beb3081da6118a12bb492d08a76b1ae3d14044c2d7d6

See more details on using hashes here.

File details

Details for the file duckpd-0.0.7-py3-none-any.whl.

File metadata

  • Download URL: duckpd-0.0.7-py3-none-any.whl
  • Upload date:
  • Size: 61.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.21 {"installer":{"name":"uv","version":"0.11.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for duckpd-0.0.7-py3-none-any.whl
Algorithm Hash digest
SHA256 b3831a3ce1834fdc6d9848c4d10efa1c169537b4cefac5618dc1fc33cd40cfc6
MD5 b59cb4004fb289a24d84181508a1d512
BLAKE2b-256 9cd13a39f065f598afefe88402432d449c66159f0135dc2100ef9e303122ed72

See more details on using hashes here.

Release history Release notifications | RSS feed

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

This release

0.0.7 This release

2 files

0.0.5

2 files

0.0.4

2 files

0.0.3

2 files

0.0.2

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page