Skip to main content

Lightweight wrapper over DuckDB with convenience helpers for DuckLake.

Project description

ducklake-client

Lightweight Python helpers for opening DuckLake connections through DuckDB.

Install

pip install ducklake-client

Open a DuckLake connection

from ducklake_client import ColumnDef, DiskStorage, DuckDBCatalog, DuckLake

lake = DuckLake(
    catalog=DuckDBCatalog("metadata.ducklake"),
    storage=DiskStorage("data"),
)

try:
    lake.schema.create("main")
    lake.table.create(
        "items",
        id=ColumnDef("INTEGER", nullable=False),
        name=ColumnDef("VARCHAR"),
    )
    rows = lake.connection.sql("SELECT * FROM lake.main.items").fetchall()
finally:
    lake.close()

DuckLake opens the underlying DuckDB connection lazily on first use. The client installs and loads the DuckDB ducklake and parquet extensions, attaches the catalog as lake, and exposes the native DuckDB connection through connection.

Context manager usage

from ducklake_client import DiskStorage, DuckDBCatalog, DuckLake

with DuckLake(
    catalog=DuckDBCatalog("metadata.ducklake"),
    storage=DiskStorage("data"),
) as lake:
    lake.connection.execute("CREATE TABLE IF NOT EXISTS lake.main.events (id INTEGER)")
    lake.connection.execute("INSERT INTO lake.main.events VALUES (?)", [1])
    print(lake.connection.sql("SELECT count(*) FROM lake.main.events").fetchone())

Modules

DuckLake-specific helpers are grouped into modules. Native DuckDB behavior stays on lake.connection.

from ducklake_client import ColumnDef, DiskStorage, DuckDBCatalog, DuckLake

with DuckLake(
    catalog=DuckDBCatalog("metadata.ducklake"),
    storage=DiskStorage("data"),
) as lake:
    lake.schema.create("main")
    lake.table.create_from_csv(
        "nl_train_stations",
        "https://blobs.duckdb.org/nl_stations.csv",
    )
    lake.table.comment("nl_train_stations", "Dutch railway stations")
    lake.table.comment(
        "nl_train_stations",
        "Full station name",
        column_name="name_long",
    )
    tables = lake.table.list()
    views = lake.view.list()
    # Metadata-only by default: this does not scan the table data.
    info = lake.table.info("nl_train_stations")

    # Summary statistics and an exact row count both scan table data.
    detailed_info = lake.table.info(
        "nl_train_stations",
        include_summary=True,
        include_row_count=True,
    )

Nested and parameterized column types can be composed without writing native SQL:

from ducklake_client import ColumnDef, ListType, MapType, StructType

lake.table.create(
    "elements",
    attributes=ColumnDef(MapType("VARCHAR", "VARCHAR")),
    tags=ColumnDef(ListType("VARCHAR")),
    location=ColumnDef(StructType({"latitude": "DOUBLE", "longitude": "DOUBLE"})),
)

Append a microbatch from mapping records without constructing native SQL:

lake.table.append(
    "elements",
    [
        {"id": 1, "attributes": {"color": "green"}},
        {"id": 2, "attributes": {"color": "blue"}},
    ],
)

table.append matches columns by name and also accepts Arrow Table, RecordBatch, and RecordBatchReader objects. A DuckDBPyRelation created from the same lake.connection is accepted directly:

batch = lake.connection.sql("SELECT id, attributes FROM staging_elements")
lake.table.append("elements", batch)

Empty record batches are a no-op. Each non-empty append is issued as one insert statement. Arrow support does not require importing PyArrow in ducklake-client; the provided object must implement an Arrow C data interface.

Fenced and idempotent appends

DuckLake tables do not enforce unique or primary-key constraints. table.append therefore provides at-least-once writes by itself. Use a cooperative fence plus a transaction when a retried microbatch must not be appended twice:

batch_id = "atlas-elements-2026-07-11T00:00:00Z"

with lake.fence("atlas", "elements", batch_id, timeout=30):
    with lake.transaction():
        already_ingested = lake.sql_scalar(
            """
            SELECT EXISTS (
                SELECT 1
                FROM lake.main.ingestion_batches
                WHERE batch_id = $batch_id
            )
            """,
            batch_id=batch_id,
        )
        if not already_ingested:
            lake.table.append("elements", records)
            lake.table.append("ingestion_batches", [{"batch_id": batch_id}])

Create the marker table once during bootstrap. The fence serializes cooperating writers for the same keys; the transaction makes the marker and data append commit or roll back together. Every writer for that idempotency domain must use the same fence namespace and keys.

Fence guarantees depend on the catalog:

  • PostgresCatalog uses PostgreSQL session advisory locks and supports remote, cross-process writers. The lock is released if its connection closes or the process exits.
  • File-backed DuckDBCatalog and SqliteCatalog use advisory sidecar file locks plus process-local locks. They coordinate processes that share a filesystem with working advisory-lock semantics and all use this client.
  • In-memory DuckDB catalogs and platforms without advisory file locking fall back to process-local exclusion only.

Fencing is cooperative rather than a DuckLake uniqueness constraint. It cannot exclude writers that bypass lake.fence, and filesystem locks may not be reliable on every network filesystem. DuckLakeFenceTimeout reports acquisition timeout; other backend failures raise DuckLakeFenceError.

Ad hoc SQL as dict rows

For quick queries with named parameters (DuckDB $param syntax), use sql_dicts:

from ducklake_client import DiskStorage, DuckDBCatalog, DuckLake

with DuckLake(
    catalog=DuckDBCatalog("metadata.ducklake"),
    storage=DiskStorage("data"),
) as lake:
    rows = lake.sql_dicts("SELECT $n AS v", n=41)

The package includes pytz, which DuckDB requires when materializing TIMESTAMPTZ values as Python objects.

Transactions

Use transaction() to automatically begin, commit, or roll back a block on the native DuckDB connection.

from ducklake_client import ColumnDef, DiskStorage, DuckDBCatalog, DuckLake

with DuckLake(
    catalog=DuckDBCatalog("metadata.ducklake"),
    storage=DiskStorage("data"),
) as lake:
    with lake.transaction():
        lake.schema.create("main")
        lake.table.create(
            "items",
            id=ColumnDef("INTEGER", nullable=False),
            name=ColumnDef("VARCHAR"),
        )
        lake.connection.execute("INSERT INTO lake.main.items VALUES (?, ?)", [1, "example"])

Configuration

DuckLake requires explicit catalog and storage config objects:

from ducklake_client import DiskStorage, DuckDBCatalog, DuckLake

lake = DuckLake(
    catalog=DuckDBCatalog("metadata.ducklake"),
    storage=DiskStorage("data"),
)

You can pass DuckDB runtime settings with DuckDBConfig:

from ducklake_client import DiskStorage, DuckDBConfig, DuckDBCatalog, DuckLake

lake = DuckLake(
    catalog=DuckDBCatalog("metadata.ducklake"),
    storage=DiskStorage("data"),
    duckdb=DuckDBConfig(
        database=":memory:",
        threads=4,
        memory_limit="2GB",
    ),
)

Common DuckLake ATTACH settings have a typed configuration object:

from ducklake_client import (
    DiskStorage,
    DuckDBCatalog,
    DuckLake,
    DuckLakeAttachConfig,
)

lake = DuckLake(
    catalog=DuckDBCatalog("metadata.ducklake"),
    storage=DiskStorage("data"),
    attach=DuckLakeAttachConfig(
        data_inlining_row_limit=50,
        automatic_migration=False,
    ),
)

data_inlining_row_limit=0 disables data inlining for the connection. Typed options also cover create_if_not_exists, encrypted, and override_data_path. Less common or extension-version-specific parameters can still be passed through attach_options; when both forms specify the same key, attach_options takes precedence.

Catalogs can be DuckDBCatalog, SqliteCatalog, or PostgresCatalog. Storage can be DiskStorage or S3Storage.

Bootstrap behavior

schema.create, table.create, and table.create_from_csv are idempotent by default: each uses IF NOT EXISTS. Pass if_not_exists=False when an existing object should be reported as an error. An idempotent create does not reconcile or migrate the definition of an object that already exists.

S3 storage

Configure S3 and S3-compatible storage through S3Storage; native secret SQL is not required:

import os

from ducklake_client import DuckDBCatalog, DuckLake, S3Storage

lake = DuckLake(
    catalog=DuckDBCatalog("metadata.ducklake"),
    storage=S3Storage(
        bucket="atlas-data",
        prefix="ducklake",
        region="eu-west-1",
        key_id=os.environ.get("AWS_ACCESS_KEY_ID"),
        secret_access_key=os.environ.get("AWS_SECRET_ACCESS_KEY"),
        session_token=os.environ.get("AWS_SESSION_TOKEN"),
    ),
)

endpoint, url_style, and use_ssl support S3-compatible services. When no secret options are supplied, the client does not create a DuckDB secret; access then depends on credentials already available in the DuckDB environment.

For a local MinIO server, provide the endpoint without embedding credentials and use path-style URLs:

from ducklake_client import ColumnDef, DuckDBCatalog, DuckLake, S3Storage

with DuckLake(
    # The catalog is durable because this is a file, not ":memory:".
    catalog=DuckDBCatalog("state/atlas.ducklake"),
    storage=S3Storage(
        bucket="atlas-data",
        prefix="ducklake",
        endpoint="http://localhost:9000",
        key_id="minioadmin",
        secret_access_key="minioadmin",
        url_style="path",
        use_ssl=False,
    ),
) as lake:
    lake.schema.create("main")
    lake.table.create(
        "events",
        id=ColumnDef("BIGINT", nullable=False),
        payload=ColumnDef("JSON"),
    )

The S3 settings have these roles:

  • endpoint selects an S3-compatible service such as MinIO. A URL is accepted; the client passes its host and optional port to DuckDB.
  • url_style="path" produces bucket paths suitable for typical local MinIO setups. Use the service's required style in other environments.
  • use_ssl controls HTTPS independently of the endpoint spelling.
  • key_id, secret_access_key, and session_token create a temporary DuckDB secret managed by this client. If they and all other secret options are absent, no secret is created; configure credentials in DuckDB or its environment before accessing private objects.

S3 stores DuckLake data files, not the DuckLake catalog itself. Catalog durability is configured separately: DuckDBCatalog and SqliteCatalog persist metadata at their filesystem paths, while PostgresCatalog persists it in PostgreSQL. Keep the catalog on durable storage and back it up independently of the S3 bucket. A catalog path inside an ephemeral container will be lost even when its data files remain in S3.

Exceptions

The package exception hierarchy is public API:

  • DuckLakeError is the base package exception.
  • DuckLakeConfigError reports invalid client configuration or helper input.
  • DuckLakeConnectionError reports connection initialization failures.
  • DuckLakeFenceError reports cooperative fencing failures, with DuckLakeFenceTimeout for acquisition timeouts.
  • DuckLakeQueryError reports failures from client and module query helpers.

The original exception is retained as __cause__. Operations performed directly through lake.connection remain native DuckDB operations and raise DuckDB's own exceptions.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ducklake_client-0.8.0.tar.gz (27.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ducklake_client-0.8.0-py3-none-any.whl (35.6 kB view details)

Uploaded Python 3

File details

Details for the file ducklake_client-0.8.0.tar.gz.

File metadata

  • Download URL: ducklake_client-0.8.0.tar.gz
  • Upload date:
  • Size: 27.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for ducklake_client-0.8.0.tar.gz
Algorithm Hash digest
SHA256 53a1022360241a963151ea9b03d58e40a9a49d3771192665d4ea3e6e64b7c5b9
MD5 77255791f3bfd20e50b651860d888231
BLAKE2b-256 e3b40c088a80b637874c9b8ec11d32ab28f337cb1d3831e01d178a773866218a

See more details on using hashes here.

File details

Details for the file ducklake_client-0.8.0-py3-none-any.whl.

File metadata

File hashes

Hashes for ducklake_client-0.8.0-py3-none-any.whl
Algorithm Hash digest
SHA256 0926d5110f4f9cb167b07353877a0ff6f4a005f42bd3e7061d857767ff75ac7d
MD5 a509826195386023eb45785ba9342473
BLAKE2b-256 50dc39505d023c0f763a579eeee70ceac54bdf1751dc366fa6b8096369784e43

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page