Lightweight wrapper over DuckDB with convenience helpers for DuckLake.
Project description
ducklake-client
Lightweight Python helpers for opening DuckLake connections through DuckDB.
Install
pip install ducklake-client
Open a DuckLake connection
from ducklake_client import ColumnDef, DiskStorage, DuckDBCatalog, DuckLake
lake = DuckLake(
catalog=DuckDBCatalog("metadata.ducklake"),
storage=DiskStorage("data"),
)
try:
lake.schema.create("main")
lake.table.create(
"items",
id=ColumnDef("INTEGER", nullable=False),
name=ColumnDef("VARCHAR"),
)
rows = lake.connection.sql("SELECT * FROM lake.main.items").fetchall()
finally:
lake.close()
DuckLake opens the underlying DuckDB connection lazily on first use. The client installs and loads the DuckDB ducklake and parquet extensions, attaches the catalog as lake, and exposes the native DuckDB connection through connection.
Context manager usage
from ducklake_client import DiskStorage, DuckDBCatalog, DuckLake
with DuckLake(
catalog=DuckDBCatalog("metadata.ducklake"),
storage=DiskStorage("data"),
) as lake:
lake.connection.execute("CREATE TABLE IF NOT EXISTS lake.main.events (id INTEGER)")
lake.connection.execute("INSERT INTO lake.main.events VALUES (?)", [1])
print(lake.connection.sql("SELECT count(*) FROM lake.main.events").fetchone())
Modules
DuckLake-specific helpers are grouped into modules. Native DuckDB behavior stays on lake.connection.
from ducklake_client import ColumnDef, DiskStorage, DuckDBCatalog, DuckLake
with DuckLake(
catalog=DuckDBCatalog("metadata.ducklake"),
storage=DiskStorage("data"),
) as lake:
lake.schema.create("main")
lake.table.create_from_csv(
"nl_train_stations",
"https://blobs.duckdb.org/nl_stations.csv",
)
lake.table.comment("nl_train_stations", "Dutch railway stations")
lake.table.comment(
"nl_train_stations",
"Full station name",
column_name="name_long",
)
tables = lake.table.list()
views = lake.view.list()
# Metadata-only by default: this does not scan the table data.
info = lake.table.info("nl_train_stations")
# Summary statistics and an exact row count both scan table data.
detailed_info = lake.table.info(
"nl_train_stations",
include_summary=True,
include_row_count=True,
)
Nested and parameterized column types can be composed without writing native SQL:
from ducklake_client import ColumnDef, ListType, MapType, StructType
lake.table.create(
"elements",
attributes=ColumnDef(MapType("VARCHAR", "VARCHAR")),
tags=ColumnDef(ListType("VARCHAR")),
location=ColumnDef(StructType({"latitude": "DOUBLE", "longitude": "DOUBLE"})),
)
Ad hoc SQL as dict rows
For quick queries with named parameters (DuckDB $param syntax), use sql_dicts:
from ducklake_client import DiskStorage, DuckDBCatalog, DuckLake
with DuckLake(
catalog=DuckDBCatalog("metadata.ducklake"),
storage=DiskStorage("data"),
) as lake:
rows = lake.sql_dicts("SELECT $n AS v", n=41)
Transactions
Use transaction() to automatically begin, commit, or roll back a block on the native DuckDB connection.
from ducklake_client import ColumnDef, DiskStorage, DuckDBCatalog, DuckLake
with DuckLake(
catalog=DuckDBCatalog("metadata.ducklake"),
storage=DiskStorage("data"),
) as lake:
with lake.transaction():
lake.schema.create("main")
lake.table.create(
"items",
id=ColumnDef("INTEGER", nullable=False),
name=ColumnDef("VARCHAR"),
)
lake.connection.execute("INSERT INTO lake.main.items VALUES (?, ?)", [1, "example"])
Configuration
DuckLake requires explicit catalog and storage config objects:
from ducklake_client import DiskStorage, DuckDBCatalog, DuckLake
lake = DuckLake(
catalog=DuckDBCatalog("metadata.ducklake"),
storage=DiskStorage("data"),
)
You can pass DuckDB runtime settings with DuckDBConfig:
from ducklake_client import DiskStorage, DuckDBConfig, DuckDBCatalog, DuckLake
lake = DuckLake(
catalog=DuckDBCatalog("metadata.ducklake"),
storage=DiskStorage("data"),
duckdb=DuckDBConfig(
database=":memory:",
threads=4,
memory_limit="2GB",
),
)
Common DuckLake ATTACH settings have a typed configuration object:
from ducklake_client import (
DiskStorage,
DuckDBCatalog,
DuckLake,
DuckLakeAttachConfig,
)
lake = DuckLake(
catalog=DuckDBCatalog("metadata.ducklake"),
storage=DiskStorage("data"),
attach=DuckLakeAttachConfig(
data_inlining_row_limit=50,
automatic_migration=False,
),
)
data_inlining_row_limit=0 disables data inlining for the connection. Typed
options also cover create_if_not_exists, encrypted, and
override_data_path. Less common or extension-version-specific parameters can
still be passed through attach_options; when both forms specify the same key,
attach_options takes precedence.
Catalogs can be DuckDBCatalog, SqliteCatalog, or PostgresCatalog. Storage can be DiskStorage or S3Storage.
Bootstrap behavior
schema.create, table.create, and table.create_from_csv are idempotent by
default: each uses IF NOT EXISTS. Pass if_not_exists=False when an existing
object should be reported as an error. An idempotent create does not reconcile
or migrate the definition of an object that already exists.
S3 storage
Configure S3 and S3-compatible storage through S3Storage; native secret SQL is
not required:
import os
from ducklake_client import DuckDBCatalog, DuckLake, S3Storage
lake = DuckLake(
catalog=DuckDBCatalog("metadata.ducklake"),
storage=S3Storage(
bucket="atlas-data",
prefix="ducklake",
region="eu-west-1",
key_id=os.environ.get("AWS_ACCESS_KEY_ID"),
secret_access_key=os.environ.get("AWS_SECRET_ACCESS_KEY"),
session_token=os.environ.get("AWS_SESSION_TOKEN"),
),
)
endpoint, url_style, and use_ssl support S3-compatible services. When no
secret options are supplied, the client does not create a DuckDB secret; access
then depends on credentials already available in the DuckDB environment.
For a local MinIO server, provide the endpoint without embedding credentials and use path-style URLs:
from ducklake_client import ColumnDef, DuckDBCatalog, DuckLake, S3Storage
with DuckLake(
# The catalog is durable because this is a file, not ":memory:".
catalog=DuckDBCatalog("state/atlas.ducklake"),
storage=S3Storage(
bucket="atlas-data",
prefix="ducklake",
endpoint="http://localhost:9000",
key_id="minioadmin",
secret_access_key="minioadmin",
url_style="path",
use_ssl=False,
),
) as lake:
lake.schema.create("main")
lake.table.create(
"events",
id=ColumnDef("BIGINT", nullable=False),
payload=ColumnDef("JSON"),
)
The S3 settings have these roles:
endpointselects an S3-compatible service such as MinIO. A URL is accepted; the client passes its host and optional port to DuckDB.url_style="path"produces bucket paths suitable for typical local MinIO setups. Use the service's required style in other environments.use_sslcontrols HTTPS independently of the endpoint spelling.key_id,secret_access_key, andsession_tokencreate a temporary DuckDB secret managed by this client. If they and all other secret options are absent, no secret is created; configure credentials in DuckDB or its environment before accessing private objects.
S3 stores DuckLake data files, not the DuckLake catalog itself. Catalog durability
is configured separately: DuckDBCatalog and SqliteCatalog persist metadata at
their filesystem paths, while PostgresCatalog persists it in PostgreSQL. Keep
the catalog on durable storage and back it up independently of the S3 bucket. A
catalog path inside an ephemeral container will be lost even when its data files
remain in S3.
Exceptions
The package exception hierarchy is public API:
DuckLakeErroris the base package exception.DuckLakeConfigErrorreports invalid client configuration or helper input.DuckLakeConnectionErrorreports connection initialization failures.DuckLakeQueryErrorreports failures from client and module query helpers.
The original exception is retained as __cause__. Operations performed directly
through lake.connection remain native DuckDB operations and raise DuckDB's own
exceptions.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ducklake_client-0.6.0.tar.gz.
File metadata
- Download URL: ducklake_client-0.6.0.tar.gz
- Upload date:
- Size: 19.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6be8a85f7f84d5bfa08d7fedefc0b0dca93256d7629ffdfdecada7c36e7120b4
|
|
| MD5 |
123710b6cfacb3d3af66ee8b3d16d1d3
|
|
| BLAKE2b-256 |
4f7369e734d2f3380604a4783c2cfaf55b2e56fb1906e05ccdd272313817151a
|
File details
Details for the file ducklake_client-0.6.0-py3-none-any.whl.
File metadata
- Download URL: ducklake_client-0.6.0-py3-none-any.whl
- Upload date:
- Size: 30.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3c7285faee2a2032d4d9b045ce7a00359f3e9265d4a4b8fae49cdb09e5b65f97
|
|
| MD5 |
c3713c9dd8b435118bb1387b3965b5f3
|
|
| BLAKE2b-256 |
b8e09f741f370cdb4955555c118d173f527068bcf161018e311a3383b58c9dbb
|