Skip to main content

dlt-duckhaven

A dlt destination that loads data into DuckHaven-governed Iceberg tables through the DuckHaven API.

It is a staged-Parquet SQL destination built on the duckhaven-sql-connector, modeled on dlt's databricks/athena destinations: dlt writes Parquet, the destination stages it to the workspace object storage, then issues the load command (COPY/INSERT … SELECT read_parquet(...)) through the DuckHaven session API, which dispatches to an agent that writes the Iceberg table. Control goes through the API; bulk data goes through storage staging — the same split Databricks and MotherDuck use.

Status: alpha. The destination, its capabilities/type mapping, staged loads (append/replace/merge), and schema evolution are implemented and tested. Loading uses the session's presigned-URL stage (POST …/sql/sessions/{id}/staging-files); a DuckHaven with that endpoint and SQL_SESSIONS_ENABLED=true is required at runtime. See CHANGELOG.md.

Install

pip install dlt-duckhaven
# pip install "dlt-duckhaven[otel]"  # optional per-load-job spans

No storage SDK (s3fs/adlfs) is needed — staged files are uploaded to a presigned URL over plain HTTPS.

Configure

destination="duckhaven" reads the following (via dlt.yml, .dlt/secrets.toml, or env):

[destination.duckhaven]
host = "https://duckhaven.internal"
workspace = "analytics"
agent = "…-uuid-…"        # optional; omit to let the API pick compute
catalog = "raw"

[destination.duckhaven.credentials]
token = "dh_pat_…"        # a DuckHaven Personal Access Token (service account)
import dlt

pipeline = dlt.pipeline(
    pipeline_name="ingest",
    destination="duckhaven",
    dataset_name="analytics",  # the Iceberg schema/namespace within `catalog`
)
pipeline.run(my_resource)

dataset_name is the Iceberg schema; catalog.dataset_name.table is the fully-qualified relation. Auth is a DuckHaven service-account PAT — every statement is authorized and audited at the API, so a dlt load is fully governed.

Config reference

Field Required Description
host yes DuckHaven API base URL (https://…).
workspace yes Workspace slug the session opens in.
credentials.token yes A DuckHaven service-account PAT (dh_pat_…). May also be passed as the credentials value directly.
catalog recommended DuckHaven (Polaris) catalog that qualifies loaded tables.
agent no Explicit compute (an agent UUID); omit to let the API auto-pick.

If the DuckHaven deployment runs elastic compute and has scaled to zero, open_connection now waits while an agent starts — up to five minutes — instead of failing the load. A deployment that cannot start compute at all still fails immediately rather than waiting out that budget.

A DuckHaven server can restrict which agents a pipeline may target. A denial fails open_connection, so the load stops before any job runs. "Agent not found" on an agent you know exists means it is restricted and you hold no grant — such an agent is hidden rather than reported as forbidden, so it reads exactly like a deleted one. "requires the 'use' tier" means it is visible to you but your grant is too low. Omitting agent narrows auto-pick to agents you may use, so it can report no agent available rather than falling back to one you cannot run on.

Write dispositions

  • append — stages Parquet and loads it into the table.
  • replaceinsert-from-staging (default) swaps via a staging dataset; truncate-and-insert clears the table first. Iceberg has no cheap TRUNCATE, so truncation is a DELETE FROM … WHERE 1=1.
  • merge — delete-insert upsert against a staging dataset; set a primary_key (and/or merge_key). A second run with the same key updates in place rather than appending.

Tables are stored as Iceberg (via the attached Polaris catalog); the type mapper constrains dlt types to Iceberg-safe DuckDB types (JSON → VARCHAR, microsecond timestamps, no 128-bit integers). Schema evolution (new columns on a later run) is applied with ALTER TABLE … ADD COLUMN, one statement per new column — DuckDB accepts only one ALTER action per statement, so several columns arriving in the same run are added one after another rather than in a single combined statement.

Concurrent loads

dlt loads in parallel by default, so several jobs can write to the same table at once — one resource whose data spans more than one Parquet file is enough. Iceberg settles those races with optimistic concurrency: one writer commits and the others are told to start again from the refreshed table metadata. The destination retries a losing commit for you, with jittered backoff, so this is invisible in normal operation and there is no need to serialize the load step with LOAD__WORKERS=1. A retry is safe because a rejected commit publishes nothing — none of its rows are visible, so re-running cannot duplicate them. Under sustained contention the retries are bounded and dlt's own job retry takes over.

Reading values back

DuckHaven returns rows as JSON, so temporal values arrive as ISO-8601 strings and are converted back to datetime on the way out. Which columns get converted is decided from the column types the server reports, so a VARCHAR column that happens to hold an ISO-8601-looking string stays the string it is. Against a server that reports no column types the destination falls back to recognizing datetimes by their shape, which can misread such a VARCHAR — the reason the typed path exists.

Observability

With the otel extra, each load job and staging upload emits an OpenTelemetry span (dlt_duckhaven.load_job, dlt_duckhaven.stage). Because the connector injects a W3C traceparent on every request, those spans parent the connector's HTTP spans and the server trace, so a dlt load traces end-to-end (client → API → agent). Without the extra the instrumentation is a no-op.

Governance & staging

Bulk Parquet is staged to the workspace object storage via a presigned-URL stage — the session's staging-files endpoint returns a short-lived put_url (upload) and get_url (read) per file, scoped to the session's staging prefix. The client uploads to put_url with a plain HTTP PUT; the agent reads get_url over httpfs. No storage credentials live on the client or the agent — all backend-specific signing (S3/MinIO SigV4, Azure SAS) happens in the API, which already owns the storage integration. Every load statement is authorized and audited at the API, so a dlt load is fully governed.

License

Apache-2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dlt_duckhaven-0.6.0.tar.gz (24.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dlt_duckhaven-0.6.0-py3-none-any.whl (25.7 kB view details)

Uploaded Python 3

File details

Details for the file dlt_duckhaven-0.6.0.tar.gz.

File metadata

  • Download URL: dlt_duckhaven-0.6.0.tar.gz
  • Upload date:
  • Size: 24.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for dlt_duckhaven-0.6.0.tar.gz
Algorithm Hash digest
SHA256 9d19c073e4aab79f8cc4e3fab1d0ae3613dacf073acd99b01863c3e47cbf6ef1
MD5 423ff836377f96c00454d00d1e896405
BLAKE2b-256 0e69ce78cbb589f3ce17df31255e6dcfd859ea8e6717c6692d3233b2e7348efd

See more details on using hashes here.

Provenance

The following attestation bundles were made for dlt_duckhaven-0.6.0.tar.gz:

Publisher: release.yml on tamasmrtn/duckhaven-clients

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file dlt_duckhaven-0.6.0-py3-none-any.whl.

File metadata

  • Download URL: dlt_duckhaven-0.6.0-py3-none-any.whl
  • Upload date:
  • Size: 25.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for dlt_duckhaven-0.6.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ed172c518b2403f7cfd1ddbe4d143ed71eaa3d7cc838a9600b027c422b8af674
MD5 694e3af707cbc3ff360f43e7645742a4
BLAKE2b-256 ede0c6ac34f8c6e440f9a3baaad4aea68012649e3542e578ac7d697df34664c8

See more details on using hashes here.

Provenance

The following attestation bundles were made for dlt_duckhaven-0.6.0-py3-none-any.whl:

Publisher: release.yml on tamasmrtn/duckhaven-clients

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 files

0.5.1

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page