Skip to main content

dlt-duckhaven

A dlt destination that loads data into DuckHaven-governed Iceberg tables through the DuckHaven API.

It is a staged-Parquet SQL destination built on the duckhaven-sql-connector, modeled on dlt's databricks/athena destinations: dlt writes Parquet, the destination stages it to the workspace object storage, then issues the load command (COPY/INSERT … SELECT read_parquet(...)) through the DuckHaven session API, which dispatches to an agent that writes the Iceberg table. Control goes through the API; bulk data goes through storage staging — the same split Databricks and MotherDuck use.

Status: alpha. The destination, its capabilities/type mapping, staged loads (append/replace/merge), and schema evolution are implemented and tested. Loading uses the session's presigned-URL stage (POST …/sql/sessions/{id}/staging-files); a DuckHaven with that endpoint and SQL_SESSIONS_ENABLED=true is required at runtime. See CHANGELOG.md.

Install

pip install dlt-duckhaven
# pip install "dlt-duckhaven[otel]"  # optional per-load-job spans

No storage SDK (s3fs/adlfs) is needed — staged files are uploaded to a presigned URL over plain HTTPS.

Configure

destination="duckhaven" reads the following (via dlt.yml, .dlt/secrets.toml, or env):

[destination.duckhaven]
host = "https://duckhaven.internal"
workspace = "analytics"
agent = "…-uuid-…"        # optional; omit to let the API pick compute
catalog = "raw"

[destination.duckhaven.credentials]
token = "dh_pat_…"        # a DuckHaven Personal Access Token (service account)
import dlt

pipeline = dlt.pipeline(
    pipeline_name="ingest",
    destination="duckhaven",
    dataset_name="analytics",   # the Iceberg schema/namespace within `catalog`
)
pipeline.run(my_resource)

dataset_name is the Iceberg schema; catalog.dataset_name.table is the fully-qualified relation. Auth is a DuckHaven service-account PAT — every statement is authorized and audited at the API, so a dlt load is fully governed.

Config reference

Field Required Description
host yes DuckHaven API base URL (https://…).
workspace yes Workspace slug the session opens in.
credentials.token yes A DuckHaven service-account PAT (dh_pat_…). May also be passed as the credentials value directly.
catalog recommended DuckHaven (Polaris) catalog that qualifies loaded tables.
agent no Explicit compute (an agent UUID); omit to let the API auto-pick.

If the DuckHaven deployment runs elastic compute and has scaled to zero, open_connection now waits while an agent starts — up to five minutes — instead of failing the load. A deployment that cannot start compute at all still fails immediately rather than waiting out that budget.

A DuckHaven server can restrict which agents a pipeline may target. A denial fails open_connection, so the load stops before any job runs. "Agent not found" on an agent you know exists means it is restricted and you hold no grant — such an agent is hidden rather than reported as forbidden, so it reads exactly like a deleted one. "requires the 'use' tier" means it is visible to you but your grant is too low. Omitting agent narrows auto-pick to agents you may use, so it can report no agent available rather than falling back to one you cannot run on.

Write dispositions

  • append — stages Parquet and loads it into the table.
  • replaceinsert-from-staging (default) swaps via a staging dataset; truncate-and-insert clears the table first. Iceberg has no cheap TRUNCATE, so truncation is a DELETE FROM … WHERE 1=1.
  • merge — delete-insert upsert against a staging dataset; set a primary_key (and/or merge_key). A second run with the same key updates in place rather than appending.

Tables are stored as Iceberg (via the attached Polaris catalog); the type mapper constrains dlt types to Iceberg-safe DuckDB types (JSON → VARCHAR, microsecond timestamps, no 128-bit integers). Schema evolution (new columns on a later run) is applied with ALTER TABLE … ADD COLUMN, one statement per new column — DuckDB accepts only one ALTER action per statement, so several columns arriving in the same run are added one after another rather than in a single combined statement.

Concurrent loads

dlt loads in parallel by default, so several jobs can write to the same table at once — one resource whose data spans more than one Parquet file is enough. Iceberg settles those races with optimistic concurrency: one writer commits and the others are told to start again from the refreshed table metadata. The destination retries a losing commit for you, with jittered backoff, so this is invisible in normal operation and there is no need to serialize the load step with LOAD__WORKERS=1. A retry is safe because a rejected commit publishes nothing — none of its rows are visible, so re-running cannot duplicate them. Under sustained contention the retries are bounded and dlt's own job retry takes over.

Reading values back

DuckHaven returns rows as JSON, so temporal values arrive as ISO-8601 strings and are converted back to datetime on the way out. Which columns get converted is decided from the column types the server reports, so a VARCHAR column that happens to hold an ISO-8601-looking string stays the string it is. Against a server that reports no column types the destination falls back to recognizing datetimes by their shape, which can misread such a VARCHAR — the reason the typed path exists.

Observability

With the otel extra, each load job and staging upload emits an OpenTelemetry span (dlt_duckhaven.load_job, dlt_duckhaven.stage). Because the connector injects a W3C traceparent on every request, those spans parent the connector's HTTP spans and the server trace, so a dlt load traces end-to-end (client → API → agent). Without the extra the instrumentation is a no-op.

Governance & staging

Bulk Parquet is staged to the workspace object storage via a presigned-URL stage — the session's staging-files endpoint returns a short-lived put_url (upload) and get_url (read) per file, scoped to the session's staging prefix. The client uploads to put_url with a plain HTTP PUT; the agent reads get_url over httpfs. No storage credentials live on the client or the agent — all backend-specific signing (S3/MinIO SigV4, Azure SAS) happens in the API, which already owns the storage integration. Every load statement is authorized and audited at the API, so a dlt load is fully governed.

License

Apache-2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dlt_duckhaven-0.5.0.tar.gz (23.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dlt_duckhaven-0.5.0-py3-none-any.whl (25.7 kB view details)

Uploaded Python 3

File details

Details for the file dlt_duckhaven-0.5.0.tar.gz.

File metadata

  • Download URL: dlt_duckhaven-0.5.0.tar.gz
  • Upload date:
  • Size: 23.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for dlt_duckhaven-0.5.0.tar.gz
Algorithm Hash digest
SHA256 e23b055991163a3d08cac83022d7bd26f0b22ced6a67e01f88a77d42f31fa7a7
MD5 a9432be9d97ee48762f60743f81932b4
BLAKE2b-256 1a7fd2369a3a7e15f7b5cb8e62d522ea5511a00a27a070b8758934b1e35f9339

See more details on using hashes here.

Provenance

The following attestation bundles were made for dlt_duckhaven-0.5.0.tar.gz:

Publisher: release.yml on tamasmrtn/duckhaven-clients

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file dlt_duckhaven-0.5.0-py3-none-any.whl.

File metadata

  • Download URL: dlt_duckhaven-0.5.0-py3-none-any.whl
  • Upload date:
  • Size: 25.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for dlt_duckhaven-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 a4e91e9a8ada6c20427e257c5e6ac0f49774b778f1fa9df70a99473e51306138
MD5 aca795ee7fea92d9ffccfd520f80519d
BLAKE2b-256 29a958677a79b20a23061c8a746dfef554b9f7c8898f7a7bf734c8263077ca49

See more details on using hashes here.

Provenance

The following attestation bundles were made for dlt_duckhaven-0.5.0-py3-none-any.whl:

Publisher: release.yml on tamasmrtn/duckhaven-clients

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.6.0

2 files

0.5.1

2 files

This release

0.5.0 This release

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page