Skip to main content

dlt-duckhaven

A dlt destination that loads data into DuckHaven-governed Iceberg tables through the DuckHaven API.

It is a staged-Parquet SQL destination built on the duckhaven-sql-connector, modeled on dlt's databricks/athena destinations: dlt writes Parquet, the destination stages it to the workspace object storage, then issues the load command (COPY/INSERT … SELECT read_parquet(...)) through the DuckHaven session API, which dispatches to an agent that writes the Iceberg table. Control goes through the API; bulk data goes through storage staging — the same split Databricks and MotherDuck use.

Status: alpha. The destination, its capabilities/type mapping, staged loads (append/replace/merge), and schema evolution are implemented and tested. Loading uses the session's presigned-URL stage (POST …/sql/sessions/{id}/staging-files); a DuckHaven with that endpoint and SQL_SESSIONS_ENABLED=true is required at runtime. See CHANGELOG.md.

Install

pip install dlt-duckhaven
# pip install "dlt-duckhaven[otel]"  # optional per-load-job spans

No storage SDK (s3fs/adlfs) is needed — staged files are uploaded to a presigned URL over plain HTTPS.

Configure

destination="duckhaven" reads the following (via dlt.yml, .dlt/secrets.toml, or env):

[destination.duckhaven]
host = "https://duckhaven.internal"
workspace = "analytics"
agent = "…-uuid-…"        # optional; omit to let the API pick compute
catalog = "raw"

[destination.duckhaven.credentials]
token = "dh_pat_…"        # a DuckHaven Personal Access Token (service account)
import dlt

pipeline = dlt.pipeline(
    pipeline_name="ingest",
    destination="duckhaven",
    dataset_name="analytics",   # the Iceberg schema/namespace within `catalog`
)
pipeline.run(my_resource)

dataset_name is the Iceberg schema; catalog.dataset_name.table is the fully-qualified relation. Auth is a DuckHaven service-account PAT — every statement is authorized and audited at the API, so a dlt load is fully governed.

Config reference

Field Required Description
host yes DuckHaven API base URL (https://…).
workspace yes Workspace slug the session opens in.
credentials.token yes A DuckHaven service-account PAT (dh_pat_…). May also be passed as the credentials value directly.
catalog recommended DuckHaven (Polaris) catalog that qualifies loaded tables.
agent no Explicit compute (an agent UUID); omit to let the API auto-pick.

If the DuckHaven deployment runs elastic compute and has scaled to zero, open_connection now waits while an agent starts — up to five minutes — instead of failing the load. A deployment that cannot start compute at all still fails immediately rather than waiting out that budget.

A DuckHaven server can restrict which agents a pipeline may target. A denial fails open_connection, so the load stops before any job runs. "Agent not found" on an agent you know exists means it is restricted and you hold no grant — such an agent is hidden rather than reported as forbidden, so it reads exactly like a deleted one. "requires the 'use' tier" means it is visible to you but your grant is too low. Omitting agent narrows auto-pick to agents you may use, so it can report no agent available rather than falling back to one you cannot run on.

Write dispositions

  • append — stages Parquet and loads it into the table.
  • replaceinsert-from-staging (default) swaps via a staging dataset; truncate-and-insert clears the table first. Iceberg has no cheap TRUNCATE, so truncation is a DELETE FROM … WHERE 1=1.
  • merge — delete-insert upsert against a staging dataset; set a primary_key (and/or merge_key). A second run with the same key updates in place rather than appending.

Tables are stored as Iceberg (via the attached Polaris catalog); the type mapper constrains dlt types to Iceberg-safe DuckDB types (JSON → VARCHAR, microsecond timestamps, no 128-bit integers). Schema evolution (new columns on a later run) is applied with ALTER TABLE … ADD COLUMN.

Reading values back

DuckHaven returns rows as JSON, so temporal values arrive as ISO-8601 strings and are converted back to datetime on the way out. Which columns get converted is decided from the column types the server reports, so a VARCHAR column that happens to hold an ISO-8601-looking string stays the string it is. Against a server that reports no column types the destination falls back to recognizing datetimes by their shape, which can misread such a VARCHAR — the reason the typed path exists.

Observability

With the otel extra, each load job and staging upload emits an OpenTelemetry span (dlt_duckhaven.load_job, dlt_duckhaven.stage). Because the connector injects a W3C traceparent on every request, those spans parent the connector's HTTP spans and the server trace, so a dlt load traces end-to-end (client → API → agent). Without the extra the instrumentation is a no-op.

Governance & staging

Bulk Parquet is staged to the workspace object storage via a presigned-URL stage — the session's staging-files endpoint returns a short-lived put_url (upload) and get_url (read) per file, scoped to the session's staging prefix. The client uploads to put_url with a plain HTTP PUT; the agent reads get_url over httpfs. No storage credentials live on the client or the agent — all backend-specific signing (S3/MinIO SigV4, Azure SAS) happens in the API, which already owns the storage integration. Every load statement is authorized and audited at the API, so a dlt load is fully governed.

License

Apache-2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dlt_duckhaven-0.3.0.tar.gz (20.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dlt_duckhaven-0.3.0-py3-none-any.whl (23.1 kB view details)

Uploaded Python 3

File details

Details for the file dlt_duckhaven-0.3.0.tar.gz.

File metadata

  • Download URL: dlt_duckhaven-0.3.0.tar.gz
  • Upload date:
  • Size: 20.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for dlt_duckhaven-0.3.0.tar.gz
Algorithm Hash digest
SHA256 78bd14884f4e2eba8531cf8ca38ba03d7734f938a4df4c32235ffc2afa416941
MD5 464fecc599490c0c17ba815562284e57
BLAKE2b-256 6a69d54e7b1b03b6487fca009f23171703b6fa43b485735caacaba1d6bf127f7

See more details on using hashes here.

Provenance

The following attestation bundles were made for dlt_duckhaven-0.3.0.tar.gz:

Publisher: release.yml on tamasmrtn/duckhaven-clients

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file dlt_duckhaven-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: dlt_duckhaven-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 23.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for dlt_duckhaven-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b7fa01ee871bbbaf8532189fa85546bfbc6f765da7512804453f15df93dbb2f6
MD5 fe2fc053e1d6422ec9e43312e132fcde
BLAKE2b-256 73332dbfe166061568da13afa8088b22ee9ce8ba5a67edfda4bd346939f1b5e3

See more details on using hashes here.

Provenance

The following attestation bundles were made for dlt_duckhaven-0.3.0-py3-none-any.whl:

Publisher: release.yml on tamasmrtn/duckhaven-clients

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.0

2 files

This release

0.3.0 This release

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page