Skip to main content

hotdata-dlt-destination

Load data into Hotdata managed databases using dlt.

dlt handles extraction, schema inference, and batching. This package handles the Hotdata side — uploading each batch as Parquet and registering it with your managed database.

Install

pip install hotdata-dlt-destination

Quickstart

import dlt
from hotdata_dlt_destination import hotdata

@dlt.resource(name="orders", write_disposition="append")
def orders_resource():
    yield [
        {"id": 1, "customer": "Alice", "total": 99.00},
        {"id": 2, "customer": "Bob",   "total": 49.50},
    ]

pipeline = dlt.pipeline(
    pipeline_name="my_pipeline",
    destination=hotdata(
        database_name="sales",
        declared_tables=["orders"],
    ),
)

pipeline.run(orders_resource())

Set your credentials as environment variables before running:

export HOTDATA_API_KEY=your_api_key
export HOTDATA_WORKSPACE=your_workspace_id

That's it. On first run, the sales managed database is created automatically and the orders table is loaded.

hotdata is a native dlt destination (JobClientBase + WithStateSync): it supports nested/child tables, preserves dlt's internal columns (_dlt_id, _dlt_load_id), and persists schema-version, load, and pipeline-state tables in the managed database so incremental sources resume correctly across runs. If an existing managed database is missing a declared table on a later run, the table is added to it in place; existing tables and their data are left untouched.

Read your data back

The same pipeline object that writes can read, through dlt's standard dataset interface — no Hotdata-specific code, no database IDs, no hand-written SQL:

ds = pipeline.dataset()

ds.table("orders").df()      # whole table -> pandas.DataFrame
ds.table("orders").arrow()   # -> pyarrow.Table

# raw SQL
ds("SELECT customer, sum(total) AS spend FROM orders GROUP BY customer").df()

# fluent
ds.table("orders").select("id", "total").where("total > 50").order_by("total").limit(10).df()

Queries run server-side on Hotdata's Apache DataFusion engine (Postgres-compatible SQL). It's the same read API you'd use with the duckdb, postgres, or bigquery destinations — enabled because hotdata advertises dlt's SQL-client interface (WithSqlClient).

You can also author queries with ibisdataset().table("orders").to_ibis() gives an ibis table; build an expression and dlt compiles it to SQL and runs it through the same client. (The live ibis backend, dataset().ibis(), is not supported — dlt wires that to a direct engine connection per destination, which Hotdata's remote engine doesn't expose.)

Feature support

Where hotdata stands against the dlt destination capability spec. ✅ supported · ⚠️ supported with caveats · ❌ not supported.

Write dispositions

Disposition Support Notes
append Existing rows kept; new batch appended (read-modify-write)
replace truncate-and-insert — table contents fully replaced
merge Upsert by primary_key — see merge strategies below

Merge strategies

Strategy Support Notes
upsert Default. Dedupes by primary_key, falling back to dlt's _dlt_id
insert-only Inserts rows whose key isn't already present; never updates existing rows
delete-insert Not supported
scd2 Not supported

Replace strategies

Strategy Support Notes
truncate-and-insert
insert-from-staging No staging dataset
staging-optimized No staging dataset

Keys & column hints

Feature Support Notes
primary_key Drives merge/upsert and insert-only de-duplication
merge_key Use primary_key
hard_delete Deletes are not propagated
dedup_sort

Loader file formats

Format Support Notes
parquet Preferred and only loader format
jsonl
insert_values
csv

Structure & lifecycle

Feature Support Notes
Nested / child tables Up to max_table_nesting (default 1000), e.g. orders__items
dlt internal columns (_dlt_id, _dlt_load_id) Preserved, never stripped
dlt system tables (_dlt_loads, _dlt_version) Persisted in the managed database
Pipeline state sync (WithStateSync) Incremental sources resume across runs
Dataset read API (pipeline.dataset()) Read loaded data as pandas / arrow / fluent SQL, server-side on DataFusion — see Read your data back
ibis expressions (.table("t").to_ibis()) Built as ibis, compiled to SQL, executed via the sql_client
Live ibis backend (dataset().ibis()) dlt maps this to a direct per-destination engine connection; Hotdata's engine is remote (REST + DuckLake), not a wire-protocol DB or local files
New columns Permissive column promotion on append/merge
New tables A table missing on a later run is declared in place on the existing database — no recreate, no data movement
Multiple tables per pipeline Pass every table name via declared_tables

Staging, transactions & identifiers

Feature Support Notes
Filesystem / remote staging Parquet is uploaded directly to Hotdata
Staging dataset
DDL transactions
Case-sensitive identifiers snake_case, case-insensitive; identifiers up to 255 chars

Configuration

Parameter Env variable Default Description
api_key HOTDATA_API_KEY required Your Hotdata API key
workspace_id HOTDATA_WORKSPACE required Your Hotdata workspace ID
database_name HOTDATA_DATABASE dlt Managed database to load into
schema HOTDATA_SCHEMA public Schema within the managed database
write_disposition HOTDATA_WRITE_DISPOSITION append Default write mode (see below)
declared_tables HOTDATA_DECLARED_TABLES All table names the pipeline will write (required for multi-table pipelines — see below)
create_database_if_missing True Create the managed database if it doesn't exist yet
max_retries HOTDATA_MAX_RETRIES 5 How many times to retry a failed request
retry_backoff_seconds HOTDATA_RETRY_BACKOFF_SECONDS 1.0 Initial wait between retries (grows with each attempt)

You can pass any of these as keyword arguments to hotdata(...), or set the corresponding environment variable. hotdata also accepts max_table_nesting (default 1000).

Write modes

Each resource can control how its data lands in the table:

Mode What it does
replace Deletes everything in the table and loads the new batch. Good for full refreshes.
append Adds new rows to the table without touching existing data. Good for event logs and immutable records.
merge (= upsert) Updates existing rows by primary key, inserts new ones. Good for syncing a source of truth.

dlt resources set write_disposition to append, replace, or merge only. merge performs upsert-by-primary-key — it is what the internal upsert disposition resolves to, so there is no separate upsert to set. The destination also implements an insert-only combine (insert rows whose key isn't already present, never updating existing rows), but dlt does not expose it as a resource write_disposition, so it cannot be selected per resource.

Set the default for all resources on the destination:

hotdata(write_disposition="replace", ...)

Or set it per resource — this takes priority:

@dlt.resource(name="customers", write_disposition="merge", primary_key="id")
def customers_resource():
    ...

Multiple tables

When a pipeline writes to more than one table, pass all table names to declared_tables. Hotdata needs to know the full list upfront to set up the managed database correctly.

pipeline = dlt.pipeline(
    pipeline_name="ecommerce",
    destination=hotdata(
        database_name="ecommerce",
        declared_tables=["customers", "orders", "products"],
    ),
)

pipeline.run([customers_resource(), orders_resource(), products_resource()])

If you add a new table later, include it in declared_tables on the next run.

Verify a load

After a pipeline runs, use the Hotdata CLI to check that the data landed:

# List your managed databases
hotdata databases list

# Check that tables are loaded and queryable
hotdata databases tables list --database sales

# Query the data
hotdata query "SELECT * FROM public.orders LIMIT 5" -d sales

Demo pipeline

The package includes a demo that downloads 9 macro-economic indicators from the Federal Reserve (FRED) and loads them into Hotdata. It's a good reference for how a real pipeline is structured.

export HOTDATA_API_KEY=your_api_key
export HOTDATA_WORKSPACE=your_workspace_id
uv run hotdata-dlt-demo

This creates a example_macro database with two tables:

  • macro_indicators_raw — one row per (date, series, value), all 9 series at their original frequency
  • macro_wide — one row per month from 1992 onward, each indicator as its own column

How it works

Each pipeline run:

  1. dlt serializes your data to Parquet
  2. The Parquet file is uploaded to Hotdata
  3. load_managed_table replaces the target table with the new data

For append, merge, upsert, and insert-only, the destination reads the current table contents first, combines in Python (by primary_key, falling back to dlt's _dlt_id), then writes the combined result back. This is done transparently — your resource just yields rows.

The destination preserves dlt's native _dlt_id / _dlt_load_id columns and persists dlt's schema-version, load, and pipeline-state tables in the managed database so incremental sources can restore their state on the next run. No extra columns are added.

Resources

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hotdata_dlt_destination-0.7.1.tar.gz (19.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

hotdata_dlt_destination-0.7.1-py3-none-any.whl (24.9 kB view details)

Uploaded Python 3

File details

Details for the file hotdata_dlt_destination-0.7.1.tar.gz.

File metadata

  • Download URL: hotdata_dlt_destination-0.7.1.tar.gz
  • Upload date:
  • Size: 19.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for hotdata_dlt_destination-0.7.1.tar.gz
Algorithm Hash digest
SHA256 f5c0662649321e7b8006831193d2075c8ea70d8488775683d733234d40d11125
MD5 bf92574c1ba1688b5ae2d0ad12770bc4
BLAKE2b-256 2c45d529946bac31594453e2492756abba72532109ab4e83d2bfd3df7fb08c8c

See more details on using hashes here.

Provenance

The following attestation bundles were made for hotdata_dlt_destination-0.7.1.tar.gz:

Publisher: publish.yml on hotdata-dev/hotdata-dlt-destination

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hotdata_dlt_destination-0.7.1-py3-none-any.whl.

File metadata

File hashes

Hashes for hotdata_dlt_destination-0.7.1-py3-none-any.whl
Algorithm Hash digest
SHA256 9def83e37d2cf08646f45c2973ae941159481399f49765bac4a5514da2276701
MD5 fcd8d1b0259a04abac77cfe64bcdd0c9
BLAKE2b-256 c719ac332cf4568e9228144eec2f3ebebb22d6cac21168528dc0c2165bf27b5a

See more details on using hashes here.

Provenance

The following attestation bundles were made for hotdata_dlt_destination-0.7.1-py3-none-any.whl:

Publisher: publish.yml on hotdata-dev/hotdata-dlt-destination

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.14.0

2 files

0.13.2

2 files

0.13.1

2 files

0.13.0

2 files

0.12.0

2 files

0.11.0

2 files

0.10.0

2 files

0.9.5

2 files

0.9.4

2 files

0.9.3

2 files

0.9.2

2 files

0.9.1

2 files

0.9.0

2 files

0.8.0

2 files

0.7.2

2 files

This release

0.7.1 This release

2 files

0.7.0

2 files

0.6.1

2 files

0.6.0

2 files

0.4.2

2 files

0.4.1

2 files

0.3.4

2 files

0.3.3

2 files

0.3.1

2 files

0.3.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page