Skip to main content

dbt-hotdata

Python versions dbt License: MIT

Transform data in Hotdata instant databases with dbt.

Hotdata is a managed analytics engine speaking HotSQL (standard SQL with analytics extensions — if you know Postgres, you already know most of it) with no DDL surface: tables are created by loading data, not by CREATE TABLE. This adapter embraces that. Every model runs as the Chain pattern, entirely against the API:

  1. the model's compiled SELECT executes server-side,
  2. the result streams back as Arrow,
  3. a native load (replace / append / upsert) applies it to the managed table.

No local database engine, no driver, no version matching — pure Python over HTTPS. The same project runs unchanged on a laptop, in CI, or in a serverless function.

Contents

Requirements

  • Python 3.11+, dbt-core 1.10+ (tested through 1.12)
  • A Hotdata workspace, an API key, and its workspace ID — from your Hotdata dashboard or the Hotdata CLI.

Install

pip install dbt-hotdata
# or
uv add dbt-hotdata

Quickstart

profiles.yml

my_project:
  target: dev
  outputs:
    dev:
      type: hotdata
      workspace_id: your_workspace_id
      # database_id: db_abc123   # pin after the first run — see below
      schema: public
      threads: 1

Set your API key in the environment (it's a secret; the workspace ID is routing, not a credential):

export HOTDATA_API_KEY=your_api_key

dbt_project.yml — Hotdata has no views, and dbt's default materialization is view, so set the project default to table:

models:
  my_project:
    +materialized: table

Then:

dbt run

On first run, an instant database labelled dbt is created automatically and its id is printed:

hotdata: created instant database db_abc123 (name='dbt'). Pin it for future runs by
setting database_id: db_abc123 in profiles.yml.

Instant databases are addressed by id — Hotdata database names are not unique, so a name can't identify one. Pin database_id in the profile to keep building into the same database; without it, each run creates a fresh one (useful for CI or per-branch runs — databases can be set to expire).

An incremental model

-- models/events_rollup.sql
{{ config(
    materialized='incremental',
    incremental_strategy='merge',
    unique_key='event_id'
) }}

select event_id, user_id, count(*) as touches, max(occurred_at) as last_seen
from {{ source('app', 'events') }}
{% if is_incremental() %}
where occurred_at > (select max(last_seen) from {{ this }})
{% endif %}
group by event_id, user_id

merge runs as a native server-side upsert matched on unique_key — updates matches, inserts the rest, no full-table read. append (the default strategy) adds the new rows. First runs and --full-refresh load with replace.

How materializations work

Materialization What happens
table Model SQL runs server-side → result loads with native replace. No temp table, no rename swap (there is no rename).
incremental Same, with append (default) or upsert (incremental_strategy: merge + unique_key, composite keys supported).
seed The CSV becomes Arrow (numbers stay exact — integers and decimals, never silently floats), then a replace load. column_types: are applied as Arrow casts.
ephemeral Standard dbt — inlined into consumers, nothing built.
view ❌ Fails up front: Hotdata has no views. The error tells you to set +materialized: table.
snapshot ❌ Fails up front: merges update rows in place, so past versions aren't kept. Keep history with an append incremental model.

Tests, dbt show, analyses, and source freshness all run as plain SELECTs on the server. dbt docs generate builds the catalog from the managed-table API plus Arrow schema probes.

Schema evolution is additive and automatic: a model that starts producing a new column just includes it in the next load — existing data is never touched, and types can widen but never silently shrink. (on_schema_change is therefore ignored.)

SQL dialect

Write models in HotSQL. It is Postgres-familiar, so SQL written for Postgres mostly runs unchanged, and the adapter overrides the cross-database macros (dateadd, datediff, convert_timezone) where HotSQL differs. Hotdata's query API also accepts the Postgres, DuckDB, and Snowflake dialects (translated to HotSQL server-side), but this adapter always submits model SQL as native HotSQL.

Feature support

Feature Support Notes
table, incremental, seed, ephemeral ✅ See above
Incremental strategies ⚠️ append, merge (native upsert by unique_key). No delete+insert, no microbatch
view, snapshot ❌ Clear error up front
Tests (generic + singular) ✅ Run server-side; store_failures supported
dbt docs generate ✅ Catalog from the managed-table API
Source freshness ✅ loaded_at_field queries run server-side
Cross-database macros ✅ dateadd, datediff, convert_timezone implemented for HotSQL (convert_timezone is DST-aware)
Hooks (pre-hook/post-hook, on-run-*) ⚠️ Run server-side — SELECT-shaped SQL only (no DDL exists)
Python models ❌
Model contracts / constraints ❌ No DDL; dbt warns they are unenforced
Grants ❌ Ignored with a warning — access is governed by workspace API keys
Transactions ❌ begin/commit are no-ops (Hotdata has no transactions)
Query cancellation ❌ An in-flight HTTPS query can't be interrupted client-side

Configuration

Profile field Env variable Default Description
api_key HOTDATA_API_KEY required API key (a secret — prefer the env var or "{{ env_var('HOTDATA_API_KEY') }}")
workspace_id HOTDATA_WORKSPACE required Workspace ID (routing, not a secret)
database_id HOTDATA_DATABASE — Id of the instant database to build into. This is how a database is targeted — names aren't unique. Printed on first-run create; pin it to reuse
database_name — dbt Display label used only when creating a new database (never to look one up)
schema — public Schema inside the instant database
create_database_if_missing — true Create a database on first run when no database_id is pinned
api_base_url HOTDATA_API_URL https://api.hotdata.dev API endpoint
max_retries — 8 Retry budget for transient errors (409/429/5xx). Loads take a catalog-level lock per database; ~42s of linear backoff outlasts a concurrent writer
retry_backoff_seconds — 1.5 Initial retry wait (grows linearly)
threads — 1 Loads into one database serialize server-side (contention is retried); more threads still help when models spend most of their time in query execution

database: stays unset — inside an instant database the SQL catalog is always literally default (relations render as "default"."schema"."table"), and the adapter rejects any other value up front.

Fields left unset in the profile resolve from the platform's own HOTDATA_* environment variables (explicit profile values always win). These are the Hotdata CLI conventions — under any orchestrator that sets them, the adapter needs no profile fields at all beyond type: hotdata. A database_id adopted from the environment is logged, since it retargets the whole build.

How it relates to hotdata-dlt-destination

hotdata-dlt-destination loads external data into Hotdata (the EL); this adapter transforms it inside Hotdata (the T). They share the same conventions — HOTDATA_API_KEY from the environment, workspace_id as a plain parameter, id-first database_id addressing, the same retry classification — and the same underlying SDK (hotdata + hotdata-framework). Point dbt at the database_id your dlt pipeline prints, add sources for the loaded tables, and build models on top.

Running after a dlt load

hotdata-dlt-destination ships a dbt bridge: after a pipeline run, one helper call executes a dbt package against the exact instant database the load just wrote — no profiles.yml to author, credentials and routing reused from the pipeline. See “Transform with dbt” in that repo's README. This adapter itself knows nothing about dlt; the bridge drives it through the HOTDATA_* environment contract above.

Development

The project uses uv for dependency management.

git clone https://github.com/hotdata-dev/dbt-hotdata.git
cd dbt-hotdata

uv sync                 # install deps (including dev group)

uv run pytest           # run the test suite (offline — no credentials needed)
uv run ruff check       # lint
uv run ruff format      # format
uv run mypy             # type-check

The test suite runs entirely offline: adapter logic is exercised against an in-memory fake client, and a real dbt parse verifies plugin registration and every macro.

License

MIT © Hotdata Inc.

Resources

Release files for dbt-hotdata 0.2.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for dbt-hotdata 0.2.2
File Size Uploaded
dbt_hotdata-0.2.2.tar.gz 144.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for dbt-hotdata 0.2.2
File Interpreter ABI Platform
dbt_hotdata-0.2.2-py3-none-any.whl Python 3 none any Details

Total release size: 177.7 kB

Release files / dbt_hotdata-0.2.2.tar.gz

Download URL dbt_hotdata-0.2.2.tar.gz
Size 144.7 kB
Tags Source
SHA-256 checksum
How to use checksums
60c4226c10d3f2ccf1dc3e10653304463c8dad7ef39553e2f6a839debdc3e3cc
BLAKE2b-256 checksum
How to use checksums
c0f941ec8c0b3d68d888e7ec67825264807e3ad63b8df786e253fbf17a53ca31
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release files / dbt_hotdata-0.2.2-py3-none-any.whl

Download URL dbt_hotdata-0.2.2-py3-none-any.whl
Size 33.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
673d8e48dcb3988ae713853d4112ff3a17a1cbd85530aae5a4c8184a952aab91
BLAKE2b-256 checksum
How to use checksums
4bc117914e5d2e89c0cba8ba9a495d8f7903b644227cc84bc3e1e9601163e38d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.2 This release

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page