dbt-hotdata
Transform data in Hotdata instant databases with dbt.
Hotdata is a managed analytics engine speaking HotSQL (standard SQL with analytics extensions — if you know Postgres, you already know most of it) with no DDL surface: tables are created by loading data, not by CREATE TABLE. This adapter embraces that. Every model runs as the Chain pattern, entirely against the API:
- the model's compiled
SELECTexecutes server-side, - the result streams back as Arrow,
- a native load (
replace/append/upsert) applies it to the managed table.
No local database engine, no driver, no version matching — pure Python over HTTPS. The same project runs unchanged on a laptop, in CI, or in a serverless function.
Contents
- Requirements
- Install
- Quickstart
- How materializations work
- Feature support
- Configuration
- How it relates to hotdata-dlt-destination
- Development
- License
Requirements
- Python 3.11+, dbt-core 1.10+ (tested through 1.12)
- A Hotdata workspace, an API key, and its workspace ID — from your Hotdata dashboard or the Hotdata CLI.
Install
pip install dbt-hotdata
# or
uv add dbt-hotdata
Quickstart
profiles.yml
my_project:
target: dev
outputs:
dev:
type: hotdata
workspace_id: your_workspace_id
# database_id: db_abc123 # pin after the first run — see below
schema: public
threads: 1
Set your API key in the environment (it's a secret; the workspace ID is routing, not a credential):
export HOTDATA_API_KEY=your_api_key
dbt_project.yml — Hotdata has no views, and dbt's default materialization is view, so set the project default to table:
models:
my_project:
+materialized: table
Then:
dbt run
On first run, an instant database labelled dbt is created automatically and its id is printed:
hotdata: created instant database db_abc123 (name='dbt'). Pin it for future runs by
setting database_id: db_abc123 in profiles.yml.
Instant databases are addressed by id — Hotdata database names are not unique, so a name can't identify one. Pin database_id in the profile to keep building into the same database; without it, each run creates a fresh one (useful for CI or per-branch runs — databases can be set to expire).
An incremental model
-- models/events_rollup.sql
{{ config(
materialized='incremental',
incremental_strategy='merge',
unique_key='event_id'
) }}
select event_id, user_id, count(*) as touches, max(occurred_at) as last_seen
from {{ source('app', 'events') }}
{% if is_incremental() %}
where occurred_at > (select max(last_seen) from {{ this }})
{% endif %}
group by event_id, user_id
merge runs as a native server-side upsert matched on unique_key — updates matches, inserts the rest, no full-table read. append (the default strategy) adds the new rows. First runs and --full-refresh load with replace.
How materializations work
| Materialization | What happens |
|---|---|
table |
Model SQL runs server-side → result loads with native replace. No temp table, no rename swap (there is no rename). |
incremental |
Same, with append (default) or upsert (incremental_strategy: merge + unique_key, composite keys supported). |
seed |
The CSV becomes Arrow (numbers stay exact — integers and decimals, never silently floats), then a replace load. column_types: are applied as Arrow casts. |
ephemeral |
Standard dbt — inlined into consumers, nothing built. |
view |
❌ Fails up front: Hotdata has no views. The error tells you to set +materialized: table. |
snapshot |
❌ Fails up front: merges update rows in place, so past versions aren't kept. Keep history with an append incremental model. |
Tests, dbt show, analyses, and source freshness all run as plain SELECTs on the server. dbt docs generate builds the catalog from the managed-table API plus Arrow schema probes.
Schema evolution is additive and automatic: a model that starts producing a new column just includes it in the next load — existing data is never touched, and types can widen but never silently shrink. (on_schema_change is therefore ignored.)
SQL dialect
Write models in HotSQL. It is Postgres-familiar, so SQL written for Postgres mostly runs unchanged, and the adapter overrides the cross-database macros (dateadd, datediff, convert_timezone) where HotSQL differs. Hotdata's query API also accepts the Postgres, DuckDB, and Snowflake dialects (translated to HotSQL server-side), but this adapter always submits model SQL as native HotSQL.
Feature support
| Feature | Support | Notes |
|---|---|---|
table, incremental, seed, ephemeral |
✅ | See above |
| Incremental strategies | ⚠️ | append, merge (native upsert by unique_key). No delete+insert, no microbatch |
view, snapshot |
❌ | Clear error up front |
| Tests (generic + singular) | ✅ | Run server-side; store_failures supported |
dbt docs generate |
✅ | Catalog from the managed-table API |
| Source freshness | ✅ | loaded_at_field queries run server-side |
| Cross-database macros | ✅ | dateadd, datediff, convert_timezone implemented for HotSQL (convert_timezone is DST-aware) |
Hooks (pre-hook/post-hook, on-run-*) |
⚠️ | Run server-side — SELECT-shaped SQL only (no DDL exists) |
| Python models | ❌ | |
| Model contracts / constraints | ❌ | No DDL; dbt warns they are unenforced |
| Grants | ❌ | Ignored with a warning — access is governed by workspace API keys |
| Transactions | ❌ | begin/commit are no-ops (Hotdata has no transactions) |
| Query cancellation | ❌ | An in-flight HTTPS query can't be interrupted client-side |
Configuration
| Profile field | Env variable | Default | Description |
|---|---|---|---|
api_key |
HOTDATA_API_KEY |
required | API key (a secret — prefer the env var or "{{ env_var('HOTDATA_API_KEY') }}") |
workspace_id |
HOTDATA_WORKSPACE |
required | Workspace ID (routing, not a secret) |
database_id |
HOTDATA_DATABASE |
— | Id of the instant database to build into. This is how a database is targeted — names aren't unique. Printed on first-run create; pin it to reuse |
database_name |
— | dbt |
Display label used only when creating a new database (never to look one up) |
schema |
— | public |
Schema inside the instant database |
create_database_if_missing |
— | true |
Create a database on first run when no database_id is pinned |
api_base_url |
HOTDATA_API_URL |
https://api.hotdata.dev |
API endpoint |
max_retries |
— | 8 |
Retry budget for transient errors (409/429/5xx). Loads take a catalog-level lock per database; ~42s of linear backoff outlasts a concurrent writer |
retry_backoff_seconds |
— | 1.5 |
Initial retry wait (grows linearly) |
threads |
— | 1 |
Loads into one database serialize server-side (contention is retried); more threads still help when models spend most of their time in query execution |
database: stays unset — inside an instant database the SQL catalog is always literally default (relations render as "default"."schema"."table"), and the adapter rejects any other value up front.
Fields left unset in the profile resolve from the platform's own HOTDATA_* environment variables (explicit profile values always win). These are the Hotdata CLI conventions — under any orchestrator that sets them, the adapter needs no profile fields at all beyond type: hotdata. A database_id adopted from the environment is logged, since it retargets the whole build.
How it relates to hotdata-dlt-destination
hotdata-dlt-destination loads external data into Hotdata (the EL); this adapter transforms it inside Hotdata (the T). They share the same conventions — HOTDATA_API_KEY from the environment, workspace_id as a plain parameter, id-first database_id addressing, the same retry classification — and the same underlying SDK (hotdata + hotdata-framework). Point dbt at the database_id your dlt pipeline prints, add sources for the loaded tables, and build models on top.
Running after a dlt load
hotdata-dlt-destination ships a dbt bridge: after a pipeline run, one helper call executes a dbt package against the exact instant database the load just wrote — no profiles.yml to author, credentials and routing reused from the pipeline. See “Transform with dbt” in that repo's README. This adapter itself knows nothing about dlt; the bridge drives it through the HOTDATA_* environment contract above.
Development
The project uses uv for dependency management.
git clone https://github.com/hotdata-dev/dbt-hotdata.git
cd dbt-hotdata
uv sync # install deps (including dev group)
uv run pytest # run the test suite (offline — no credentials needed)
uv run ruff check # lint
uv run ruff format # format
uv run mypy # type-check
The test suite runs entirely offline: adapter logic is exercised against an in-memory fake client, and a real dbt parse verifies plugin registration and every macro.
License
MIT © Hotdata Inc.
Resources
Release files for dbt-hotdata 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dbt_hotdata-0.2.1.tar.gz | 144.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dbt_hotdata-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 177.4 kB
Release files / dbt_hotdata-0.2.1.tar.gz
| Download URL | dbt_hotdata-0.2.1.tar.gz |
|---|---|
| Size | 144.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
32b8e982d2f0c1010460ee58ece6c0e456e5f2981ca45aff7a21ed2475e5acef
|
|
BLAKE2b-256 checksum How to use checksums |
c85d371541a158dc42c09b2865ab7b1c725aff05efc49828d7af2570ee1b61d4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.
Transparency logRelease files / dbt_hotdata-0.2.1-py3-none-any.whl
| Download URL | dbt_hotdata-0.2.1-py3-none-any.whl |
|---|---|
| Size | 33.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3af43d6415dadc1e5be9ea8dd97613c3b04828afcb782878fef8ef44c06611d3
|
|
BLAKE2b-256 checksum How to use checksums |
de90e3f184d64e19a0fcfe704710dc74528fe5355847384acedef33ac8866057
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.
Transparency log