dbt-molinia
A dbt adapter for Molinia — a DuckDB-dialect managed data warehouse.
Molinia speaks standard DuckDB SQL, so this adapter reuses the dbt-duckdb macro library (Apache-2.0) with targeted divergences documented below.
Installation
pip install dbt-molinia
This adapter runs on dbt v1 (dbt-core ≥ 1.7.0, < 2.0). dbt v2 — the Rust
rewrite that pip install dbt now gives you — reaches warehouses through ADBC
drivers compiled into its own binary and loads no Python adapters at all, so it
cannot use this one. Install into a virtual environment and check what you got:
dbt --version
It must report dbt-core 1.x alongside molinia: <version>. If it reports 2.x,
a v2 installation is earlier on your PATH and every command will fail on
type: molinia.
Profile configuration
# ~/.dbt/profiles.yml
my_molinia_project:
target: dev
outputs:
dev:
type: molinia
host: https://app.molinia.eu # the console/API host (dev: https://app.dev.molinia.eu)
org_id: org_0123456789abcdef01234567 # the organisation's public id (or its legacy integer id)
token: rd_sa_<your-service-account-token> # service account token
schema: main # DuckDB schema (default: main)
threads: 1 # see "Rate limits" below
# optional:
warehouse_id: null # numeric id of a dedicated warehouse
requests_per_minute: 25 # client-side pacing, shared by all threads
retry_max_wait_seconds: 600 # longest one statement waits on retries
| Field | Required | Default | Meaning |
|---|---|---|---|
host |
yes | API base URL, no /api suffix. https://app.molinia.eu in production, https://app.dev.molinia.eu for the dev environment. The apex molinia.eu is the marketing site and answers every API route with 404. |
|
org_id |
yes | Public id org_… (or the legacy integer id). |
|
token |
yes | Service-account token, rd_sa_…. |
|
schema |
no | main |
Target schema. |
warehouse_id |
no | none | A dedicated warehouse's numeric id, as 19 or "19"; sent to the API as a number. Unset or empty routes queries to the org engine. |
requests_per_minute |
no | none (off) | Client-side ceiling on HTTP requests per minute to this host + org_id. One sliding-window limiter is shared by every dbt thread in the process, so threads: 4 cannot send four times this rate. |
retry_max_wait_seconds |
no | 600 |
Total seconds one statement may spend waiting to be retried before the run fails: the 429/5xx back-off or server-requested waits, plus the requests_per_minute pacing wait before each retry. Pacing before a statement's first attempt is not counted. 0 disables retries. |
These are plain YAML integers. When they come from an environment variable,
convert them: requests_per_minute: "{{ env_var('MOLINIA_RPM') | as_number }}".
Important:
org_idis the organisation's public id (org_followed by 24 hex characters), which the console shows under Settings → General. The legacy integer id (e.g.123) is still accepted. Anything else — including the DuckDB catalog nameorg-123or a slug — is answered with 404 Organization not found, on purpose: an unknown identifier is indistinguishable from a missing organisation.
Obtain a service account token from Settings → Service Accounts in the
Molinia UI or via POST /api/orgs/{orgId}/service-accounts (org owner/admin only).
A new service account has no privileges. It can authenticate but every call returns 403 until a role is granted to it. For dbt, the role needs at least:
| Privilege | Why |
|---|---|
api:query |
POST /query/execute — every model, seed and test |
table:select |
read statements |
table:create |
CREATE TABLE AS for materialised models |
table:insert, table:update, table:delete |
incremental and snapshot strategies |
table:drop |
dbt run --full-refresh, which drops before recreating |
Grant it from Settings → Roles → Grants → Grant to service account, or:
curl -X POST "$HOST/api/orgs/$ORG/roles/$ROLE_ID/service-account-grants" \
-H "Authorization: Bearer $ADMIN_JWT" -H 'Content-Type: application/json' \
-d '{"serviceAccountId": 42}'
A read-only account (api:query + table:select) works for dbt docs,
dbt ls and dbt test against pre-built models, but not for dbt run.
To revoke a key, call PATCH /api/orgs/{orgId}/service-accounts/{id}/deactivate
(there is no /revoke endpoint).
Rate limits
dbt sends one HTTP request per statement: every model is several
(introspection, create schema, create or replace, …) and every test is one
more. Molinia limits that traffic in two independent places, and both answer
HTTP 429:
| Limit | Scope | Free plan | Paid plan | 429 response |
|---|---|---|---|---|
| API request limit | per client IP, per API route | 60 requests/min | 60 requests/min | {"statusCode": 429, "message": "ThrottlerException: Too Many Requests"} plus a Retry-After: <seconds> header. Once tripped, the route stays blocked until the window ends. |
| Org-engine query limit | per organisation, org-engine queries only (not warehouse_id queries) |
30 queries/min | 300 queries/min | {"code": "org_engine_rate_limit", "limitPerMin": 30, "retryAfterSeconds": N, "message": …}, no Retry-After header |
| Org-engine daily budget | per organisation, UTC day | 3600 engine-seconds | 36000 engine-seconds | {"code": "org_engine_daily_budget", "limitSeconds": …, "usedSeconds": …, "message": …} |
The API request limit counts every client behind the same IP address together,
per route: a dbt run and a script on the same machine that both call
/query/execute share one budget of 60 a minute. The paid-plan org-engine
numbers apply to a paid organisation with a valid payment mandate.
What the adapter does about each:
- 429 from either per-minute limit — waits and retries the statement. It
waits for as long as the server asks (
Retry-After, in seconds or as an HTTP date, else the body'sretryAfterSeconds; at least 1 s). When the server gives no hint it follows a fixed schedule: 2, 4, 8, 16, then 30 s per further retry. Every wait is logged, e.g.Molinia adapter: API rate limited, waiting 12s: org-engine query limit (30 queries/min for this org); retry 1; 12s of retry_max_wait_seconds=600 used. When the next wait would push the statement's total pastretry_max_wait_seconds, the statement fails with an error naming the limit. Withrequests_per_minuteset, a retry also waits for a pacing slot, and that wait counts toward the same total. - 429
org_engine_daily_budget— fails immediately with the server's message. Minutes of waiting cannot clear a budget that resets at 00:00 UTC; run the workload on a warehouse (warehouse_id) instead. - 503 and connections that could not be made — retried up to 3 times on the
same schedule (2, 4, 8 s), within the same
retry_max_wait_seconds. The server raises every 503 before it runs the statement, and a connection that was never made carried no statement, so any statement is re-sent. - 502, 504 and connections that dropped after the request was sent — the
statement may already have run, or may still be running: the server does
not cancel a statement when its client goes away. Only read-only statements
(
SELECT,WITH … SELECT,SHOW,DESCRIBE) are re-sent, on the same schedule. Everything else fails with an error that says so, and nothing is re-sent automatically. A re-sentINSERT,UPDATE,MERGE,COPYorCALLcould apply its changes twice. A re-sentCREATE OR REPLACE,DROPorDELETEwould leave the right end state, but it would run concurrently with the original on the org engine. That can fail with a write-write conflict, and it doubles the engine time and memory the statement uses. Check the target, then rundbt retry, which re-runs only the failed nodes.
Retrying keeps a run alive; pacing keeps it from hitting the limits in the
first place. On a free-plan org use threads: 1 and
requests_per_minute: 25: 25 stays under the 30 org-engine queries a minute
with room for another tool (a parity check, the console) on the same org, and
far under the 60 requests a minute the API allows one IP. On a paid plan the
org-engine limit is 300 a minute, so the API request limit (60) is the binding
one: requests_per_minute: 50.
Local development quickstart
# ~/.dbt/profiles.yml
my_molinia_local:
target: dev
outputs:
dev:
type: molinia
host: http://localhost:8080 # local Molinia server
org_id: 1 # numeric org ID of your local org
token: rd_sa_<local-sa-token>
schema: main
Verify the connection
dbt debug
Expected output (truncated):
Connection:
host: http://localhost:8080
org_id: 1
schema: main
warehouse_id: None
requests_per_minute: None
retry_max_wait_seconds: 600
Connection test: [OK connection ok]
All checks passed!
If dbt debug reports 404 Organization not found, org_id is neither the
org_… public id nor the integer id: check it against Settings → General.
A 403 names the privilege the service account is missing.
Run your models
dbt run # compile + execute all models
dbt test # run schema and data tests
dbt docs serve # generate and open documentation
Typed results
The query API answers JSON, and JSON cannot express what DuckDB returned. So
every response carries a columnTypeIds array next to columns and rows:
one DuckDB logical type id per column. The adapter uses it to hand dbt real
Python values.
This matters because several types cross the wire as strings on purpose.
A BIGINT or a DECIMAL is exact in DuckDB and would lose digits as a JSON
number (9007199254740993 comes back as 9007199254740992), so the server
sends them as text. Temporal values arrive as ISO-8601 text for the same
reason. Passing those strings on to dbt is not harmless: dbt test selects
count(*) as failures, which is a BIGINT, and dbt's run-results schema
requires an integer — so every test in a project errored with
'0' is not of type 'integer' while every one of them was green.
| DuckDB type | Python value |
|---|---|
BOOLEAN |
bool |
TINYINT, SMALLINT, INTEGER, BIGINT, HUGEINT and the unsigned variants |
int (exact at any width) |
FLOAT, DOUBLE |
float |
DECIMAL, BIGNUM |
decimal.Decimal — exact, never via float |
DATE |
datetime.date |
TIME, TIME_NS |
datetime.time (naive) |
TIME WITH TIME ZONE |
datetime.time (aware) |
TIMESTAMP, TIMESTAMP_S/_MS/_NS |
datetime.datetime (naive) |
TIMESTAMP WITH TIME ZONE |
datetime.datetime (aware) |
VARCHAR, JSON, BLOB, ENUM, UUID, LIST, STRUCT, MAP, INTERVAL, everything else |
unchanged, as received |
Notes:
JSONis an alias overVARCHARin DuckDB and reports the same type id, so JSON text stays text — which is what dbt's agate helper expects.- A plain
TIMESTAMPis naive in DuckDB. The response renders it with a trailingZ, which is an ISO-8601 artefact and not a zone; onlyTIMESTAMP WITH TIME ZONEproduces an awaredatetime. datetimeresolves to microseconds, so aTIMESTAMP_NSvalue is truncated after 6 fractional digits.- Conversion never fails a statement. An unknown type id, a response without
columnTypeIds(an older server), or a value Python cannot represent (infinity, a year past 9999) leaves the value exactly as it arrived.
{{ load_result(...) }} in a macro therefore sees int, Decimal and
datetime values. agate has no time type, so a datetime.time is rendered as
text in an agate table ('03:04:05'); every other type above survives with
its Python type intact.
Common errors
| Message | Cause |
|---|---|
404 Organization not found |
org_id is neither the org_… public id nor the integer id |
403 … lacks 'table:create' |
the service account's role is missing a privilege (see the table above) |
400 … LAKE_ENABLED_ORGS |
the organisation has no lakehouse enabled, so it has nowhere durable to write. Reads and plain DDL work; anything that grows data (CREATE TABLE AS, INSERT, MERGE) is refused until an operator enables it |
Could not find adapter type molinia |
dbt v2 is running (see Installation) |
Relation naming
Molinia DuckDB catalogs are named org-{orgId} (engine path) or
warehouse-{warehouseId} (dedicated pool path). These names contain
hyphens which require special quoting in three-part DuckDB identifiers.
To avoid ambiguity, dbt-molinia renders all relation identifiers as
two-part names: "schema"."identifier". The catalog/database is
determined by the connection and is intentionally omitted from generated SQL.
| dbt field | Maps to | Example |
|---|---|---|
database |
DuckDB catalog (omitted from SQL) | org-123 |
schema |
DuckDB schema | analytics |
identifier |
Table or view name | orders |
Generated SQL reference: "analytics"."orders" ✓
(Not: "org-123"."analytics"."orders")
Supported SQL dialect macros
All macros dispatch via adapter.dispatch so user projects can override any
of them.
| dbt macro | Molinia / DuckDB implementation |
|---|---|
current_timestamp() |
now() |
dateadd(datepart, interval, from_date) |
from_date + INTERVAL (interval) datepart |
datediff(datepart, from, to) |
datediff('datepart', from, to) |
concat(fields) |
concat(field1, field2, …) |
hash(field) |
md5(cast(field as varchar)) |
last_day(date, datepart) |
last_day(date) (month); computed boundary for quarter/year |
safe_cast(field, type) |
TRY_CAST(field AS type) |
snapshot_string_as_time(id) |
CAST(id AS TIMESTAMP) |
type_bigint() |
BIGINT |
type_boolean() |
BOOLEAN |
type_float() |
DOUBLE |
type_int() |
INTEGER |
type_numeric() |
DECIMAL(28, 6) |
type_string() |
VARCHAR |
type_timestamp() |
TIMESTAMP |
generate_schema_name
-- Custom schema set in model config (+schema: reporting)? Use it directly.
-- Otherwise fall back to target.schema from profiles.yml.
Unlike the dbt default, Molinia does not prefix the custom schema with the
target schema. If you set +schema: reporting in dbt_project.yml, the model
lands in reporting, not <target_schema>_reporting.
Divergences from dbt-duckdb
The following dbt-duckdb features are not supported in dbt-molinia.
Attempting to use them raises a DbtRuntimeError with an explanation.
| dbt-duckdb feature | Molinia status | Reason |
|---|---|---|
external_table materialization |
NOT SUPPORTED | No local filesystem access from dbt; use Molinia ingest API |
read_csv() / read_parquet() path references |
NOT SUPPORTED | Molinia manages storage via S3/ingestion API |
ATTACH DATABASE |
NOT SUPPORTED | Multi-org isolation is server-managed; use Data Sharing API |
| Extension install macros | NOT SUPPORTED | httpfs/azure managed by Molinia server |
| Python UDFs from dbt side | NOT SUPPORTED | UDFs are managed via the Molinia UDF API |
database in 3-part relation |
Maps to org-{orgId} catalog |
DuckDB catalog naming — omitted from rendered SQL |
Catalog introspection (information_schema)
dbt uses information_schema queries extensively for catalog commands (dbt docs generate,
dbt test --store-failures, source freshness checks). The behaviour of DuckDB's
information_schema through the Molinia query API is documented here.
Verified query patterns
-- List user-visible schemas (internal schemas excluded by the adapter)
SELECT schema_name
FROM information_schema.schemata
WHERE catalog_name = current_database()
AND schema_name NOT IN ('information_schema', 'pg_catalog', 'temp');
-- List tables in a schema
SELECT table_name, table_type
FROM information_schema.tables
WHERE table_schema = 'main'; -- returns table_type = 'BASE TABLE' or 'VIEW'
-- Get columns for a table
SELECT column_name, data_type, character_maximum_length,
numeric_precision, numeric_scale
FROM information_schema.columns
WHERE table_schema = 'main' AND table_name = 'my_model'
ORDER BY ordinal_position;
-- Check if a relation exists
SELECT count(*)
FROM information_schema.tables
WHERE table_schema = 'main' AND table_name = 'my_model';
DuckDB internal schemas
DuckDB exposes information_schema, pg_catalog, and temp in
information_schema.schemata in addition to user-created schemas. The
molinia__list_schemas macro filters these out so dbt's list_schemas()
only returns user-visible schemas like main, analytics, or any custom
schema created by your project.
RLS temp views and information_schema
Molinia's RLS enforcement creates TEMP VIEW objects (e.g.
CREATE OR REPLACE TEMP VIEW "orders" AS SELECT * FROM "org-123".main."orders" WHERE …).
In DuckDB, temp objects live in the temp schema of the current session's
catalog. Querying information_schema.tables WHERE table_schema = 'main'
always returns the underlying base table (table_type = 'BASE TABLE'),
NOT the temp view — even when an RLS view shadows it at query time.
Consequently:
dbt docs generatesees accuratetable_typevalues.get_columns_in_relationreads columns from the base-table entry; no special handling for RLS views is needed.- Source freshness row-count checks query through the RLS view (correct — they should respect access policy), while catalog checks see the real table.
DuckDB → dbt type mapping
Column metadata, as information_schema.columns.data_type reports it. For
the Python type a query result arrives as, see Typed results.
DuckDB data_type |
dbt dtype |
|---|---|
VARCHAR, TEXT |
string |
INTEGER, INT4 |
int |
BIGINT, INT8 |
bigint |
DOUBLE, FLOAT8 |
float |
DECIMAL(p,s) |
numeric |
BOOLEAN |
boolean |
TIMESTAMP, TIMESTAMPTZ |
timestamp |
DATE |
date |
BLOB |
binary |
Types not in the table are passed through verbatim.
Development
pip install -e ".[dev]"
python -m unittest discover # from this directory; pytest tests/ works too
The suite needs no network: HTTP behaviour (429 back-off, 5xx retries, pacing)
is tested against injected clocks and a local http.server, and result typing
against the response bodies the API really sends.
License
Apache-2.0 — see LICENSE and NOTICE.
The SQL macro implementations in dbt/include/molinia/macros/ are derived from
dbt-duckdb, which is licensed under the
Apache License 2.0. Earlier revisions of this file called that project
MIT-licensed; it never was.
Release files for dbt-molinia 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dbt_molinia-0.1.1.tar.gz | 67.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dbt_molinia-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 111.5 kB
Release files / dbt_molinia-0.1.1.tar.gz
| Download URL | dbt_molinia-0.1.1.tar.gz |
|---|---|
| Size | 67.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f43b0dea8ff138d9d5ea6e163ac94080ed629c0df8c15702246b6b276e98333c
|
|
BLAKE2b-256 checksum How to use checksums |
ecb5d0fc07a35b1e81cd454fe39467527a643fff53a85b3e6f2f6338d7ee8d8b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency logRelease files / dbt_molinia-0.1.1-py3-none-any.whl
| Download URL | dbt_molinia-0.1.1-py3-none-any.whl |
|---|---|
| Size | 43.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f467b83c7cc2cc6e7cac78c40674406a573df901dae6f33d5af2286f9b195682
|
|
BLAKE2b-256 checksum How to use checksums |
acc97e2cf52d7a04dc1135276d51278428c17513433bb09fe927c6513cbd6a32
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency log