Skip to main content

dbt-molinia

A dbt adapter for Molinia — a DuckDB-dialect managed data warehouse.

Molinia speaks standard DuckDB SQL, so this adapter reuses the dbt-duckdb macro library (Apache-2.0) with targeted divergences documented below.


Installation

pip install dbt-molinia

This adapter runs on dbt v1 (dbt-core ≥ 1.7.0, < 2.0). dbt v2 — the Rust rewrite that pip install dbt now gives you — reaches warehouses through ADBC drivers compiled into its own binary and loads no Python adapters at all, so it cannot use this one. Install into a virtual environment and check what you got:

dbt --version

It must report dbt-core 1.x alongside molinia: <version>. If it reports 2.x, a v2 installation is earlier on your PATH and every command will fail on type: molinia.


Profile configuration

# ~/.dbt/profiles.yml
my_molinia_project:
  target: dev
  outputs:
    dev:
      type: molinia
      host: https://app.molinia.eu              # the console/API host (dev: https://app.dev.molinia.eu)
      org_id: org_0123456789abcdef01234567      # the organisation's public id (or its legacy integer id)
      token: rd_sa_<your-service-account-token> # service account token
      schema: main                              # DuckDB schema (default: main)
      threads: 1                                # see "Rate limits" below
      # optional:
      warehouse_id: null                        # numeric id of a dedicated warehouse
      requests_per_minute: 25                   # client-side pacing, shared by all threads
      retry_max_wait_seconds: 600               # longest one statement waits on retries
Field Required Default Meaning
host yes API base URL, no /api suffix. https://app.molinia.eu in production, https://app.dev.molinia.eu for the dev environment. The apex molinia.eu is the marketing site and answers every API route with 404.
org_id yes Public id org_… (or the legacy integer id).
token yes Service-account token, rd_sa_….
schema no main Target schema.
warehouse_id no none A dedicated warehouse's numeric id, as 19 or "19"; sent to the API as a number. Unset or empty routes queries to the org engine.
requests_per_minute no none (off) Client-side ceiling on HTTP requests per minute to this host + org_id. One sliding-window limiter is shared by every dbt thread in the process, so threads: 4 cannot send four times this rate.
retry_max_wait_seconds no 600 Total seconds one statement may spend waiting to be retried before the run fails: the 429/5xx back-off or server-requested waits, plus the requests_per_minute pacing wait before each retry. Pacing before a statement's first attempt is not counted. 0 disables retries.

These are plain YAML integers. When they come from an environment variable, convert them: requests_per_minute: "{{ env_var('MOLINIA_RPM') | as_number }}".

Important: org_id is the organisation's public id (org_ followed by 24 hex characters), which the console shows under Settings → General. The legacy integer id (e.g. 123) is still accepted. Anything else — including the DuckDB catalog name org-123 or a slug — is answered with 404 Organization not found, on purpose: an unknown identifier is indistinguishable from a missing organisation.

Obtain a service account token from Settings → Service Accounts in the Molinia UI or via POST /api/orgs/{orgId}/service-accounts (org owner/admin only).

A new service account has no privileges. It can authenticate but every call returns 403 until a role is granted to it. For dbt, the role needs at least:

Privilege Why
api:query POST /query/execute — every model, seed and test
table:select read statements
table:create CREATE TABLE AS for materialised models
table:insert, table:update, table:delete incremental and snapshot strategies
table:drop dbt run --full-refresh, which drops before recreating

Grant it from Settings → Roles → Grants → Grant to service account, or:

curl -X POST "$HOST/api/orgs/$ORG/roles/$ROLE_ID/service-account-grants" \
  -H "Authorization: Bearer $ADMIN_JWT" -H 'Content-Type: application/json' \
  -d '{"serviceAccountId": 42}'

A read-only account (api:query + table:select) works for dbt docs, dbt ls and dbt test against pre-built models, but not for dbt run.

To revoke a key, call PATCH /api/orgs/{orgId}/service-accounts/{id}/deactivate (there is no /revoke endpoint).


Rate limits

dbt sends one HTTP request per statement: every model is several (introspection, create schema, create or replace, …) and every test is one more. Molinia limits that traffic in two independent places, and both answer HTTP 429:

Limit Scope Free plan Paid plan 429 response
API request limit per client IP, per API route 60 requests/min 60 requests/min {"statusCode": 429, "message": "ThrottlerException: Too Many Requests"} plus a Retry-After: <seconds> header. Once tripped, the route stays blocked until the window ends.
Org-engine query limit per organisation, org-engine queries only (not warehouse_id queries) 30 queries/min 300 queries/min {"code": "org_engine_rate_limit", "limitPerMin": 30, "retryAfterSeconds": N, "message": …}, no Retry-After header
Org-engine daily budget per organisation, UTC day 3600 engine-seconds 36000 engine-seconds {"code": "org_engine_daily_budget", "limitSeconds": …, "usedSeconds": …, "message": …}

The API request limit counts every client behind the same IP address together, per route: a dbt run and a script on the same machine that both call /query/execute share one budget of 60 a minute. The paid-plan org-engine numbers apply to a paid organisation with a valid payment mandate.

What the adapter does about each:

  • 429 from either per-minute limit — waits and retries the statement. It waits for as long as the server asks (Retry-After, in seconds or as an HTTP date, else the body's retryAfterSeconds; at least 1 s). When the server gives no hint it follows a fixed schedule: 2, 4, 8, 16, then 30 s per further retry. Every wait is logged, e.g. Molinia adapter: API rate limited, waiting 12s: org-engine query limit (30 queries/min for this org); retry 1; 12s of retry_max_wait_seconds=600 used. When the next wait would push the statement's total past retry_max_wait_seconds, the statement fails with an error naming the limit. With requests_per_minute set, a retry also waits for a pacing slot, and that wait counts toward the same total.
  • 429 org_engine_daily_budget — fails immediately with the server's message. Minutes of waiting cannot clear a budget that resets at 00:00 UTC; run the workload on a warehouse (warehouse_id) instead.
  • 503 and connections that could not be made — retried up to 3 times on the same schedule (2, 4, 8 s), within the same retry_max_wait_seconds. The server raises every 503 before it runs the statement, and a connection that was never made carried no statement, so any statement is re-sent.
  • 502, 504 and connections that dropped after the request was sent — the statement may already have run, or may still be running: the server does not cancel a statement when its client goes away. Only read-only statements (SELECT, WITH … SELECT, SHOW, DESCRIBE) are re-sent, on the same schedule. Everything else fails with an error that says so, and nothing is re-sent automatically. A re-sent INSERT, UPDATE, MERGE, COPY or CALL could apply its changes twice. A re-sent CREATE OR REPLACE, DROP or DELETE would leave the right end state, but it would run concurrently with the original on the org engine. That can fail with a write-write conflict, and it doubles the engine time and memory the statement uses. Check the target, then run dbt retry, which re-runs only the failed nodes.

Retrying keeps a run alive; pacing keeps it from hitting the limits in the first place. On a free-plan org use threads: 1 and requests_per_minute: 25: 25 stays under the 30 org-engine queries a minute with room for another tool (a parity check, the console) on the same org, and far under the 60 requests a minute the API allows one IP. On a paid plan the org-engine limit is 300 a minute, so the API request limit (60) is the binding one: requests_per_minute: 50.


Local development quickstart

# ~/.dbt/profiles.yml
my_molinia_local:
  target: dev
  outputs:
    dev:
      type: molinia
      host: http://localhost:8080   # local Molinia server
      org_id: 1                     # numeric org ID of your local org
      token: rd_sa_<local-sa-token>
      schema: main

Verify the connection

dbt debug

Expected output (truncated):

Connection:
  host: http://localhost:8080
  org_id: 1
  schema: main
  warehouse_id: None
  requests_per_minute: None
  retry_max_wait_seconds: 600

Connection test: [OK connection ok]

All checks passed!

If dbt debug reports 404 Organization not found, org_id is neither the org_… public id nor the integer id: check it against Settings → General. A 403 names the privilege the service account is missing.

Run your models

dbt run          # compile + execute all models
dbt test         # run schema and data tests
dbt docs serve   # generate and open documentation

Typed results

The query API answers JSON, and JSON cannot express what DuckDB returned. So every response carries a columnTypeIds array next to columns and rows: one DuckDB logical type id per column. The adapter uses it to hand dbt real Python values.

This matters because several types cross the wire as strings on purpose. A BIGINT or a DECIMAL is exact in DuckDB and would lose digits as a JSON number (9007199254740993 comes back as 9007199254740992), so the server sends them as text. Temporal values arrive as ISO-8601 text for the same reason. Passing those strings on to dbt is not harmless: dbt test selects count(*) as failures, which is a BIGINT, and dbt's run-results schema requires an integer — so every test in a project errored with '0' is not of type 'integer' while every one of them was green.

DuckDB type Python value
BOOLEAN bool
TINYINT, SMALLINT, INTEGER, BIGINT, HUGEINT and the unsigned variants int (exact at any width)
FLOAT, DOUBLE float
DECIMAL, BIGNUM decimal.Decimal — exact, never via float
DATE datetime.date
TIME, TIME_NS datetime.time (naive)
TIME WITH TIME ZONE datetime.time (aware)
TIMESTAMP, TIMESTAMP_S/_MS/_NS datetime.datetime (naive)
TIMESTAMP WITH TIME ZONE datetime.datetime (aware)
VARCHAR, JSON, BLOB, ENUM, UUID, LIST, STRUCT, MAP, INTERVAL, everything else unchanged, as received

Notes:

  • JSON is an alias over VARCHAR in DuckDB and reports the same type id, so JSON text stays text — which is what dbt's agate helper expects.
  • A plain TIMESTAMP is naive in DuckDB. The response renders it with a trailing Z, which is an ISO-8601 artefact and not a zone; only TIMESTAMP WITH TIME ZONE produces an aware datetime.
  • datetime resolves to microseconds, so a TIMESTAMP_NS value is truncated after 6 fractional digits.
  • Conversion never fails a statement. An unknown type id, a response without columnTypeIds (an older server), or a value Python cannot represent (infinity, a year past 9999) leaves the value exactly as it arrived.

{{ load_result(...) }} in a macro therefore sees int, Decimal and datetime values. agate has no time type, so a datetime.time is rendered as text in an agate table ('03:04:05'); every other type above survives with its Python type intact.


Common errors

Message Cause
404 Organization not found org_id is neither the org_… public id nor the integer id
403 … lacks 'table:create' the service account's role is missing a privilege (see the table above)
400 … LAKE_ENABLED_ORGS the organisation has no lakehouse enabled, so it has nowhere durable to write. Reads and plain DDL work; anything that grows data (CREATE TABLE AS, INSERT, MERGE) is refused until an operator enables it
Could not find adapter type molinia dbt v2 is running (see Installation)

Relation naming

Molinia DuckDB catalogs are named org-{orgId} (engine path) or warehouse-{warehouseId} (dedicated pool path). These names contain hyphens which require special quoting in three-part DuckDB identifiers.

To avoid ambiguity, dbt-molinia renders all relation identifiers as two-part names: "schema"."identifier". The catalog/database is determined by the connection and is intentionally omitted from generated SQL.

dbt field Maps to Example
database DuckDB catalog (omitted from SQL) org-123
schema DuckDB schema analytics
identifier Table or view name orders

Generated SQL reference: "analytics"."orders" ✓
(Not: "org-123"."analytics"."orders")


Supported SQL dialect macros

All macros dispatch via adapter.dispatch so user projects can override any of them.

dbt macro Molinia / DuckDB implementation
current_timestamp() now()
dateadd(datepart, interval, from_date) from_date + INTERVAL (interval) datepart
datediff(datepart, from, to) datediff('datepart', from, to)
concat(fields) concat(field1, field2, …)
hash(field) md5(cast(field as varchar))
last_day(date, datepart) last_day(date) (month); computed boundary for quarter/year
safe_cast(field, type) TRY_CAST(field AS type)
snapshot_string_as_time(id) CAST(id AS TIMESTAMP)
type_bigint() BIGINT
type_boolean() BOOLEAN
type_float() DOUBLE
type_int() INTEGER
type_numeric() DECIMAL(28, 6)
type_string() VARCHAR
type_timestamp() TIMESTAMP

generate_schema_name

-- Custom schema set in model config (+schema: reporting)?  Use it directly.
-- Otherwise fall back to target.schema from profiles.yml.

Unlike the dbt default, Molinia does not prefix the custom schema with the target schema. If you set +schema: reporting in dbt_project.yml, the model lands in reporting, not <target_schema>_reporting.


Divergences from dbt-duckdb

The following dbt-duckdb features are not supported in dbt-molinia. Attempting to use them raises a DbtRuntimeError with an explanation.

dbt-duckdb feature Molinia status Reason
external_table materialization NOT SUPPORTED No local filesystem access from dbt; use Molinia ingest API
read_csv() / read_parquet() path references NOT SUPPORTED Molinia manages storage via S3/ingestion API
ATTACH DATABASE NOT SUPPORTED Multi-org isolation is server-managed; use Data Sharing API
Extension install macros NOT SUPPORTED httpfs/azure managed by Molinia server
Python UDFs from dbt side NOT SUPPORTED UDFs are managed via the Molinia UDF API
database in 3-part relation Maps to org-{orgId} catalog DuckDB catalog naming — omitted from rendered SQL

Catalog introspection (information_schema)

dbt uses information_schema queries extensively for catalog commands (dbt docs generate, dbt test --store-failures, source freshness checks). The behaviour of DuckDB's information_schema through the Molinia query API is documented here.

Verified query patterns

-- List user-visible schemas (internal schemas excluded by the adapter)
SELECT schema_name
FROM information_schema.schemata
WHERE catalog_name = current_database()
  AND schema_name NOT IN ('information_schema', 'pg_catalog', 'temp');

-- List tables in a schema
SELECT table_name, table_type
FROM information_schema.tables
WHERE table_schema = 'main';        -- returns table_type = 'BASE TABLE' or 'VIEW'

-- Get columns for a table
SELECT column_name, data_type, character_maximum_length,
       numeric_precision, numeric_scale
FROM information_schema.columns
WHERE table_schema = 'main' AND table_name = 'my_model'
ORDER BY ordinal_position;

-- Check if a relation exists
SELECT count(*)
FROM information_schema.tables
WHERE table_schema = 'main' AND table_name = 'my_model';

DuckDB internal schemas

DuckDB exposes information_schema, pg_catalog, and temp in information_schema.schemata in addition to user-created schemas. The molinia__list_schemas macro filters these out so dbt's list_schemas() only returns user-visible schemas like main, analytics, or any custom schema created by your project.

RLS temp views and information_schema

Molinia's RLS enforcement creates TEMP VIEW objects (e.g. CREATE OR REPLACE TEMP VIEW "orders" AS SELECT * FROM "org-123".main."orders" WHERE …). In DuckDB, temp objects live in the temp schema of the current session's catalog. Querying information_schema.tables WHERE table_schema = 'main' always returns the underlying base table (table_type = 'BASE TABLE'), NOT the temp view — even when an RLS view shadows it at query time.

Consequently:

  • dbt docs generate sees accurate table_type values.
  • get_columns_in_relation reads columns from the base-table entry; no special handling for RLS views is needed.
  • Source freshness row-count checks query through the RLS view (correct — they should respect access policy), while catalog checks see the real table.

DuckDB → dbt type mapping

Column metadata, as information_schema.columns.data_type reports it. For the Python type a query result arrives as, see Typed results.

DuckDB data_type dbt dtype
VARCHAR, TEXT string
INTEGER, INT4 int
BIGINT, INT8 bigint
DOUBLE, FLOAT8 float
DECIMAL(p,s) numeric
BOOLEAN boolean
TIMESTAMP, TIMESTAMPTZ timestamp
DATE date
BLOB binary

Types not in the table are passed through verbatim.


Development

pip install -e ".[dev]"
python -m unittest discover        # from this directory; pytest tests/ works too

The suite needs no network: HTTP behaviour (429 back-off, 5xx retries, pacing) is tested against injected clocks and a local http.server, and result typing against the response bodies the API really sends.


License

Apache-2.0 — see LICENSE and NOTICE.

The SQL macro implementations in dbt/include/molinia/macros/ are derived from dbt-duckdb, which is licensed under the Apache License 2.0. Earlier revisions of this file called that project MIT-licensed; it never was.

Release files for dbt-molinia 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for dbt-molinia 0.1.1
File Size Uploaded
dbt_molinia-0.1.1.tar.gz 67.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for dbt-molinia 0.1.1
File Interpreter ABI Platform
dbt_molinia-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 111.5 kB

Release files / dbt_molinia-0.1.1.tar.gz

Download URL dbt_molinia-0.1.1.tar.gz
Size 67.6 kB
Tags Source
SHA-256 checksum
How to use checksums
f43b0dea8ff138d9d5ea6e163ac94080ed629c0df8c15702246b6b276e98333c
BLAKE2b-256 checksum
How to use checksums
ecb5d0fc07a35b1e81cd454fe39467527a643fff53a85b3e6f2f6338d7ee8d8b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release files / dbt_molinia-0.1.1-py3-none-any.whl

Download URL dbt_molinia-0.1.1-py3-none-any.whl
Size 43.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f467b83c7cc2cc6e7cac78c40674406a573df901dae6f33d5af2286f9b195682
BLAKE2b-256 checksum
How to use checksums
acc97e2cf52d7a04dc1135276d51278428c17513433bb09fe927c6513cbd6a32
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page