Skip to main content

dlt-firebolt

Community dlt destination for Firebolt.

Load data into Firebolt with dlt using direct HTTP upload (default) or S3 staging + COPY INTO for large loads.

Requires dlt-firebolt 0.3.0+ for upload mode and simplified Core/managed configuration.

Two ways to load

Mode Best for Firebolt load path
upload (default) Firebolt Core, local dev, quick starts HTTP multipart → READ_PARQUET('upload://…')
s3 Managed Firebolt production, large bulk loads Parquet on S3 → COPY INTO

On managed Firebolt today, set FIREBOLT_STAGING_MODE=s3. Upload is the code default but the managed engine does not accept multipart upload yet; you will get a clear error if you use upload mode there.

Install

pip install "dlt-firebolt>=0.3.0"

Requires Python 3.10+.

Register the destination once in your project:

import firebolt_dest  # registers destination="firebolt"

Prerequisites

Upload mode (Core / local)

  1. Firebolt Core running locally (or another environment that supports upload://), or a managed account once upload is supported there.

No S3 bucket, external location, or AWS credentials required for the Firebolt load step. dlt still writes Parquet to a local staging directory during normalize.

S3 mode (managed production)

  1. Firebolt service account with access to your database and engine.
  2. S3 bucket for dlt filesystem staging.
  3. Firebolt external location for that bucket prefix:
CREATE LOCATION "your_location_name" WITH
  SOURCE = 'CLOUD_STORAGE'
  URL = 's3://your-bucket/your-prefix/'
  CREDENTIALS = (AWS_ROLE_ARN = 'arn:aws:iam::...:role/...');

Set FIREBOLT_S3_LOCATION_NAME to the exact location name. Set S3_PREFIX to the same key prefix as the LOCATION URL path for classic single-tenant setups (COPY PATTERN is relative to the LOCATION path). Set FIREBOLT_STAGING_MODE=s3.

Multi-tenant (one LOCATION for all tenants): create the LOCATION at the bucket root (URL = 's3://your-bucket/'), set FIREBOLT_S3_LOCATION_URL=s3://your-bucket/ (or s3_location_url in TOML — see below), and set a per-tenant staging prefix so objects land under s3://bucket/<tenant>/.... With make_firebolt_pipeline(from_secrets=False), that prefix is S3_PREFIX=tenant-a. With from_secrets=True / a bare firebolt() destination, steer staging with a tenant-specific [destination.filesystem] bucket_url instead — S3_PREFIX is not read on that path. COPY PATTERN keeps the tenant segment (e.g. tenant-a/dlt/staging/...). Classic prefix-LOCATION setups that omit FIREBOLT_S3_LOCATION_URL still fall back to s3_prefix as the LOCATION path. Migration note: loads that previously relied on the old silent basename fallback (object key not under the LOCATION/s3_prefix path) now hard-fail with TerminalValueError — fix the LOCATION URL / prefix alignment rather than expecting a filename-only PATTERN. A LOCATION at the bucket root grants its credentials read access to every prefix in the bucket; PATTERN is not an authorization boundary. Bucket-root suits multiple prefixes inside a single trust boundary — for real tenant isolation use per-tenant LOCATIONs or buckets.

The machine running dlt needs AWS credentials that can write to the staging bucket (via environment variables, an AWS profile, or an attached IAM role). Firebolt reads the staged files from the external location you configured, not the runner's AWS identity. Staging Parquet is not automatically deleted after COPY INTO, so add an S3 lifecycle rule or a periodic cleanup if you don't want objects to accumulate.

Quick start

Firebolt Core (upload, no S3)

export FIREBOLT_CORE_URL=http://localhost:3473
import dlt
import firebolt_dest
from firebolt_dest.configuration import make_firebolt_pipeline

@dlt.resource(name="orders", write_disposition="append")
def orders():
    yield {"order_id": 1, "customer": "Acme"}

pipeline = make_firebolt_pipeline(
    pipeline_name="my_pipeline",
    dataset_name="my_dataset",
)

pipeline.run(orders(), loader_file_format="parquet")

Managed Firebolt (S3)

Set service-account credentials and FIREBOLT_STAGING_MODE=s3 (see Configuration). The same pipeline code applies; only env vars change.

Tables are created as {dataset}_{table} in the public schema (for example my_dataset_orders). To use a real Firebolt schema instead (my_dataset.orders), set FIREBOLT_USE_SCHEMA_PER_DATASET=true — see Configuration.

Using .dlt/secrets.toml

Environment variables are the primary path used in the e2e scripts. You can also load configuration from .dlt/secrets.toml with from_secrets=True (verified on Firebolt Core with upload mode; see .dlt/secrets.toml.example).

Note: In dlt's credential block, host is the Firebolt database name and database is the Firebolt engine name (matching FIREBOLT_DATABASE and FIREBOLT_ENGINE, not swapped).

pipeline = make_firebolt_pipeline(
    pipeline_name="my_pipeline",
    dataset_name="my_dataset",
    from_secrets=True,
)

Example for managed Firebolt (S3 mode):

[destination.firebolt]
staging_mode = "s3"
s3_location_name = "your_location_name"
s3_prefix = "dlt-landing"
# use_schema_per_dataset = false  # opt-in real schemas; default off
# Multi-tenant bucket-root LOCATION (PATTERN relative to this URL):
# s3_location_url = "s3://example-bucket/"
# Steer staging with a tenant-specific filesystem bucket_url (S3_PREFIX env is
# NOT read when from_secrets=True / destination="firebolt"):
# (see bucket_url below — e.g. s3://example-bucket/tenant-a/dlt/staging)

[destination.firebolt.credentials]
host = "YOUR_FIREBOLT_DATABASE"
database = "YOUR_FIREBOLT_ENGINE"
username = "YOUR_CLIENT_ID"
password = "YOUR_CLIENT_SECRET"
account_name = "YOUR_ACCOUNT_NAME"

[destination.filesystem]
bucket_url = "s3://example-bucket/dlt-landing/dlt/staging"
# Multi-tenant example: bucket_url = "s3://example-bucket/tenant-a/dlt/staging"

Which config path reads which variable

Setting make_firebolt_pipeline(from_secrets=False) from_secrets=True / bare firebolt()
LOCATION URL for PATTERN FIREBOLT_S3_LOCATION_URL → s3_location_url s3_location_url in [destination.firebolt] (or DESTINATION__FIREBOLT__S3_LOCATION_URL)
Staging write prefix S3_BUCKET + S3_PREFIX → filesystem bucket_url [destination.filesystem] bucket_url only (S3_PREFIX is not read)
LOCATION name FIREBOLT_S3_LOCATION_NAME s3_location_name

See .dlt/secrets.toml.example for a Firebolt Core upload template and the managed S3 fields including s3_location_url.

With from_secrets=True (or a plain destination="firebolt" pipeline), set the flag in TOML as use_schema_per_dataset = true under [destination.firebolt], or via the dlt env alias DESTINATION__FIREBOLT__USE_SCHEMA_PER_DATASET=true. FIREBOLT_USE_SCHEMA_PER_DATASET is read only when make_firebolt_pipeline(..., from_secrets=False) builds the destination from env.

Configuration

Managed Firebolt

Variable Required Description
FIREBOLT_CLIENT_ID yes Service account client ID
FIREBOLT_CLIENT_SECRET yes Service account secret
FIREBOLT_ACCOUNT_NAME yes Firebolt account name
FIREBOLT_DATABASE yes Target database
FIREBOLT_ENGINE yes Engine name
FIREBOLT_STAGING_MODE no s3 for managed production (recommended)
FIREBOLT_S3_LOCATION_NAME s3 mode Firebolt external location name
FIREBOLT_S3_LOCATION_URL no LOCATION URL for PATTERN (e.g. s3://bucket/); empty → use s3_prefix as LOCATION path. Read by make_firebolt_pipeline(from_secrets=False); for secrets/TOML use s3_location_url / DESTINATION__FIREBOLT__S3_LOCATION_URL
S3_BUCKET s3 mode Staging bucket (from_secrets=False only; secrets path uses filesystem bucket_url)
S3_PREFIX no Staging key prefix (default: dlt-landing); per-tenant prefix when LOCATION is at bucket root (from_secrets=False only)
FIREBOLT_USE_SCHEMA_PER_DATASET no true to map dataset_name to a real Firebolt schema (default: off — tables stay public.{dataset}_{table}). Read only by make_firebolt_pipeline(from_secrets=False); with from_secrets=True / bare firebolt() set use_schema_per_dataset in TOML or DESTINATION__FIREBOLT__USE_SCHEMA_PER_DATASET

The destination resolves the engine URL from your account; you do not set an HTTP endpoint manually.

Schema-per-dataset (opt-in)

Default is off. Existing pipelines keep writing public.{dataset}_{table} and store dlt state tables (_dlt_loads, _dlt_pipeline_state, _dlt_version) under that public-prefix naming. Enabling the flag is a deliberate breaking change for that layout:

Flag off (default) Flag on
Table name public.tenant_a_orders tenant_a.orders
Staging public.tenant_a_staging_* schema tenant_a_staging
dlt state public.{dataset}__dlt_* {dataset}._dlt_*

Where dlt state actually lives: the incremental cursor / load state that “fresh vs retained” refers to is primarily local pipeline state on the machine running dlt (under the pipeline working directory). Destination {dataset}._dlt_* tables are a separate store; dropping them alone does not reset the local cursor.

First flag-ON run — two different situations:

  • Fresh machine (no local pipeline state): the destination schema being pre-created or not does not change local state — you still get a fresh local cursor. Destination _dlt_* tables appear only after the first successful load.
  • Same machine with existing local state: that local state is what continues the cursor. Pre-creating the destination schema (for CLONE/CTAS) does not by itself “retain” local state; genuine destination-state continuity across machines requires cloning/migrating _dlt_* as well if you rely on those tables.

has_dataset() (schema exists in information_schema.schemata) only gates whether dlt runs CREATE SCHEMA; it is not the mechanism that resets or retains the local cursor.

Do not set dataset_name="public" (or point staging at public via staging_dataset_name_layout) with the flag on — the connector rejects those reserved names. Prefer a dedicated dataset/schema name such as tenant_a or demo.

Replace disposition: in schema mode, truncate/replace uses DELETE … WHERE 1=1 (not TRUNCATE) on both Core and managed, so older Core builds that silently no-op schema-qualified TRUNCATE stay correct. The destination role therefore needs DELETE privilege, and full-table deletes create delete logs.

Firebolt does not support ALTER TABLE … SET SCHEMA, and ALTER TABLE … RENAME TO cannot rename across schemas. To bring existing data from the old public-prefix layout into the new schema, Firebolt’s recommended path is a zero-copy CREATE TABLE … CLONE (metadata-level; near-instant; no extra storage), then drop the old table:

-- CLONE requires the target schema to exist first. Pre-creating it does not
-- reset or retain local dlt state by itself (see state notes above). On a fresh
-- machine you still start with a fresh local cursor; clone destination _dlt_*
-- only if you need those tables elsewhere.
CREATE SCHEMA IF NOT EXISTS tenant_a;
CREATE TABLE tenant_a.orders CLONE public.tenant_a_orders;
-- verify row counts / spot-check, then:
DROP TABLE public.tenant_a_orders;

Repeat per table you want to keep. External tables cannot be cloned — recreate them in the target schema. Firebolt Core does not support CLONE (verified on Core 5.0.1); on Core use CTAS instead:

CREATE TABLE tenant_a.orders AS SELECT * FROM public.tenant_a_orders;
DROP TABLE public.tenant_a_orders;

To start with a fresh local incremental cursor, clear or omit the local pipeline working directory (or run a new pipeline_name) before the first flag-ON run. Dropping only {dataset}._dlt_* in the destination is not equivalent.

Firebolt Core

Variable Required Description
FIREBOLT_CORE_URL yes Core HTTP endpoint (selects Core when set)
FIREBOLT_CORE_DATABASE no Database name (default: firebolt)
FIREBOLT_STAGING_MODE no upload (default)
Variable Description
DLT_LOCAL_STAGING_DIR Local path for upload-mode normalize staging
DLT_DATASET_NAME Default dataset name (default: demo)

Set credentials via environment variables or .dlt/secrets.toml. Do not commit secrets.

Supported capabilities

Feature Support
Loader format Parquet only
Staging upload (HTTP, default) or s3 (COPY INTO)
append Yes
replace truncate-and-insert, insert-from-staging
merge delete-insert (single-table and nested)

Development

Clone the repository and install in editable mode with dev dependencies:

git clone https://github.com/firebolt-db/dlt-firebolt.git
cd dlt-firebolt
pip install -e ".[dev]"
cp .env.example .env   # fill in Firebolt credentials (and S3 for integration tests)
pytest -m "not integration"

Core e2e (upload, no S3):

bash scripts/core_e2e.sh my_dataset          # nested merge
bash scripts/hn_core_e2e.sh hn_blog 30     # blog example

Optional live integration tests (requires Firebolt, S3, and AWS credentials):

FIREBOLT_RUN_INTEGRATION=1 pytest -m integration -v

License

Apache License 2.0. See LICENSE.

Status

Community package maintained by Firebolt. Not part of core dlt.

  • Published on PyPI (pip install dlt-firebolt)
  • Append, merge, and replace dispositions
  • Nested multi-table merge
  • HTTP upload mode (0.3.0+, Firebolt Core)
  • Website integration docs
  • Optional listing on dlt community destinations page

Release files for dlt-firebolt 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for dlt-firebolt 0.4.0
File Size Uploaded
dlt_firebolt-0.4.0.tar.gz 46.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for dlt-firebolt 0.4.0
File Interpreter ABI Platform
dlt_firebolt-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 76.6 kB

Release files / dlt_firebolt-0.4.0.tar.gz

Download URL dlt_firebolt-0.4.0.tar.gz
Size 46.9 kB
Tags Source
SHA-256 checksum
How to use checksums
08ceacc7ecde2f9ebf1a9e6bf5f202b86214e9d87b5eafffdd47531074c1b6dd
BLAKE2b-256 checksum
How to use checksums
31195ba86b8184f635df678beb7d928542d8eae06e510aff88b63af3f25b9a53
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release files / dlt_firebolt-0.4.0-py3-none-any.whl

Download URL dlt_firebolt-0.4.0-py3-none-any.whl
Size 29.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cbf1c85ecd05e7c5f483c064059353478a927f5425f5220a1853c06afdbaa0a1
BLAKE2b-256 checksum
How to use checksums
e26883c3286cccf0ac3eb3f682a4040bd0ea5f808fd64ef65ba23d8f1d3de584
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page