dlt-firebolt
Community dlt destination for Firebolt.
Load data into Firebolt with dlt using direct HTTP upload (default) or S3 staging + COPY INTO for large loads.
Requires dlt-firebolt 0.3.0+ for upload mode and simplified Core/managed configuration.
Two ways to load
| Mode | Best for | Firebolt load path |
|---|---|---|
upload (default) |
Firebolt Core, local dev, quick starts | HTTP multipart → READ_PARQUET('upload://…') |
s3 |
Managed Firebolt production, large bulk loads | Parquet on S3 → COPY INTO |
On managed Firebolt today, set FIREBOLT_STAGING_MODE=s3. Upload is the code default but the managed engine does not accept multipart upload yet; you will get a clear error if you use upload mode there.
Install
pip install "dlt-firebolt>=0.3.0"
Requires Python 3.10+.
Register the destination once in your project:
import firebolt_dest # registers destination="firebolt"
Prerequisites
Upload mode (Core / local)
- Firebolt Core running locally (or another environment that supports
upload://), or a managed account once upload is supported there.
No S3 bucket, external location, or AWS credentials required for the Firebolt load step. dlt still writes Parquet to a local staging directory during normalize.
S3 mode (managed production)
- Firebolt service account with access to your database and engine.
- S3 bucket for dlt filesystem staging.
- Firebolt external location for that bucket prefix:
CREATE LOCATION "your_location_name" WITH
SOURCE = 'CLOUD_STORAGE'
URL = 's3://your-bucket/your-prefix/'
CREDENTIALS = (AWS_ROLE_ARN = 'arn:aws:iam::...:role/...');
Set FIREBOLT_S3_LOCATION_NAME to the exact location name. Set S3_PREFIX to the same key prefix as the LOCATION URL path for classic single-tenant setups (COPY PATTERN is relative to the LOCATION path). Set FIREBOLT_STAGING_MODE=s3.
Multi-tenant (one LOCATION for all tenants): create the LOCATION at the bucket root (URL = 's3://your-bucket/'), set FIREBOLT_S3_LOCATION_URL=s3://your-bucket/ (or s3_location_url in TOML — see below), and set a per-tenant staging prefix so objects land under s3://bucket/<tenant>/.... With make_firebolt_pipeline(from_secrets=False), that prefix is S3_PREFIX=tenant-a. With from_secrets=True / a bare firebolt() destination, steer staging with a tenant-specific [destination.filesystem] bucket_url instead — S3_PREFIX is not read on that path. COPY PATTERN keeps the tenant segment (e.g. tenant-a/dlt/staging/...). Classic prefix-LOCATION setups that omit FIREBOLT_S3_LOCATION_URL still fall back to s3_prefix as the LOCATION path. Migration note: loads that previously relied on the old silent basename fallback (object key not under the LOCATION/s3_prefix path) now hard-fail with TerminalValueError — fix the LOCATION URL / prefix alignment rather than expecting a filename-only PATTERN. A LOCATION at the bucket root grants its credentials read access to every prefix in the bucket; PATTERN is not an authorization boundary. Bucket-root suits multiple prefixes inside a single trust boundary — for real tenant isolation use per-tenant LOCATIONs or buckets.
The machine running dlt needs AWS credentials that can write to the staging bucket (via environment variables, an AWS profile, or an attached IAM role). Firebolt reads the staged files from the external location you configured, not the runner's AWS identity. Staging Parquet is not automatically deleted after COPY INTO, so add an S3 lifecycle rule or a periodic cleanup if you don't want objects to accumulate.
Quick start
Firebolt Core (upload, no S3)
export FIREBOLT_CORE_URL=http://localhost:3473
import dlt
import firebolt_dest
from firebolt_dest.configuration import make_firebolt_pipeline
@dlt.resource(name="orders", write_disposition="append")
def orders():
yield {"order_id": 1, "customer": "Acme"}
pipeline = make_firebolt_pipeline(
pipeline_name="my_pipeline",
dataset_name="my_dataset",
)
pipeline.run(orders(), loader_file_format="parquet")
Managed Firebolt (S3)
Set service-account credentials and FIREBOLT_STAGING_MODE=s3 (see Configuration). The same pipeline code applies; only env vars change.
Tables are created as {dataset}_{table} in the public schema (for example my_dataset_orders). To use a real Firebolt schema instead (my_dataset.orders), set FIREBOLT_USE_SCHEMA_PER_DATASET=true — see Configuration.
Using .dlt/secrets.toml
Environment variables are the primary path used in the e2e scripts. You can also load configuration from .dlt/secrets.toml with from_secrets=True (verified on Firebolt Core with upload mode; see .dlt/secrets.toml.example).
Note: In dlt's credential block, host is the Firebolt database name and database is the Firebolt engine name (matching FIREBOLT_DATABASE and FIREBOLT_ENGINE, not swapped).
pipeline = make_firebolt_pipeline(
pipeline_name="my_pipeline",
dataset_name="my_dataset",
from_secrets=True,
)
Example for managed Firebolt (S3 mode):
[destination.firebolt]
staging_mode = "s3"
s3_location_name = "your_location_name"
s3_prefix = "dlt-landing"
# use_schema_per_dataset = false # opt-in real schemas; default off
# Multi-tenant bucket-root LOCATION (PATTERN relative to this URL):
# s3_location_url = "s3://example-bucket/"
# Steer staging with a tenant-specific filesystem bucket_url (S3_PREFIX env is
# NOT read when from_secrets=True / destination="firebolt"):
# (see bucket_url below — e.g. s3://example-bucket/tenant-a/dlt/staging)
[destination.firebolt.credentials]
host = "YOUR_FIREBOLT_DATABASE"
database = "YOUR_FIREBOLT_ENGINE"
username = "YOUR_CLIENT_ID"
password = "YOUR_CLIENT_SECRET"
account_name = "YOUR_ACCOUNT_NAME"
[destination.filesystem]
bucket_url = "s3://example-bucket/dlt-landing/dlt/staging"
# Multi-tenant example: bucket_url = "s3://example-bucket/tenant-a/dlt/staging"
Which config path reads which variable
| Setting | make_firebolt_pipeline(from_secrets=False) |
from_secrets=True / bare firebolt() |
|---|---|---|
| LOCATION URL for PATTERN | FIREBOLT_S3_LOCATION_URL → s3_location_url |
s3_location_url in [destination.firebolt] (or DESTINATION__FIREBOLT__S3_LOCATION_URL) |
| Staging write prefix | S3_BUCKET + S3_PREFIX → filesystem bucket_url |
[destination.filesystem] bucket_url only (S3_PREFIX is not read) |
| LOCATION name | FIREBOLT_S3_LOCATION_NAME |
s3_location_name |
See .dlt/secrets.toml.example for a Firebolt Core upload template and the managed S3 fields including s3_location_url.
With from_secrets=True (or a plain destination="firebolt" pipeline), set the flag in TOML as use_schema_per_dataset = true under [destination.firebolt], or via the dlt env alias DESTINATION__FIREBOLT__USE_SCHEMA_PER_DATASET=true. FIREBOLT_USE_SCHEMA_PER_DATASET is read only when make_firebolt_pipeline(..., from_secrets=False) builds the destination from env.
Configuration
Managed Firebolt
| Variable | Required | Description |
|---|---|---|
FIREBOLT_CLIENT_ID |
yes | Service account client ID |
FIREBOLT_CLIENT_SECRET |
yes | Service account secret |
FIREBOLT_ACCOUNT_NAME |
yes | Firebolt account name |
FIREBOLT_DATABASE |
yes | Target database |
FIREBOLT_ENGINE |
yes | Engine name |
FIREBOLT_STAGING_MODE |
no | s3 for managed production (recommended) |
FIREBOLT_S3_LOCATION_NAME |
s3 mode | Firebolt external location name |
FIREBOLT_S3_LOCATION_URL |
no | LOCATION URL for PATTERN (e.g. s3://bucket/); empty → use s3_prefix as LOCATION path. Read by make_firebolt_pipeline(from_secrets=False); for secrets/TOML use s3_location_url / DESTINATION__FIREBOLT__S3_LOCATION_URL |
S3_BUCKET |
s3 mode | Staging bucket (from_secrets=False only; secrets path uses filesystem bucket_url) |
S3_PREFIX |
no | Staging key prefix (default: dlt-landing); per-tenant prefix when LOCATION is at bucket root (from_secrets=False only) |
FIREBOLT_USE_SCHEMA_PER_DATASET |
no | true to map dataset_name to a real Firebolt schema (default: off — tables stay public.{dataset}_{table}). Read only by make_firebolt_pipeline(from_secrets=False); with from_secrets=True / bare firebolt() set use_schema_per_dataset in TOML or DESTINATION__FIREBOLT__USE_SCHEMA_PER_DATASET |
The destination resolves the engine URL from your account; you do not set an HTTP endpoint manually.
Schema-per-dataset (opt-in)
Default is off. Existing pipelines keep writing public.{dataset}_{table} and store dlt state tables (_dlt_loads, _dlt_pipeline_state, _dlt_version) under that public-prefix naming. Enabling the flag is a deliberate breaking change for that layout:
| Flag off (default) | Flag on | |
|---|---|---|
| Table name | public.tenant_a_orders |
tenant_a.orders |
| Staging | public.tenant_a_staging_* |
schema tenant_a_staging |
| dlt state | public.{dataset}__dlt_* |
{dataset}._dlt_* |
Where dlt state actually lives: the incremental cursor / load state that “fresh vs retained” refers to is primarily local pipeline state on the machine running dlt (under the pipeline working directory). Destination {dataset}._dlt_* tables are a separate store; dropping them alone does not reset the local cursor.
First flag-ON run — two different situations:
- Fresh machine (no local pipeline state): the destination schema being pre-created or not does not change local state — you still get a fresh local cursor. Destination
_dlt_*tables appear only after the first successful load. - Same machine with existing local state: that local state is what continues the cursor. Pre-creating the destination schema (for CLONE/CTAS) does not by itself “retain” local state; genuine destination-state continuity across machines requires cloning/migrating
_dlt_*as well if you rely on those tables.
has_dataset() (schema exists in information_schema.schemata) only gates whether dlt runs CREATE SCHEMA; it is not the mechanism that resets or retains the local cursor.
Do not set dataset_name="public" (or point staging at public via staging_dataset_name_layout) with the flag on — the connector rejects those reserved names. Prefer a dedicated dataset/schema name such as tenant_a or demo.
Replace disposition: in schema mode, truncate/replace uses DELETE … WHERE 1=1 (not TRUNCATE) on both Core and managed, so older Core builds that silently no-op schema-qualified TRUNCATE stay correct. The destination role therefore needs DELETE privilege, and full-table deletes create delete logs.
Firebolt does not support ALTER TABLE … SET SCHEMA, and ALTER TABLE … RENAME TO cannot rename across schemas. To bring existing data from the old public-prefix layout into the new schema, Firebolt’s recommended path is a zero-copy CREATE TABLE … CLONE (metadata-level; near-instant; no extra storage), then drop the old table:
-- CLONE requires the target schema to exist first. Pre-creating it does not
-- reset or retain local dlt state by itself (see state notes above). On a fresh
-- machine you still start with a fresh local cursor; clone destination _dlt_*
-- only if you need those tables elsewhere.
CREATE SCHEMA IF NOT EXISTS tenant_a;
CREATE TABLE tenant_a.orders CLONE public.tenant_a_orders;
-- verify row counts / spot-check, then:
DROP TABLE public.tenant_a_orders;
Repeat per table you want to keep. External tables cannot be cloned — recreate them in the target schema. Firebolt Core does not support CLONE (verified on Core 5.0.1); on Core use CTAS instead:
CREATE TABLE tenant_a.orders AS SELECT * FROM public.tenant_a_orders;
DROP TABLE public.tenant_a_orders;
To start with a fresh local incremental cursor, clear or omit the local pipeline working directory (or run a new pipeline_name) before the first flag-ON run. Dropping only {dataset}._dlt_* in the destination is not equivalent.
Firebolt Core
| Variable | Required | Description |
|---|---|---|
FIREBOLT_CORE_URL |
yes | Core HTTP endpoint (selects Core when set) |
FIREBOLT_CORE_DATABASE |
no | Database name (default: firebolt) |
FIREBOLT_STAGING_MODE |
no | upload (default) |
| Variable | Description |
|---|---|
DLT_LOCAL_STAGING_DIR |
Local path for upload-mode normalize staging |
DLT_DATASET_NAME |
Default dataset name (default: demo) |
Set credentials via environment variables or .dlt/secrets.toml. Do not commit secrets.
Supported capabilities
| Feature | Support |
|---|---|
| Loader format | Parquet only |
| Staging | upload (HTTP, default) or s3 (COPY INTO) |
append |
Yes |
replace |
truncate-and-insert, insert-from-staging |
merge |
delete-insert (single-table and nested) |
Development
Clone the repository and install in editable mode with dev dependencies:
git clone https://github.com/firebolt-db/dlt-firebolt.git
cd dlt-firebolt
pip install -e ".[dev]"
cp .env.example .env # fill in Firebolt credentials (and S3 for integration tests)
pytest -m "not integration"
Core e2e (upload, no S3):
bash scripts/core_e2e.sh my_dataset # nested merge
bash scripts/hn_core_e2e.sh hn_blog 30 # blog example
Optional live integration tests (requires Firebolt, S3, and AWS credentials):
FIREBOLT_RUN_INTEGRATION=1 pytest -m integration -v
License
Apache License 2.0. See LICENSE.
Status
Community package maintained by Firebolt. Not part of core dlt.
- Published on PyPI (
pip install dlt-firebolt) - Append, merge, and replace dispositions
- Nested multi-table merge
- HTTP upload mode (0.3.0+, Firebolt Core)
- Website integration docs
- Optional listing on dlt community destinations page
Release files for dlt-firebolt 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dlt_firebolt-0.4.0.tar.gz | 46.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dlt_firebolt-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 76.6 kB
Release files / dlt_firebolt-0.4.0.tar.gz
| Download URL | dlt_firebolt-0.4.0.tar.gz |
|---|---|
| Size | 46.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
08ceacc7ecde2f9ebf1a9e6bf5f202b86214e9d87b5eafffdd47531074c1b6dd
|
|
BLAKE2b-256 checksum How to use checksums |
31195ba86b8184f635df678beb7d928542d8eae06e510aff88b63af3f25b9a53
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.5
|
Release files / dlt_firebolt-0.4.0-py3-none-any.whl
| Download URL | dlt_firebolt-0.4.0-py3-none-any.whl |
|---|---|
| Size | 29.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
cbf1c85ecd05e7c5f483c064059353478a927f5425f5220a1853c06afdbaa0a1
|
|
BLAKE2b-256 checksum How to use checksums |
e26883c3286cccf0ac3eb3f682a4040bd0ea5f808fd64ef65ba23d8f1d3de584
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.5
|