dbx-flame
A configuration-driven data engineering framework for Databricks. Onboarding a new dataset means writing a workflow YAML — never Python.
Why
Most Databricks pipelines grow one notebook per dataset. This framework inverts that: a single typed engine reads a workflow YAML and carries every dataset through the same five layers — Start, Pipeline, Typing, Policies, Output. A new source is a new YAML file, not a new code path.
The project is alpha — see Status for what's implemented today.
Key Capabilities
- Config-driven onboarding — declare an origin, a write verb and a target; the framework validates the combination and runs it. No per-dataset Python.
- Five typed layers — Start → Pipeline → Typing → Policies → Output, each reachable only
through a typed
Contextand a DataFrame, so every layer is independently testable. - Five write verbs —
APPEND,FULL,UPSERT,SCD2,COMPLETE_DELTA— layer-agnostic, so the same verb serves Inbound→Bronze or Bronze→Silver. - Gold is SQL — each Gold table is a materialized view over Silver, in its own
.sqlfile. Silver's MERGEs never block it. See Gold. - A data quality gate, not a bolt-on — every batch is checked against a
Databricks DQX ruleset before it's written;
errorrefuses the batch,warnlogs and lets it through. - Fail fast, fail loud — invalid config, unknown origins, and failed checks stop the run with one aggregated error report. Nothing logs an error and reports success.
- Ships as a wheel — src-layout package deployed via Databricks Asset Bundles and run
through
python_wheel_taskentry points. No notebook logic, nosys.pathhacks.
Quick Look
Two tasks, two YAML blocks — CSV into Bronze, then Bronze into Silver with SCD Type 1
upsert on CLAIM_ID:
# Inbound → Bronze
source.origin: csv
source.path: /Volumes/dev/source_data/inbound/
source.directory: CLAIMS
output.verb: append
output.schema_name: claims
output.table: CLAIMS_BRONZE
# Bronze → Silver
source.origin: delta
source.schema_name: claims
source.table: CLAIMS_BRONZE
output.verb: upsert
output.schema_name: claims
output.table: CLAIMS_SILVER
output.keys: CLAIM_ID
output.event_time.column: __EXPORT_DATE
No code changes for either step — both are entries in a workflow YAML deployed through the bundle. See 03_write_verbs.md for the full verb matrix.
Documentation
- 00_overview.md — goals, principles, glossary, layer diagram
- 01_architecture.md — layers, protocols, execution flow, extension points
- 02_config_schema.md — the typed config schema, parameter by parameter
- 03_write_verbs.md — verb semantics with worked examples
- 04_policies.md — the data quality gate, driven by a DQX ruleset
- 05_testing.md — unit and platform testing
- 06_roadmap.md — phased implementation plan and current status
Development
Targets Linux (native or WSL). Dependencies and the virtualenv are managed with Poetry (2.x).
Prerequisites: a JDK (11 or 17 — required by PySpark) and Python 3.11 exactly (matches Databricks Runtime 15.4 LTS). If your system Python isn't 3.11, install one (e.g. via pyenv or a standalone build) and point Poetry at it.
poetry env use python3.11 # once, to pin the interpreter (path to a 3.11 binary if not on PATH)
make install # runtime + dev dependencies
make test # unit tests with coverage
Poetry keeps the virtualenv outside the project (~/.cache/pypoetry/virtualenvs), so there is
no .venv/ directory to get out of sync with the interpreter actually running the tests.
Make targets
| Target | What it does |
|---|---|
make install |
poetry install — create the venv and install everything |
make test |
poetry run pytest — unit tests with coverage (70% gate) |
make qa |
black ., then flake8, then yamllint . |
make format |
import sort (ruff --select I --fix) + black . |
make build |
poetry build — wheel + sdist into dist/ |
make clean |
remove dist/, caches, __pycache__ |
make with no target prints this list.
Note that make qa rewrites files — black . formats in place rather than checking.
Use poetry run black --check . for a read-only pass.
Tool configuration
All tool config lives in pyproject.toml except flake8, which cannot read it — flake8's
settings are in .flake8.
mypy and ruff are installed and configured but are not part of make qa yet.
Testing
make test runs unit tests against a local Spark + Delta session. Platform tests are real
Databricks jobs under platform_tests/ — each one generates its own
fixtures and asserts the resulting tables:
databricks bundle deploy -t dev_01 -p <profile>
databricks bundle run integration_test_suite -t dev_01 -p <profile>
See 05_testing.md for the full strategy.
On Windows (WSL)
Spark doesn't run natively on Windows, so do all of the above inside WSL (e.g. Ubuntu 24.04), not PowerShell/cmd — the Databricks CLI profile lives there too:
wsl -d Ubuntu-24.04 -- bash -lc 'cd /path/to/dbx-flame && poetry run pytest'
Deploying to Databricks
The bundle in databricks.yml is configured to work with a Python wheel: it
builds the wheel, then deploys the bundle. Authenticate with a CLI profile or with
DATABRICKS_HOST / DATABRICKS_TOKEN:
databricks bundle deploy
Traceability
Every run writes to an audit log — not vendor telemetry, a trail in your own Databricks
catalog (monitoring_{env}.audit.logs). It's not a bolt-on: Context carries the logger as a
required field, and the audit write is deliberately the first thing a task does — if it can't
land, the run fails before touching any data, rather than processing a batch it can't account
for.
Each row ties a log level and an event to the exact job, run and task that produced it:
| Columns | |
|---|---|
| What ran | name, source |
| When | __workflow_id / __workflow_run_id, __task_key / __task_run_id, time_stamp |
| Where | catalog, schema, table |
| Outcome | type (INFO / WARNING / ERROR), total, description, metadata |
Nothing leaves the workspace, and there's no flag to turn it off. Full contract: 00_overview.md.
Status
Start, Pipeline (CSV + JSON + Delta), Typing, Policies (DQX), and all five write verbs are implemented and tested. The SAS origin and the Gold example view are not implemented yet. Full phase-by-phase status: 06_roadmap.md.
License
Apache 2.0 — see LICENSE.
Release files for dbx-flame 0.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dbx_flame-0.0.1.tar.gz | 43.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dbx_flame-0.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 102.9 kB
Release files / dbx_flame-0.0.1.tar.gz
| Download URL | dbx_flame-0.0.1.tar.gz |
|---|---|
| Size | 43.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1c83e5dd95b69771af7d403200bbd8d4f709c5535bec30534a770f05500fd9e6
|
|
BLAKE2b-256 checksum How to use checksums |
eec1844c244f41354f264b68c8e5b18dfba2c8687fd9c090cc9b4ba2d7abd6c2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / dbx_flame-0.0.1-py3-none-any.whl
| Download URL | dbx_flame-0.0.1-py3-none-any.whl |
|---|---|
| Size | 59.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
07844b73cc7280831377125c1624f471aaf286e756325be3e69c245ee5556ad2
|
|
BLAKE2b-256 checksum How to use checksums |
be448d9c90c6994cd28648907fc8306a12b7fa4493f862bde61d900c2311ea43
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log