Declarative ClickHouse streaming pipeline deployment for staged backfill, audit, and publish workflows.
streambuild is aimed at streaming data teams who want dbt-like authored models, but with deployment semantics that fit live ClickHouse pipelines:
- plan rebuilds conservatively
- create staged shadow objects
- backfill history into the staged path
- audit staged readiness
- publish by switching stable logical views
The current product is centered on ClickHouse and streaming replay semantics. Kafka-backed sources work today, and adopted external streaming tables are now supported as replay roots.
Current Status
Current implemented workflow:
stb planstb buildstb backfillstb audit backfillstb publishstb doctorstb repair active-viewstb reconcilestb compilestb janitor
Current rollout model:
planis read-onlybackfillstarts a real staged deploymentaudit backfillinspects staged readinesspublishswitches stable logical views to staged physical tables
Installation
Requirements:
- Python
>=3.12 - ClickHouse
Local dev install:
uv sync
Run the CLI with:
uv run stb --help
Project Shape
StreamBuild projects are authored as a project root plus pipeline folders.
streambuild_project.toml
sources/
orders.yml
macros/
common.py
pipelines/
orders/
orders_enriched.sql
order_rollups.sql
Rules:
- each direct child folder under
pipelines/is one pipeline - recursive
*.sqlfiles under that folder belong to that pipeline - pipeline name is inferred from the folder name
- pipeline source is inferred transitively from model driving inputs
- model name is inferred from the SQL filename stem
- optional
pipeline.tomlstores pipeline-wide virtual-environment policy
Macros
Public Python modules under macros/ are loaded once per project analysis. Functions
defined by those modules are available in authored model, test, and audit SQL as
@function_name(...). Imported functions, async functions, __init__.py, and modules
or directories whose names start with _ are not registered.
from streambuild.compiler.macros.models import MacroContext
def qualified_source(ctx: MacroContext, table_name: str) -> str:
return f"{ctx.database}.{table_name}"
SELECT * FROM @qualified_source("orders")
Macro modules are trusted project code, not a sandbox: module-level code runs during
analysis, and a macro may perform anything allowed to that Python process. Calls accept
only nested Python literals (str, bool, int, float, None, lists, tuples, and
dictionaries with scalar keys) plus nested macro results. A first parameter named ctx
must be annotated as MacroContext; StreamBuild supplies its immutable project target,
adapter, database, virtual-environment, and variable values. Direct SQL macro calls must
return strings. Errors report both the authored SQL call and the defining macro source.
Project Config
Committed project configuration lives in streambuild_project.toml. Developer-specific
overrides may live in the gitignored streambuild_local.toml.
name = "orders_project"
default_target = "dev"
[settings]
virtual_environments = true
[connection]
host = "localhost"
port = 8123
username = "clickhouse"
password = "${ENV:CLICKHOUSE_PASSWORD}"
[defaults]
managed_source_ttl = "_replay_landed_at + INTERVAL 14 DAY"
[naming]
table_prefix = "tbl__"
view_prefix = "view__"
[targets.dev]
database = "analytics"
Notes:
nameanddefault_targetare required;adapterdefaults toclickhouse- target selection is CLI
--target, localtarget, then projectdefault_target - CLI
--varsaccepts one JSON object for${name}interpolation - connection templates are expanded only for commands that connect
- metadata lives in the same database by default
- model relation names use the model's exact
relation_name, then pipeline, project, and built-in kind-specific prefixes - connection precedence is CLI flags, fixed
STREAMBUILD_CLICKHOUSE_*environment variables, local config, selected target, then project config
Pipeline Sources
Reusable replay-driving sources live under sources/*.yml. StreamBuild follows each table model's
__source(...) or untyped __ref(...) driving input until it reaches a registered source. Every
pipeline containing tables must resolve to exactly one source. Terminal views do not participate in
source inference, so a view-only pipeline is valid and source-less.
Managed Kafka Landing
sources:
- kind: kafka
name: orders
broker_list: kafka:9092
topic: source.orders.created
ttl: _replay_landed_at + INTERVAL 30 DAY
replay_boundary:
mode: offsets
This is the managed source shape:
- StreamBuild creates the Kafka table
- StreamBuild creates the raw landing table and landing MV
- source
ttloverrides[defaults].managed_source_ttl; omitting both keeps data indefinitely - downstream models usually read the source via
__source("orders")
Adopted External Source
sources:
- kind: stream_table
name: orders
table_name: orders_existing
replay_boundary:
mode: offsets
columns:
_replay_partition: event_partition
_replay_offset: event_offset
_replay_timestamp: event_timestamp
This is the adopted-source shape:
- StreamBuild does not create the source table
- the source table must already exist in the resolved project database
table_namemust currently be a bare table name- replay boundary columns are validated against the live table schema during planning/runtime commands
Current replay-boundary rules for adopted sources:
mode: offsetsrequirespartition,offset, andtimestampmode: offsetsdoes not allowlanded_atmode: timestamprequirestimestampmode: timestampdoes not allowlanded_atmode: cursorrequirescursorandtimestamp
Currently supported external-source replay boundary modes:
-
offsets -
timestamp -
cursor
Virtual-environment projects can choose change-driven replay independently from the fallback used when bounded replay cannot preserve aggregate history:
bounded_replay_fallback = "bounded_without_history"
[replay_on_change]
breaking = "full"
non_breaking = "bounded-7d"
This optional pipeline.toml sits directly in the pipeline directory. The same policies can be
defaults in streambuild_project.toml and overrides in a model MODEL(...) header. They are
rejected when settings.virtual_environments is false.
Models
Each SQL model starts with a MODEL (...) header. Models default to streaming tables.
MODEL (
engine "MergeTree()",
order_by ["order_id", "_replay_partition", "_replay_offset"],
partition_by "toYYYYMM(event_at)",
ttl "event_at + INTERVAL 30 DAY",
settings (
index_granularity 8192,
),
replay_anchor auto,
);
SELECT
CAST(order_id AS UInt64) AS order_id,
CAST(event_at AS DateTime64(3)) AS event_at,
CAST(_replay_partition AS Int32) AS _replay_partition,
CAST(_replay_offset AS Int64) AS _replay_offset
FROM __source("orders")
Notes:
- the driving replay input may be declared with
__source(...)for source roots or__ref(...)for managed upstream models - additional managed dependencies are declared with
__ref(...) - for table models only, additional
__ref(...)dependencies must declareref_type - header fields use SQLBuild syntax: whitespace-separated
key valueentries, lists in[...], and nested mappings in(...) - omitted SQL storage settings default to
engine "MergeTree()"andorder_by ["_replay_timestamp"] - both
CAST(expr AS Type)andexpr::Typeare accepted
Terminal Views
An ordinary query view uses kind view and may read any number of upstream sources or models:
MODEL (
kind view,
relation_name customer_orders,
);
SELECT
orders.order_id::UInt64 AS order_id,
payments.amount_cents::UInt64 AS amount_cents
FROM __ref("orders") AS orders
JOIN __ref("payments") AS payments USING (order_id)
Views have no driving input, storage settings, replay policy, or replay work. View refs reject
ref_type; every __source(...) and __ref(...) is an ordinary query dependency. A view must be a
terminal node across the complete project graph: no table or view model may reference it. Tests and
audits may target it. relation_name is an exact warehouse relation override for either model kind;
without one, table and view names use the effective table_prefix or view_prefix from optional
pipeline [naming], project [naming], then the tbl__ and view__ defaults. kafka__, raw__,
and mv__ remain framework-reserved.
Replay Lineage
StreamBuild exposes a normalized replay lineage surface.
Current intent:
_replay_*is the normalized source-agnostic replay vocabulary
Current generic replay columns:
_replay_partition_replay_offset_replay_timestamp_replay_landed_at_replay_cursor
Current behavior:
- managed Kafka landing populates the normalized
_replay_*lineage columns directly - adopted sources map declared physical source columns into the normalized replay surface
- downstream managed outputs should preserve
_replay_*when they need replay lineage
Core Commands
From a project directory:
uv run stb plan
uv run stb build
uv run stb backfill
uv run stb audit backfill
uv run stb publish
uv run stb doctor
uv run stb repair active-view --table tbl__orders
uv run stb reconcile
uv run stb compile
uv run stb janitor
From outside the project directory:
uv run stb plan --project-dir examples/orders_demo
Compile Artifacts
stb compile writes artifacts under project-level target/.
Static compile products and runtime evidence have separate owners:
target/
manifest.json
streambuild_dag.json
compiled/
models/<pipeline>/
resources/
sources/<source>/
models/<pipeline>/
workflows/<pipeline>/
steps/
workflow.sql
workflow.json
audits/
tests/
run/
tests/
stb compile atomically replaces only the static owners and never writes under
target/run/. Runtime commands own their command-specific subtrees.
The compile manifest includes:
- resolved database
- relations
- source metadata
- model specs
- logical tests and audits
- realized adapter resources
- every emitted static artifact path
- workflow paths and logical DAG identity
Example
See examples/orders_demo/ for a runnable local demo using:
- Redpanda
- ClickHouse
- a synthetic producer
- a real
streambuildproject
Demo README:
examples/orders_demo/README.md
Development
Useful commands:
make format
make lint
make type
make test
make test-all
make check
make verify
Current meanings:
make check: fast structural and static validationmake verify: full validation including tests
Testing
The repo uses:
- unit tests under
tests/unit - integration tests under
tests/integration - end-to-end tests under
tests/e2e
Recent coverage includes:
- staged backfill / audit / publish flows
- active-view diagnosis and repair
- adopted external replay sources
- normalized replay lineage behavior
Scope Notes
Current intentional limitations:
- ClickHouse-only runtime
- external adopted sources must resolve in the project database
- managed Kafka sources support
offsets,timestamp, andlanded_at; adopted relations supportoffsets,timestamp, andcursor
This repo is actively evolving around staged rollout correctness, replay semantics, and migration/adoption support for existing ClickHouse streaming tables.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file streambuild-0.3.0.tar.gz.
File metadata
- Download URL: streambuild-0.3.0.tar.gz
- Upload date:
- Size: 727.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
27faa7de41517de4e7ec6c9bbce03876a3165d5ae7c9e56ecf550e9914aa8322
|
|
| MD5 |
fc994e44a9e0df71f1942d8228c6d2cf
|
|
| BLAKE2b-256 |
377ec1b821a42fc166e8b6820c00e8513372ab75201a267e35c8219b270f6e32
|
Provenance
The following attestation bundles were made for streambuild-0.3.0.tar.gz:
Publisher:
publish.yml on chio-labs/streambuild
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
streambuild-0.3.0.tar.gz -
Subject digest:
27faa7de41517de4e7ec6c9bbce03876a3165d5ae7c9e56ecf550e9914aa8322 - Sigstore transparency entry: 2309775723
- Sigstore integration time:
-
Permalink:
chio-labs/streambuild@81054642b0f89eb07fee04142ff18b99a9f99507 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/chio-labs
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@81054642b0f89eb07fee04142ff18b99a9f99507 -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file streambuild-0.3.0-py3-none-any.whl.
File metadata
- Download URL: streambuild-0.3.0-py3-none-any.whl
- Upload date:
- Size: 462.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f60a450918b0fb889c221faa8a99b4354cd457010d62f2ec086f4361214a9309
|
|
| MD5 |
d66b0bf20df385c297ff0b2eebd5f346
|
|
| BLAKE2b-256 |
ace9be7770c2f46a55f32ebf1cc4a0091842b4756d35f857511127a922ace066
|
Provenance
The following attestation bundles were made for streambuild-0.3.0-py3-none-any.whl:
Publisher:
publish.yml on chio-labs/streambuild
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
streambuild-0.3.0-py3-none-any.whl -
Subject digest:
f60a450918b0fb889c221faa8a99b4354cd457010d62f2ec086f4361214a9309 - Sigstore transparency entry: 2309775728
- Sigstore integration time:
-
Permalink:
chio-labs/streambuild@81054642b0f89eb07fee04142ff18b99a9f99507 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/chio-labs
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@81054642b0f89eb07fee04142ff18b99a9f99507 -
Trigger Event:
workflow_dispatch
-
Statement type: