Python/SQL-first data platform: transformation, built-in orchestration, and durable streaming ingestion
Project description
interlace
Python/SQL-first data platform: transformation, orchestration, and durable streaming — one process.
interlace is an independent, MIT-licensed alternative to dbt/SQLMesh that also replaces the
orchestrator (no Airflow) and the ingestion layer (Cloudflare-Pipelines-style durable streams).
Models are .sql files or Python functions; state is versioned snapshots with virtual
environments and a terraform-style plan/apply; everything runs in a single daemon on
DuckDB + DuckLake by default.
Status: 1.0. Requires Python 3.12+. The package is published to PyPI as
interlaced; the import name and CLI areinterlace.
uv pip install "interlaced[service]" # extras: service, adbc, postgres, polars, pandas, all
Sixty seconds
interlace init my-project && cd my-project
interlace plan # terraform-style preview: added / breaking / non-breaking / reuse
interlace apply # build changed models, run checks, promote the environment
interlace serve # the daemon: web UI (/ui) + HTTP API + scheduler + streams, one process
Every model builds into a fingerprinted physical table (interlace__main.orders__a1b2c3);
environments are views over those tables, so promotion and rollback are atomic view swaps and a
dev environment reuses prod's tables for free. Production is the unprefixed namespace —
consumers query main.orders; sandboxes are prefixed (dev__main.orders). Commands default to
prod; pass --env dev while developing.
Models
SQL — a file per model; upstreams referenced by model name, dependencies inferred by parsing (sqlglot), config in a leading comment block:
/* interlace:
strategy: scd_type_2
key: customer_id
schedule: {cron: "0 * * * *"}
checks:
- not_null: customer_id
- unique: customer_id
*/
SELECT customer_id, name, tier FROM raw_customers
Python — functions whose parameters name their upstreams; data crosses as Arrow (never pandas), streamed with bounded memory:
from interlace import model
@model(strategy="merge_by_key", key="order_id", cursor="updated_at")
def orders(cursor, this):
"""Incremental API extract: `cursor` is max(updated_at) already in the
warehouse (None on first run); `this` is the previous materialisation."""
rows = fetch_orders(since=cursor) # your code
return pyarrow.Table.from_pylist(rows) # or RecordBatchReader / generator of batches
Strategies: full, view, ephemeral (CTE-inlined), merge_by_key (upsert),
full_merge (full-state source applied as a minimal diff), incremental_by_time
(windowed, interval-ledger backfill/catchup), scd_type_2 (history with validity windows).
Plan / apply
$ interlace plan
Model Change Category Build
orders modified non_breaking rebuild
order_stats modified non_breaking reuse <- output provably identical: not rebuilt
- Changes classify breaking / non-breaking / forward-only; a plan with breaking changes
refuses to apply without
--force. Downstream models whose output is provably identical (column-pruned impact analysis) reuse their existing tables instead of rebuilding — an improvement over model-granular invalidation. apply --forward-onlylets history-keeping models (scd2/merge/incremental) survive a definition change: the existing table is copied to the new version, the new logic applies to the copy going forward, and checks gate before views move.- Checks gate promotion: 10 built-in types (not_null, unique, accepted_values, row_count,
freshness, expression, relationships, pattern, range, sql) plus
@checkPython functions — an error-severity failure blocks before the environment view moves.interlace checks runre-runs them ad hoc against any environment's promoted tables. interlace gcremoves snapshots no environment references (reference-aware: tables shared through reuse survive).
Streaming
Declare a stream; POST to it; rows are durable (SQLite WAL log) before the 200, deduplicated by
idempotency key, and materialized exactly-once into streams.<name> — a micro-batch flusher
commits the data and the watermark in one warehouse transaction, and SQL models just read the
table. A flush triggers the models that consume the stream.
from interlace import stream
@stream("orders", schema={"order_id": "string", "total": "double"},
idempotency_key="order_id", retention="7d", on_schema_drift="evolve")
def orders(event): ...
curl -X POST localhost:8000/streams/orders -d '{"order_id": "o1", "total": 49.5}'
Schema drift is yours to choose: reject (400), evolve (new columns appear), or
quarantine (bad events divert to <stream>__quarantine). When the warehouse falls behind,
publishes get 429 backpressure instead of unbounded backlog.
Reverse ETL
Attach external databases and deliver model results into them — the live table is never dropped, keyed modes reuse the same merge strategies:
# interlace.yaml
attach:
crm: "postgres:host=... dbname=crm"
/* interlace: {export: {to: table, target: crm.public.accounts, mode: merge_by_key, key: id}} */
SELECT id, tier, lifetime_value FROM account_summary
File exports (to: parquet|csv|json) work the same way. Sinks are environment-gated: by
default the side effect fires only from prod — a dev apply never writes to a live external
table (opt in with environments: [dev, prod]).
Multi-engine
Models run on named engines: DuckDB/DuckLake by default, Postgres natively over ADBC
(pip install 'interlaced[adbc]'), with per-model pinning:
engines:
pg: {type: postgres, database: "${PG_DSN}"}
/* interlace: {engine: pg, strategy: merge_by_key, key: id} */
Strategies execute inside the pinned engine (no DuckDB middleman); cross-engine dependencies
appear as explicit transfer lines in the plan and move as Arrow (or a federated ATTACH
fast lane when possible). Contract: docs/architecture/MULTI_ENGINE.md.
The daemon
interlace serve runs everything in one process:
- the web UI at
/ui(in-package, zero build step) — ten views: overview, lineage canvas with column-level tracing, models, plan/apply with SQL diffs, live runs, query console, streams, checks, environments, and system — live over SSE; - the HTTP API (Litestar + msgspec, OpenAPI at
/schema/scalar) with the same surface as the CLI: plan/apply, runs, checks, streams, engines, schedules, lineage, query, gc; - the scheduler: cron/interval triggers enqueue onto a durable run queue (leases,
retries, cooperative cancellation —
interlace cancel <id>orPOST /runs/{id}/cancel); - stream ingestion and retention.
Scoped API keys (interlace apikey create ci --scope read) lock it down; a durable event log
backs GET /events/stream (SSE with Last-Event-ID replay).
Add --quack quack:localhost:4213 to serve the warehouse itself over DuckDB's quack protocol —
other processes (CLI runs, ad-hoc DuckDB clients) then share it concurrently by setting
database: quack:localhost:4213.
Architecture in five lines
- The IR is a sqlglot AST; the wire format is an Arrow RecordBatchReader; strategies
are AST builders and dialect appears only at
transpile(). - Storage defaults to DuckLake (Parquet + SQL catalog) opened as DuckDB's primary database.
- Control plane (snapshots, intervals, queue, events, keys) is SQLite WAL; Postgres is the scale-out swap.
- Streams live in their own durable log; the materializer commits data + watermark in one warehouse transaction — exactly-once without distributed coordination.
- No Jinja, no pandas in core, no external orchestrator.
The full design rationale lives in docs/architecture/v2-design.md.
Development
Toolchain is pinned with proto, tasks run via
moon, uv owns the virtualenv:
proto install
moon run interlace:sync # install deps
moon run interlace:test # 350+ tests
moon run interlace:check # black + ruff (CI equivalent)
moon run interlace:typecheck # mypy
MIT licensed.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file interlaced-1.0.1.tar.gz.
File metadata
- Download URL: interlaced-1.0.1.tar.gz
- Upload date:
- Size: 176.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3c0a81e8ade0ae4b84f65f856aae3706fa1eeba176c1fb353681c4cf91b198b6
|
|
| MD5 |
b545c342e99462c72d97ff409dc0f393
|
|
| BLAKE2b-256 |
1a933bfa05f528d879a454a0640051e2475947f67b4a0865b36d9649eb5c6dfe
|
Provenance
The following attestation bundles were made for interlaced-1.0.1.tar.gz:
Publisher:
publish.yml on interlace-sh/interlace
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
interlaced-1.0.1.tar.gz -
Subject digest:
3c0a81e8ade0ae4b84f65f856aae3706fa1eeba176c1fb353681c4cf91b198b6 - Sigstore transparency entry: 2335601195
- Sigstore integration time:
-
Permalink:
interlace-sh/interlace@1348055fde2748fba65b24af55c95d600495980d -
Branch / Tag:
refs/tags/v1.0.1 - Owner: https://github.com/interlace-sh
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@1348055fde2748fba65b24af55c95d600495980d -
Trigger Event:
push
-
Statement type:
File details
Details for the file interlaced-1.0.1-py3-none-any.whl.
File metadata
- Download URL: interlaced-1.0.1-py3-none-any.whl
- Upload date:
- Size: 211.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a3703a3133f46039b0d4838ee65add608e807cba59cded2012101ea39ecd9beb
|
|
| MD5 |
c9cd04e52dd55c64f587756de64906a6
|
|
| BLAKE2b-256 |
8f94dadc084a3846ee6d91339ebf6a113fe22524f0b0bfbbd273cb2f410200db
|
Provenance
The following attestation bundles were made for interlaced-1.0.1-py3-none-any.whl:
Publisher:
publish.yml on interlace-sh/interlace
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
interlaced-1.0.1-py3-none-any.whl -
Subject digest:
a3703a3133f46039b0d4838ee65add608e807cba59cded2012101ea39ecd9beb - Sigstore transparency entry: 2335601215
- Sigstore integration time:
-
Permalink:
interlace-sh/interlace@1348055fde2748fba65b24af55c95d600495980d -
Branch / Tag:
refs/tags/v1.0.1 - Owner: https://github.com/interlace-sh
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@1348055fde2748fba65b24af55c95d600495980d -
Trigger Event:
push
-
Statement type: