Skip to main content

Reble

tests PyPI License: Apache-2.0 CLA assistant

Your models are just SQL files. Your bucket is the warehouse. Any machine with the CLI and bucket credentials is a full query engine.

Branch your warehouse like you branch your code — no warehouse server to run, scale, or pay for.

-- models/demo/orders_clean.sql      ← filename = table name. That's the whole format.
SELECT id, amount FROM raw.orders WHERE amount > 0

The reble loop in 40 seconds: a one-line metric fix on a 1.2M-row warehouse — git switch -c, then reble run creates the data branch with inferred scope, status shows git provenance, row-level diff, query on the branch, promote

Prefer stills? The full CLI output design shows every command's output on one page (HTML source).

No MODEL(...) headers. No Jinja {{ ref('...') }}. No per-model YAML. Dependencies, column lineage, and change detection are read from the SQL you already wrote — powered by SQLGlot, the parser underneath the ecosystem's lineage tooling. A model with no config is a FULL rebuild; the few models that need more (incremental, in v0.2) get a couple of lines in the one reble.yml the project already has — never boilerplate per model.

Reble is a single CLI that gives a data team — three people or three hundred — a complete analytics platform: what's sized is the run, not the org. Each run is one fast node over just the tables it touches — DuckDB + Apache Iceberg + SQLGlot-powered transforms, pre-wired — with subset branching of the warehouse and branch-per-PR CI for data pipelines. Data lives in S3/GCS (or on local disk while you're trying it out); the query engine is DuckDB embedded in the CLI, so your laptop and your CI runners are the compute — for exactly as long as a run takes, and never a second of billing after.

⚠️ Status: pre-alpha, but real. The full loop works today — init → run → branch → run → diff → promote, both git orders, with 54 passing tests — pip install reble and go. The engine is the spike-validated SQLGlot-direct core: models are plain SQL files, exactly as described below. Developed and CI-tested on macOS and Linux; Windows is untested (reports welcome). Feedback is the most valuable contribution — open a Discussion.

Why Reble exists — the full story: the four gaps in data engineering workflows, why existing tools don't close them, and why now. Coming from dbt? — dbt's own jaffle shop ported to Reble, with the translation table and FAQ (short version: branches replace environments — there's nothing to configure).

The idea

Testing a data pipeline change today means cloning or rebuilding an entire dev warehouse — even when your change touches three tables. Reble branches just the tables you're changing:

pip install reble
reble init my-warehouse
git switch -c fix-orders                # branch your code like you always do…
# edit your models…
reble run                               # …and the data branch appears: named after
                                        # your git branch, scope + pins inferred
                                        # (edited models, downstream cascade,
                                        # upstream inputs frozen at the epoch);
                                        # writes go to zero-copy Iceberg branch refs
reble diff                              # schema + row-level diff vs your branch base
reble promote                           # atomic fast-forward to main, clean up

One branch gesture, two artifacts. Your git branch tracks the code change; the data branch that follows it holds the change's blast radius. No environments to configure, no branch names to invent twice. Not in a git repo (or set git_sync: false in reble.yml)? reble branch create <name> does the same thing explicitly — reble reads git state but never runs a git command for you.

Both git orders work: edit-first (scope inferred from the diff) or branch-first (empty scope + frozen epoch; the scope grows automatically at first run, and reads resolve as of the moment you branched). And when you come back to a branch after two weeks, reble status is the "where was I?" answer: what you edited but haven't run, which pinned inputs moved on main under you, what commit the data reflects, when the branch expires. Every read command takes --json for scripts and bots.

  • Zero-copy branches — branched tables use native Iceberg refs (copy-on-write); a branch of a 10GB table costs ~nothing until you write.
  • Pinned inputs — unbranched tables are read at their snapshot from branch-creation time, so your test inputs don't drift while prod keeps ingesting.
  • Row-level diffs — answer the question every reviewer actually has: what rows does this change?
  • Column-level lineage & change detection — inferred from your SQL via SQLGlot: only changed models run (cosmetic edits don't count — hashing is on the canonical AST), and downstream impact is shown before you apply.
  • No merge, ever — branches are ephemeral: create, test, promote (fast-forward or re-run) or discard. Reble refuses to build last-write-wins data merges.

Four days you've had

The same loop, in scenarios every data engineer has lived through. All CLI output below is the real tool's output format. The example warehouse:

raw.orders                 ← ingested hourly by Airbyte
raw.customers              ← ingested nightly
core.stg_orders            ← staging model
core.stg_customers         ← staging model
core.fct_revenue_daily     ← the table finance actually looks at
core.mart_exec_dashboard   ← reads fct_revenue_daily

1. "Finance says revenue is wrong" — changing a metric definition

Cancelled orders are being counted as revenue. The fix is one line in stg_orders — but fct_revenue_daily and mart_exec_dashboard are downstream, and finance will ask exactly one question: how much does this change the numbers?

gitGraph
    commit id: "prod (hourly ingest continues)"
    branch fix-cancelled-revenue
    commit id: "exclude cancelled orders"
    commit id: "run + diff: -3.2% revenue"
    checkout main
    merge fix-cancelled-revenue id: "promote"
$ vim models/core/stg_orders.sql   # ... WHERE status != 'cancelled'

$ git switch -c fix-cancelled-revenue
$ reble run
⎇ main → fix-cancelled-revenue · data branch created from your git branch
  scope  core.fct_revenue_daily, core.mart_exec_dashboard, core.stg_orders  inferred from your edits · grows as you work
  pins   raw.orders  upstream inputs, frozen now
⎇ fix-cancelled-revenue
  changed    core.stg_orders, core.fct_revenue_daily, core.mart_exec_dashboard
  published  core.stg_orders, core.fct_revenue_daily, core.mart_exec_dashboard → branch ref
✓ 3 models in 0.2s

Notice what you didn't do: create a branch in reble, or enumerate the downstream cascade. The data branch followed your git branch, and the scope came off the model graph — your one-line edit touches three tables, and the raw input feeding them is pinned so hourly ingestion can't shift your numbers mid-analysis. (Prefer it explicit? reble branch create still works, and it's the flow for projects without a git repo.)

$ reble diff
⎇ fix-cancelled-revenue vs base

  core.stg_orders  1,204,331 → 1,168,210 rows · key order_id
    −36,121 removed

  core.fct_revenue_daily  730 → 730 rows · key date_id
    ~214 changed

There's finance's answer, before anything touched prod: 36,121 cancelled orders excluded, revenue restated on 214 of 730 days. Screenshot the diff, get the sign-off, then:

$ reble promote
⎇ fix-cancelled-revenue → main
  ✓ core.fct_revenue_daily
  ✓ core.mart_exec_dashboard
  ✓ core.stg_orders
✓ promoted · branch deleted · on main
  your data is on main; git is still on fix-cancelled-revenue — merge or switch when ready

Without branches: you'd have run this in a shared dev schema (numbers drifting under you with every hourly ingest), eyeballed two spreadsheet exports, and pushed to prod hoping.

Three more, told in full in docs/scenarios.md

  • Building a brand-new mart — branch-first on a clean tree: frozen inputs while you iterate for days, and a profile instead of a diff (those 5 nulls get caught here, not in the exec's dashboard).
  • Two engineers, two branches, zero coordination — disjoint scopes work in parallel; overlap is warned at creation, and promote refuses to silently merge data.
  • The save — a "simplified" join fans out into a 64% row explosion, caught on a laptop on frozen inputs, deleted for free.

The pattern

All four are the same loop:

(edit ↔ branch, either order)  →  run  →  diff or profile  →  promote or discard

Branches are metadata only — zero-copy Iceberg refs plus a frozen epoch. Creating one is free; deleting one is guilt-free. Inputs never drift, prod is never at risk, and the diff answers the question reviewers actually ask.


Measured, not promised

The design is validated by reproducible spikes in spikes/, including a full-scale performance run — 140M rows / 10.22GB on an Apple M4 Pro laptop (pyiceberg 0.11.1, DuckDB 1.5.5):

Operation at 10GB scale Time
Create a branch of the 140M-row table < 10ms (zero-copy, size-independent)
Pinned full-table scan → Arrow 4.0s
Projected scan (2 of 6 columns) 0.47s
Full diff — both refs scanned, added + changed rows 5.9s
Branch append (5M rows) 1.3s
Bulk load throughput ~3.5M rows/s

Peak RAM 12.3GB, 3.5GB on disk (Parquet ≈ 2.9× compression). Details and the scripts to reproduce: spike 1 — branch lifecycle · spike 2 — performance · spike 4 — the SQLGlot-direct core.

The killer workflow: branch-per-PR — live

See it running on dbt's jaffle shop: a one-line fix to customer_lifetime_value, and the review comment answers the only question that matters — "~6 customers changed, nothing added, nothing removed" — computed on a zero-copy branch with epoch-pinned inputs.

On every pull request, a ~70-line hermetic workflow (no services, no state, copy it into any repo):

  1. rebuilds the main baseline from seeds
  2. creates a branch scoped to your changed models (inferred — no config)
  3. runs only the changed models against epoch-pinned inputs
  4. posts the row-level diff as a PR comment

The jaffle-shop-classic port is also the dbt side-by-side: the original Jinja project and the plain-SQL Reble port live in the same repo.

Two modes, one tool

On-ramp (zero services): everything on your laptop — DuckDB embedded, Iceberg on local disk, SQLite catalog, transforms in-process. No Docker, no daemons. This is how you try Reble in about ten minutes, run its test suite, or run a solo project.

Production (team mode): the same project pointed at S3/GCS + a shared Postgres or REST catalog — because that's where real warehouses live. Your laptop (or a CI runner) stays the query engine.

⚠️ Team mode status — be clear-eyed here. What's validated today is the single-writer shape: one person (or one CI job) running the loop over S3 (spike 06). Multiple people writing through a shared catalog at once is not supported yet. The multi-user design (local branches, remote main — spike 07) is validated but not wired. If you're a team today: give exactly one identity write access and treat everyone else as readers.

The team flow is the dbt flow. Edit SQL on a git branch, reble run locally (the data branch appears, inputs frozen), open a PR (the bot posts the row-level diff), merge — and a prod job runs reble run on main, rebuilding exactly the changed models against current inputs. Nobody types promote on a team; it's the solo shortcut for when you are your own prod job. Where this is headed: branches local, main remote, exactly like git — developers hold read-only bucket credentials, branch as zero-copy overlays of pinned prod snapshots (measured: 6ms, nothing copied), and only the merge gate writes main.

Query flow: the reble CLI on your laptop or a CI runner runs SQLGlot, pyiceberg and embedded DuckDB; it GETs Iceberg metadata and only the needed column chunks from your S3 bucket, computes locally in RAM, and PUTs results back as a branch-ref commit. No database server anywhere.

"Local compute" does not mean copying the warehouse. Each reble run streams just what that run needs — the tables its changed models touch, the columns their SQL references, at the pinned snapshots — through memory and writes results back to the bucket. Nothing is replicated or stored locally; it's the same I/O a remote warehouse does internally, with the CPU (and the bill) relocated. Branches and pins are catalog metadata, so branching a 500GB table in S3 is the same instant, zero-copy operation as locally. The full mechanics, measured S3 numbers, and what's still unmeasured: architecture.md.

# reble.yml — the entire difference between modes
warehouse: s3://my-bucket/warehouse
catalog:
  type: sql
  uri: postgresql://...

What Reble is not

  • Not a query engine, storage engine, or table format — it composes DuckDB, Iceberg, and SQLGlot and adds the branching layer and glue.
  • Not a dbt/SQLMesh replacement you must migrate to all at once — importers ({{ ref('...') }} and MODEL(...) translation) are on the roadmap.
  • Reble doesn't compete with Snowflake the product. It competes for the workloads that never needed it — and for most pipelines, the honest question isn't "which warehouse?" but "does this workload need a warehouse vendor at all?" Compute is one fast node per run over data in your bucket, not a cluster.
  • Not a full-catalog branching system (see Nessie/lakeFS for that) — Reble branches subsets over standard Iceberg catalogs, no migration required.

Where this is going

Shipped in v0.1.0: reble follows git — implicit data branches on reble run, git provenance in reble status, --json on every read command. Next, in rough order (opinions welcome in Discussions):

  • reble serve — your branch in your own tools — a local Iceberg REST catalog proxy that answers with branch-resolved snapshots, so DBeaver, DataGrip, DuckDB, Spark, or a notebook connect to localhost and see the warehouse exactly as your branch sees it. Read-only, no plugin required.
  • Agent-native operation — the branching machinery is designed to disappear under the hood: an MCP server exposing run/diff/status/promote as structured tools, so an AI agent can take "exclude cancelled orders from revenue and show me the impact," edit the model, get a zero-copy sandbox branch automatically, and hand back the row-level diff — with prod physically out of reach. The --json output on every read command is the substrate for this.
  • reble branch refresh — re-pin a long-lived branch's inputs to now and rerun, so a two-week-old branch can catch up to today's data before promote.
  • SQL-defined data testsunique, not_null, accepted-values as plain SQL assertions that run with the models and gate promote; the diff answers "what changed", tests should answer "is it still correct".
  • Team mode: local branches, remote main — developers get read-only bucket credentials and branch locally as zero-copy overlays of pinned prod snapshots (spike-validated: 6ms, nothing copied, shared warehouse untouched by local writes); only the merge gate writes main. AWS S3 Tables (managed Iceberg with a REST catalog) is a candidate shared catalog that would mean nothing to host at all.
  • Merge-driven promote — later, as a pure optimization: skip recomputing an expensive model in the prod run when its inputs haven't moved.
  • Incremental models (v0.2) — a couple of lines in reble.yml, never per-model boilerplate.
  • dbt/SQLMesh importers{{ ref('...') }} and MODEL(...) translation for gradual migration.

Design docs

Contributing

Feedback beats code right now — try the loop on your own models and open a Discussion or an issue. See CONTRIBUTING.md.

License

Apache 2.0 — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

reble-0.1.1.tar.gz (49.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

reble-0.1.1-py3-none-any.whl (37.9 kB view details)

Uploaded Python 3

File details

Details for the file reble-0.1.1.tar.gz.

File metadata

  • Download URL: reble-0.1.1.tar.gz
  • Upload date:
  • Size: 49.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for reble-0.1.1.tar.gz
Algorithm Hash digest
SHA256 0736b38bce1bdaa9204fd7cc862f7037f161479b30ada39903510136492d69ab
MD5 ef3864ad2f234e22f31501032216a801
BLAKE2b-256 65dd1cb15f12ea0416b46ae76a8a3e02cf5801d4127b34051d72f8e2722ee04f

See more details on using hashes here.

Provenance

The following attestation bundles were made for reble-0.1.1.tar.gz:

Publisher: release.yml on satya1395/reble

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file reble-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: reble-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 37.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for reble-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 b881a42a5b1c0fe25b4a4cfc98ab045f4fbdbfd4faa84e3881c1dc54766d4ba6
MD5 ccca617bec17e8023b1bcb9fac0708d7
BLAKE2b-256 948d2e3ac03cbc7ab54fbd28063a9f1fad8dbc9ebd31aefc2bed1a3f8bd7793c

See more details on using hashes here.

Provenance

The following attestation bundles were made for reble-0.1.1-py3-none-any.whl:

Publisher: release.yml on satya1395/reble

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.6.1

2 files

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.2

2 files

0.4.1

2 files

0.4.0

2 files

0.3.2

2 files

0.3.1

2 files

0.3.0

2 files

0.2.1

2 files

0.2.0

2 files

This release

0.1.1 This release

2 files

0.1.0

2 files

0.0.9

2 files

0.0.8

2 files

0.0.7

2 files

0.0.6

2 files

0.0.5

2 files

0.0.4

2 files

0.0.3

2 files

0.0.2

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page