Skip to main content

Data, connected. Content-addressable storage with cross-table ACID transactions.

Project description

Rhizo

From the rhizome — a root system with no center, where any point connects to any other.

In 1980, Deleuze and Guattari contrasted the rhizome with the tree: hierarchies vs networks, central authority vs emergent coherence. Traditional databases are trees — leaders, followers, coordination. Rhizo is rhizomatic: every node commits locally, consistency emerges mathematically.

The first database where coordination is optional.

Metric Rhizo Industry Standard Improvement
Transaction latency 0.022ms 100ms (consensus) 31,000x faster
Energy per transaction 2.2e-11 kWh 2.1e-6 kWh 97,943x less
Branch overhead 280 bytes 14.7 MB (Delta Lake) 52,500x smaller
OLAP queries 0.9ms 23ms (DuckDB) 26x faster

CI License: MIT Rust Python


The Problem

Distributed databases require consensus. Consensus adds latency, complexity, and energy cost — even when the operation doesn't need it.

Lakehouses (Delta Lake, Iceberg, Hudi) improve storage but can't solve:

Limitation What It Costs You
Single-table transactions No atomic updates across related data
No deduplication Paying for storage you don't need
No real branching Can't experiment without copying everything
Consensus overhead 100ms+ latency on every write

Rhizo can.


Benchmarks

OLAP Performance vs Industry (100K rows)

With the new DataFusion-powered OLAP engine, Rhizo delivers industry-leading query performance:

Metric Rhizo OLAP DuckDB Delta Lake Parquet Winner
Read 0.9ms 23.8ms 23.9ms 6.3ms Rhizo (26x)
Filtered (5%) 1.2ms 1.8ms 19.9ms 7.0ms Rhizo
Projection 0.7ms 1.4ms 14.5ms 3.5ms Rhizo (2x)
Complex Query 2.9ms 6.6ms 30.5ms 18.5ms Rhizo (2.3x)
Storage 3.67MB 6.26MB 63.10MB 3.73MB Rhizo (17x vs Delta)

Rhizo wins 4/6 performance categories with built-in lakehouse features no competitor matches.

JOIN Performance (10K users x 100K orders)

Operation Rhizo OLAP DuckDB Delta Lake
Simple JOIN 2.9ms 7.5ms 31.5ms
JOIN + Filter 3.0ms 5.9ms 33.4ms
JOIN + Aggregate 4.2ms 5.6ms 34.0ms

Rhizo wins all JOIN categories.

Scale Performance (1M rows)

Metric Rhizo OLAP DuckDB Speedup
Read 5.1ms 257.2ms 50x faster
Filter 1.9ms 14.2ms 7.5x faster
Write 415ms 625ms 1.5x faster

Unique Features (No Competitor Has All)

Feature Rhizo Delta Lake DuckDB Iceberg
OLAP Query Speed Yes No Yes No
Time Travel SQL Yes (VERSION 5) API only No API only
Branch Queries Yes (@branch) No No No
Changelog SQL Yes (__changelog) No No No
Cross-table ACID Yes No No No
Content Dedup Yes No No No
Merkle Integrity Yes No No No
Arrow Chunk Cache Yes (15x speedup) No No No
Algebraic Merge Yes (4M+ ops/sec) No No No

Core Operations

Operation Performance Notes
OLAP read (cached) 0.9ms 26x faster than DuckDB
Arrow cache read 0.24ms 15x faster than uncached
Write throughput 211 MB/s Native Rust Parquet encoding
Branch creation <10 ms Zero-copy, 280 bytes overhead
Time travel query 0.5ms O(1) version lookup
Cache hit rate 97.2% LRU eviction, no invalidation

Incremental Deduplication (Merkle Tree Storage)

Change Percentage Chunk Reuse Storage Savings
1% change 98.8% reuse ~49% vs naive
5% change 95.0% reuse ~47% vs naive
10% change 90.0% reuse ~45% vs naive

O(change) storage instead of O(n) per version.

Feature Comparison

Feature Rhizo Delta Lake Iceberg Hudi
Cross-table ACID Yes No No No
Zero-copy branching Yes No No* No
Global deduplication Yes No No No
Merkle tree dedup Yes No No No
Corruption detection Built-in External External External
Time travel Yes Yes Yes Yes
SQL Query Engine Yes Yes Yes Yes
Cloud Storage Planned Yes Yes Yes

*Iceberg branching requires Nessie catalog


Quick Start

Installation

# Coming soon: pip install rhizo

# For now, install from source:
git clone https://github.com/rhizodata/rhizo.git
cd rhizo
pip install -e .

Basic Usage

import rhizo
from rhizo import QueryEngine
import pandas as pd

# Initialize
store = rhizo.PyChunkStore("./data/chunks")
catalog = rhizo.PyCatalog("./data/catalog")
engine = QueryEngine(store, catalog)

# Write data
df = pd.DataFrame({
    "id": [1, 2, 3],
    "name": ["Alice", "Bob", "Charlie"],
    "score": [85.5, 92.0, 78.5]
})
engine.write_table("users", df)

# Query with DuckDB
result = engine.query("SELECT * FROM users WHERE score > 80")

Time Travel

# Query historical versions
result_v1 = engine.query(
    "SELECT AVG(score) FROM users",
    versions={"users": 1}
)

# Compare versions
diff = engine.diff_versions("users", 1, 2, key_columns=["id"])

Branching

# Create branch (instant, zero-copy)
engine.create_branch("experiment/new-scoring")
engine.checkout("experiment/new-scoring")

# Modify on branch (production unchanged)
engine.write_table("scores", updated_df)

# Query both branches
main_result = engine.query("SELECT * FROM scores", branch="main")
exp_result = engine.query("SELECT * FROM scores", branch="experiment/new-scoring")

# Merge when ready
engine.merge_branch("experiment/new-scoring", into="main")

Cross-Table Transactions

with engine.transaction() as tx:
    tx.write_table("customers", updated_customers)
    tx.write_table("orders", new_order)
    tx.write_table("audit_log", audit_entry)
    # All commit together, or all rollback

Architecture

Application Layer
    Python (rhizo) | Rust | CLI (planned)
    TableWriter | TableReader | QueryEngine (DuckDB)
                            |
                            v
                      FileCatalog
    Versioned table metadata | Time travel queries | Atomic commits
                            |
                            v
                       ChunkStore
    Content-addressed storage (BLAKE3) | Automatic deduplication
    Atomic writes | Integrity verification
                            |
                            v
                       File System
    2-level directory tree | JSON metadata | Parquet chunks

Current Status

Phase Description Status
Phase 1: Storage Content-addressable chunk store with BLAKE3 hashing Complete
Phase 2: Catalog Versioned file catalog with time travel Complete
Phase 3: Query DuckDB integration with SQL and time travel Complete
Phase 4: Branching Git-like branching with zero-copy semantics Complete
Phase 5: Transactions Cross-table ACID with recovery Complete
Phase 6: Changelog Unified batch/stream via subscriptions Complete
Phase A: Merkle Storage O(change) deduplication via Merkle trees Complete
Phase P: Performance Native Rust Parquet, parallel I/O Complete

All phases complete. 632 tests passing (370 Rust + 262 Python).

Performance Optimization Journey

Phase Optimization Result
P.1 Parallel chunk I/O (Rayon) 3-5x batch throughput
P.2 Memory-mapped reads Infrastructure for zero-copy
P.3 Parallel Parquet parsing 2.1x multi-chunk speedup
P.4 Native Rust Parquet encoder 2.3x write improvement
P.5 Arrow chunk cache 15x faster repeated reads

Phase P.5 leverages content-addressed storage for cache-friendly reads. Since chunk hashes never change, cached Arrow RecordBatches require no invalidation and are shared across tables, versions, and branches. Cache hits bypass both disk I/O and Parquet decoding.


Project Structure

rhizo/
├── rhizo_core/                 # Rust core library
│   └── src/
│       ├── chunk_store/      # Content-addressable storage
│       ├── catalog/          # Versioned catalog
│       ├── branch/           # Git-like branching
│       ├── transaction/      # Cross-table ACID
│       ├── changelog/        # Change tracking
│       └── merkle/           # Merkle tree deduplication
│
├── rhizo_python/               # PyO3 bindings
├── python/rhizo/  # Python query layer
├── tests/                    # Test suites
└── examples/                 # Interactive demos

Design Principles

  1. Immutability — All data is immutable once written. Updates create new versions.
  2. Content Addressing — Data identified by BLAKE3 hash enables automatic deduplication.
  3. Atomic Operations — Write-to-temp-rename pattern prevents corruption.
  4. Layered Architecture — ChunkStore, FileCatalog, and BranchManager are independent and composable.
  5. Time Travel by Default — Every version is preserved and queryable.
  6. Zero-Copy Branching — Branches are pointers to table versions, not data copies.

Documentation


License

MIT — See LICENSE for details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

rhizo_core-0.1.1-cp312-cp312-win_amd64.whl (4.0 MB view details)

Uploaded CPython 3.12Windows x86-64

rhizo_core-0.1.1-cp312-cp312-macosx_11_0_arm64.whl (4.1 MB view details)

Uploaded CPython 3.12macOS 11.0+ ARM64

rhizo_core-0.1.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (4.7 MB view details)

Uploaded CPython 3.8manylinux: glibc 2.17+ x86-64

File details

Details for the file rhizo_core-0.1.1-cp312-cp312-win_amd64.whl.

File metadata

  • Download URL: rhizo_core-0.1.1-cp312-cp312-win_amd64.whl
  • Upload date:
  • Size: 4.0 MB
  • Tags: CPython 3.12, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for rhizo_core-0.1.1-cp312-cp312-win_amd64.whl
Algorithm Hash digest
SHA256 40c89cdb572004a5c30ada4149e0b23ee1ec1270e243606c28715a322cc91103
MD5 f0c5f7676622b9cd5f75e755abfd9775
BLAKE2b-256 e88c7baf01a4cbab9527c96e3ad16f857ab6c5b2f7300385508b04e0d1f228ec

See more details on using hashes here.

Provenance

The following attestation bundles were made for rhizo_core-0.1.1-cp312-cp312-win_amd64.whl:

Publisher: publish.yml on rhizodata/rhizo

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file rhizo_core-0.1.1-cp312-cp312-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for rhizo_core-0.1.1-cp312-cp312-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 daf9ad1118bbaab7b4f20703a274de9d48e208c08170a1c2dc18c95739676b6b
MD5 27a89eb4f47769358828e1c07899e209
BLAKE2b-256 e9c64939323602adcc3d5dce87cf1ca9b3e94ce7cc522dc34ec62fabed56dc98

See more details on using hashes here.

Provenance

The following attestation bundles were made for rhizo_core-0.1.1-cp312-cp312-macosx_11_0_arm64.whl:

Publisher: publish.yml on rhizodata/rhizo

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file rhizo_core-0.1.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for rhizo_core-0.1.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 e992383893e8f040c9eb1c1f99a127255e0ccf2b78b156505c1e68ca1aa91bb1
MD5 2243012b2b8f1f72072641327e300537
BLAKE2b-256 b39ac8f83130a0ea112024de622ee08dfcb1d504e84b1790f5735690ee33d88b

See more details on using hashes here.

Provenance

The following attestation bundles were made for rhizo_core-0.1.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: publish.yml on rhizodata/rhizo

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page