Skip to main content

recordstore

tests PyPI license

A versioned key→record store over a content-addressed bytes store — a thin database kernel between an immutable blob store (such as Ethereum Swarm) and an application that wants to think in records and versions rather than blobs and references.

from recordstore import RecordStore, MemoryBytesStore

blobs = MemoryBytesStore()
store = RecordStore(blobs)
store.put("users/alice", {"name": "Alice", "role": "admin"})
store.put("users/bob", {"name": "Bob"})
root = store.commit()          # one reference identifies this entire version

store.get("users/alice")       # {'name': 'Alice', 'role': 'admin'}
list(store.keys("users/"))     # ['users/alice', 'users/bob']

snapshot = RecordStore.at(root, blobs)   # frozen view of that version

Concurrent writers converge without a lock server — if the shared pointer moved under a commit, it three-way merges and retries:

from recordstore import MemoryPointer

pointer = MemoryPointer(root)                  # a shared "latest version" name
a = RecordStore(blobs, pointer=pointer)        # two writers open the same version
b = RecordStore(blobs, pointer=pointer)
a.put("users/carol", {"name": "Carol"}); a.commit(reconcile=True)
b.put("users/dave",  {"name": "Dave"});  b.commit(reconcile=True)   # folds in a's change
# the pointer now names a version containing both carol and dave

Why

Content-addressed stores give you immutable put(bytes) → ref / get(ref) → bytes and nothing else: no keys, no typed records, no transactions, no snapshots. recordstore adds exactly that missing layer and nothing more:

  • Records instead of raw bytes — values are any JSON-compatible object, stored under string keys.
  • Atomic, versioned commits — mutations are staged in memory; commit() lands all of them as one new root reference. A reader either sees all of a commit or none of it.
  • Snapshot isolationRecordStore.at(root, blobs) pins one root and sees a frozen, self-consistent dataset for arbitrarily long reads, with no locking: the whole dataset-at-a-version is one reference.
  • Canonical roots — encodings are deterministic, so equal content produces an equal root reference, regardless of the insertion/deletion history that produced it. Versions are content-addressable, comparable with a string equality check, and cheap to diff.
  • Three-way mergeRecordStore.merge(base, ours, theirs) reconciles two divergent versions; canonicity makes unchanged subtrees merge for free and equal edits conflict-free, with conflicts raised or settled by a resolver(key, base, ours, theirs) you supply. The diff is O(divergence), not O(dataset).
  • Verifiable proofsstore.prove(key) produces a small, JSON-ready inclusion or absence proof against the committed root (absence is provable because the canonical encoding gives a key exactly one possible location); verify_proof(proof, root) checks it with no store access at all — hold the 32-byte root, verify any claim about the dataset.
  • Version comparisonstore.diff(other_root) answers "what changed between these two published versions?" directly: (key, mine, theirs) per differing key, the same O(divergence) structural walk, so equal roots read nothing at all.
  • Local-firstlocal_first_store(path, api_url) stops the choice between disk and Swarm: commits land on local disk instantly (offline is the normal mode), a background worker pushes them to Swarm and confirms arrival peer-to-peer, sync() is the certainty barrier, and local disk is a budgeted working set — unpushed data is pinned, only Swarm-confirmed blobs evict, evicted reads heal by verified re-fetch. Shape the working set with pin(name, prefix) / fetch(prefix); collapse local history with squash_history(); a key whose bytes are temporarily unreachable raises RecordUnavailable, never a false KeyError.
  • Multi-writer, no lock servercommit(reconcile=True) makes concurrent writers converge: if the pointer moved under you it three-way merges and retries instead of overwriting. Race-free in-process; best-effort across processes over a Swarm feed. A commutative resolver keeps 3+ writers order-independent.
  • Structural sharing — versions are stored as a persistent (copy-on-write) compacted radix trie; a commit writes only the blobs along the changed paths, and unchanged subtrees are shared between versions.
  • Concurrent I/Oitems() (bulk read), get_many/put_many, and a pooled keep-alive HTTP session let BeeBytesStore parallelise round trips instead of paying one per record — the difference between usable and not on a high-latency link.

Install

pip install recordstore

# with the Bee (Swarm) bytes backend's HTTP dependency:
pip install "recordstore[bee]"

# with the Swarm feed pointer (adds swarm-bee for SOC/secp256k1 signing):
pip install "recordstore[feeds]"

# with postage_batch_id="auto" and batch-health reporting (adds swarmfs):
pip install "recordstore[stamps]"

Python ≥ 3.11. The core imports only the standard library; both extra dependencies are imported lazily — requests only by BeeBytesStore ([bee]), swarm-bee only by SwarmFeedPointer ([feeds]).

The pieces

Layer What it does Implementations
BytesStore put(bytes) → ref, get(ref) → bytes MemoryBytesStore (in-memory, testing), DirBytesStore (durable local directory), FsspecBytesStore (S3/GCS/HTTP/… via fsspec), BeeBytesStore (Swarm Bee node over /bytes — the blob endpoint, not the raw /chunks/{address} primitive), CachedBytesStore (byte-budgeted LRU wrapper over any of them), swarmfs's LocalStore (local-first store directory)
trie (internal) canonical persistent radix trie mapping keys to value blobs
RecordStore staging, commit() / commit(reconcile=True), snapshots, sorted keys()/items(), three-way merge(), structural diff()
Pointer mutable name for the latest root MemoryPointer, FilePointer (atomic local file), SwarmFeedPointer (owner-signed Swarm feed, over swarm-bee)

| swarm_store(topic, ...) | assembles the two Swarm pieces into a store | the one place Swarm is chosen: BeeBytesStore blobs and a SwarmFeedPointer head | | local_first_store(path, api_url) | disk now, Swarm in the background | swarmfs LocalStore + journal (the reflog) + background push/confirm; sync(), sync_status(), pin/fetch, publish(pointer), squash_history() ([local] extra, swarmfs ≥ 0.7) |

Nothing above RecordStore ever sees a stored blob or a trie node — and nothing below it needs to know Swarm exists unless you asked for it:

from recordstore import swarm_store

store = swarm_store("my-notes", signer=key)   # publish: blobs + feed on Swarm
store = swarm_store("my-notes", owner=addr)   # follow someone else's feed

Documentation

  • User guide — the tutorial: concepts, the canonicity contract, running against a real Bee node, versioning patterns, error handling, and current limitations.
  • Reference — compact and definition-first: every export, signature, error, and extra in tables, pinned against the code by tests/test_reference.py. The right document to hand to an AI agent.

Testing

python3 -m pytest tests/                                 # unit + fuzz + boundary tests

BEE_API=http://<node>:1633 BEE_BATCH=<batchID> \
    python3 -m pytest tests/test_recordstore_bee.py -v   # bytes backend, live node

pip install "recordstore[feeds]"                         # needs swarm-bee
BEE_API=http://<node>:1633 BEE_BATCH=<batchID> \
    python3 -m pytest tests/test_recordstore_feed.py -v  # feed pointer, live node

The fuzz suite runs randomized put/delete histories against a plain-dict oracle and asserts the canonical-root property throughout. The Bee integration tests skip automatically unless BEE_API is set (the feed test also needs swarm-bee installed); against a real (non-dev) node always provide BEE_BATCH with a purchased postage batch id.

Background

This is a Python re-implementation of an old idea — content-addressed, canonical-root, versioned key-value storage — best known from Ethereum's Merkle Patricia Trie and from Noms/Dolt's "prolly trees." The value here is fit, not novelty: a much simpler canonical encoding than MPT (avoiding the exact bug class that once caused a chain split), a far smaller scope than Dolt/Irmin (no query language; a single three-way merge primitive, not a merge engine), and — as far as we could find — the first implementation of this pattern for Python with a Swarm/Bee backend. See the user guide's background section for the full comparison.

Status

Extracted from petfold/ontodag (July 2026) with history preserved; validated against a live Bee 2.8.1 light node on Gnosis mainnet (roundtrips, canonical roots on real BMT references, network retrievability). SwarmFeedPointer (owner-signed Swarm feed, over swarm-bee) landed in v0.4.0; three-way merge in v0.8.0; auto-reconciling commit(reconcile=True) in v0.9.0; a best-effort feed compare_and_set (cross-process reconcile) in v0.10.0. Known gaps — a Swarm feed has no atomic index claim, so exactly-simultaneous same-index writes can still race (in-process reconcile is race-free); one blob per record — are detailed in the user guide.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

recordstore-0.18.2.tar.gz (64.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

recordstore-0.18.2-py3-none-any.whl (36.4 kB view details)

Uploaded Python 3

File details

Details for the file recordstore-0.18.2.tar.gz.

File metadata

  • Download URL: recordstore-0.18.2.tar.gz
  • Upload date:
  • Size: 64.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for recordstore-0.18.2.tar.gz
Algorithm Hash digest
SHA256 e3463a8e349bf2675c090a5f69555622565f15abb997739ed946e5a49b74c5ff
MD5 162b553900aca0d4d8d295b3f4180873
BLAKE2b-256 e9de7c49021670338e9826c3ecc46419194940e662667a47b326eb5a214fd5b9

See more details on using hashes here.

Provenance

The following attestation bundles were made for recordstore-0.18.2.tar.gz:

Publisher: publish.yml on petfold/recordstore

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file recordstore-0.18.2-py3-none-any.whl.

File metadata

  • Download URL: recordstore-0.18.2-py3-none-any.whl
  • Upload date:
  • Size: 36.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for recordstore-0.18.2-py3-none-any.whl
Algorithm Hash digest
SHA256 97b7d4038a491d4ae187bbbc72e35d95a9c6079b89ded414c75abd91eee73457
MD5 0d3f979d0ae230910abc5a17ca28b29c
BLAKE2b-256 bd865dc6db7b56738c98e5327064e5cbceb5e0ae99c716b6c3e46b31b534fd96

See more details on using hashes here.

Provenance

The following attestation bundles were made for recordstore-0.18.2-py3-none-any.whl:

Publisher: publish.yml on petfold/recordstore

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.20.2

2 files

0.20.1

2 files

0.20.0

2 files

0.19.0

2 files

This release

0.18.2 This release

2 files

0.18.1

2 files

0.18.0

2 files

0.17.0

2 files

0.16.0

2 files

0.15.0

2 files

0.14.0

2 files

0.13.2

2 files

0.13.1

2 files

0.13.0

2 files

0.12.1

2 files

0.12.0

2 files

0.11.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page