Skip to main content

dr-store

CI PyPI

Definitions · Terms · Contracts · Changelog · dr-serialize

dr-store provides domain-neutral storage primitives for immutable records and document artifacts:

  • Content addressing identifies complete records by their declared schemas and SHA-256 hashes of their Canonical JSON Text under dr-serialize's frozen profile.
  • Object Store provides immutable puts, verified reads, and atomic bindings from opaque caller-owned keys to object references.
  • Storage backends supply the Object Store's atomic, append-only point and batch operations. MemoryBackend is process-local; SqliteBackend persists committed data for cross-process use.
  • Record Cache memoizes records under opaque caller-owned keys. Reads return typed hits; absent, missing, or unverifiable stored values are misses, while invalid requested schemas and operational backend faults raise. Entries are never rebound, so callers invalidate by selecting a new key; derive_cache_key provides a canonical scheme using a versioned namespace and payload. Single and bulk methods share these per-key semantics; SqliteRecordCache(path) is the managed persistent lifecycle.
  • Canonical JSON document files publish and read one standalone, bounded canonical document in an existing directory through descriptor-pinned filesystem operations.
  • Document Directory delegates one bounded canonical Manifest to that file capability beside streamed binary Sidecars.

Installation

dr-store requires Python 3.12 or newer.

python -m pip install dr-store

Usage

from dr_store import MemoryBackend, ObjectStore

store = ObjectStore(MemoryBackend())
reference, _ = store.put("example.note.v1", {"title": "hello"})
store.bind("notes/latest", reference)

assert store.resolve("notes/latest") == reference
assert store.get(reference) == {"title": "hello"}

SqliteRecordCache(path) is the paved persistent Record Cache. It initializes its database before returning and closes its current-process resources on normal or exceptional context exit. When cleanup succeeds, an exception from the context body is not suppressed; cleanup failure raises SqliteRecordCacheCloseError:

from dr_store import CacheEntry, CacheHit, SqliteRecordCache, derive_cache_key

key = derive_cache_key("example.summary.v1", {"document": "note-42"})
with SqliteRecordCache("records.sqlite3") as cache:
    winners = cache.put_many(
        {
            key: CacheEntry(
                schema="example.summary.v1",
                record={"summary": "hello"},
            )
        }
    )
    assert winners[key].schema == "example.summary.v1"
    assert cache.get_many(
        [key, "missing"], schema="example.summary.v1"
    ) == {
        key: CacheHit(record={"summary": "hello"}),
        "missing": None,
    }

CanonicalJsonFile publishes one standalone document in an existing directory. The caller must declare the maximum accepted canonical byte length:

from pathlib import Path

from dr_store import CanonicalJsonFile

artifact_directory = Path("artifacts")
artifact_directory.mkdir(exist_ok=True)
metadata = CanonicalJsonFile(
    artifact_directory,
    "metadata.json",
    max_bytes=1 << 20,
)
metadata.publish({"state": "complete"})
assert metadata.read() == {"state": "complete"}

Use the lower-level SqliteBackend(path) when assembling an ObjectStore directly whose objects and bindings must persist across processes. The rendered definitions, authoritative terms, and binding contracts describe the vocabulary, public-export mappings, and behavioral boundaries.

Content addressing

Content addressing validates each schema-qualified reference and derives its content hash through dr-serialize's canonical JSON profile. Its stable public shape is:

@dataclass(frozen=True, slots=True)
class ObjectReference:
    schema: str
    content_hash: str

    @classmethod
    def for_record(cls, schema: str, record: Jsonable) -> ObjectReference: ...
    def verify_record(self, record: Jsonable) -> None: ...

def compute_content_hash(record: Jsonable) -> str: ...
def is_content_hash(value: str) -> bool: ...

Object Store

The Object Store owns immutable record operations and opaque key bindings. Its statuses and store surface keep storage outcomes distinct from stored records:

class PutStatus(Enum):
    STORED = "stored"
    IDEMPOTENT = "idempotent"

class BindStatus(Enum):
    BOUND = "bound"
    IDEMPOTENT = "idempotent"

class ObjectStore:
    def __init__(self, backend: Backend) -> None: ...
    def put(
        self, schema: str, record: Jsonable
    ) -> tuple[ObjectReference, PutStatus]: ...
    def get(self, reference: ObjectReference) -> Jsonable: ...
    def bind(
        self, key: str, reference: ObjectReference
    ) -> BindStatus: ...
    def resolve(self, key: str) -> ObjectReference | None: ...

Storage backends

Storage backends implement one atomic protocol beneath the Object Store. Outcome objects carry the existing row when an append-only operation does not insert. Batch value objects carry prepared writes and joined binding/object read results:

@dataclass(frozen=True, slots=True)
class PutOutcome:
    inserted: bool
    stored_schema: str
    stored_canonical: str

@dataclass(frozen=True, slots=True)
class BindOutcome:
    bound: bool
    existing_schema: str
    existing_content_hash: str

@dataclass(frozen=True, slots=True)
class BoundObjectWrite:
    key: str
    schema: str
    content_hash: str
    canonical: str

@dataclass(frozen=True, slots=True)
class BoundObjectRow:
    binding_schema: str
    binding_content_hash: str
    object_schema: str | None
    canonical: str | None
class Backend(Protocol):
    def put_object(
        self, *, schema: str, content_hash: str, canonical: str
    ) -> PutOutcome: ...
    def get_object(
        self, *, schema: str, content_hash: str
    ) -> tuple[str, str] | None: ...
    def bind(
        self, *, key: str, schema: str, content_hash: str
    ) -> BindOutcome: ...
    def get_binding(self, *, key: str) -> tuple[str, str] | None: ...
    def get_bound_objects(
        self, *, keys: tuple[str, ...]
    ) -> dict[str, BoundObjectRow]: ...
    def put_bound_objects(
        self, *, entries: tuple[BoundObjectWrite, ...]
    ) -> dict[str, BindOutcome]: ...

class MemoryBackend: ...
class SqliteBackend:
    def __init__(self, path: str | Path) -> None: ...

Batch reads address only the supplied exact keys. SQLite performs chunked joined binding/object queries and does not promise one snapshot across the chunks. Every non-empty SQLite write batch uses one immediate transaction; a failure rolls back that transaction. Committed rows persist across reopen, but the backend does not promise power-loss durability.

Record Cache

The Record Cache is a memoization facade over an existing ObjectStore. It accepts opaque caller-owned keys, with derive_cache_key as the canonical helper for content-derived memoization. A typed hit keeps every strict JSON record, including null, distinct from a miss. SqliteRecordCache supplies the managed persistent form:

def derive_cache_key(namespace: str, payload: Jsonable) -> str: ...

@dataclass(frozen=True, slots=True)
class CacheHit:
    record: Jsonable

@dataclass(frozen=True, slots=True)
class CacheEntry:
    schema: str
    record: Jsonable

class RecordCache:
    def __init__(self, store: ObjectStore) -> None: ...
    def get(self, key: str, *, schema: str) -> CacheHit | None: ...
    def get_many(
        self, keys: Iterable[str], *, schema: str
    ) -> dict[str, CacheHit | None]: ...
    def put(
        self, key: str, schema: str, record: Jsonable
    ) -> ObjectReference: ...
    def put_many(
        self, entries: Mapping[str, CacheEntry]
    ) -> dict[str, ObjectReference]: ...

class SqliteRecordCache(RecordCache):
    def __init__(self, path: str | Path) -> None: ...
    def close(self) -> None: ...
    def __enter__(self) -> SqliteRecordCache: ...
    def __exit__(self, ...) -> bool: ...

get_many deduplicates requested keys and returns exactly those distinct keys, with each value independently classified as a hit or miss under the requested schema. It parses and verifies each returned bound object once; it does not promise that a multi-chunk backend read observes one snapshot. put_many validates, canonicalizes, and hashes each proposed entry once before invoking one backend write batch, then returns the first binding winner for every input key. Single-key get and put use the same paths and semantics.

The cache intentionally provides no scheduler, dirty tracking, key enumeration, prefix query, delete, expiry, eviction, or size cap. Callers own those policies and choose new keys for invalidation.

Construction captures a non-transient absolute filesystem path and establishes the SQLite schema before returning; empty and :memory: paths are rejected. Initialize a new database path with one constructor before starting concurrent constructors. Closing rejects new cache operations, waits for every admitted get, get_many, put, or put_many to finish, and then closes all operational connections tracked by that cache instance in the current process. A successful close is idempotent for repeated and concurrent callers. SqliteRecordCacheClosedError reports operations requested after closing begins, before their inputs are validated; SqliteRecordCacheCloseError reports a terminal cleanup failure to close callers, including a context exit. An interruption before connection cleanup begins restores the open lifecycle and wakes another closer; a process-level cleanup interruption terminalizes it before propagating to the elected caller. Committed records remain available after close and reopen. Closing one cache does not close a separate instance or coordinate another process, even when both use the same database path. These persistence semantics do not promise power-loss durability.

Canonical JSON document files

A canonical JSON document file owns publication and verified reads for one caller-named document in an existing directory. The byte bound is required; the nesting-depth bound defaults to the dr-serialize canonical profile maximum and applies to both publication and read:

class CanonicalJsonFile:
    def __init__(
        self,
        directory: str | Path,
        name: str,
        *,
        max_bytes: int,
        max_depth: int = CANONICAL_JSON_MAX_CONTAINER_DEPTH,
    ) -> None: ...

    @property
    def path(self) -> Path: ...
    def publish(self, document: Jsonable) -> None: ...
    def read(self) -> Jsonable: ...
@verify(UNIQUE)
class PublicationStage(StrEnum):
    ENCODE = "encode"
    CREATE_TEMP = "create_temp"
    WRITE_TEMP = "write_temp"
    FLUSH_TEMP = "flush_temp"
    REPLACE_TARGET = "replace_target"
    FLUSH_DIRECTORY = "flush_directory"

@verify(UNIQUE)
class ReplacementState(StrEnum):
    NOT_REPLACED = "not_replaced"
    REPLACED = "replaced"
    UNKNOWN = "unknown"

DocumentPublishError reports a PublicationStage and ReplacementState. NOT_REPLACED means replacement was not invoked and any prior target remains authoritative. REPLACED means replacement returned before later finalization failed. UNKNOWN means the replacement operation itself failed and cannot prove whether the target changed, so callers must inspect or coordinate before treating either value as authoritative. DocumentReadError reports the requested path. Both errors derive from DocumentFileError and preserve the originating failure as their cause.

Document Directory

A Document Directory groups one canonical JSON Manifest with streamed binary Sidecars. The directory owns allocation and publication while SidecarWriter owns bounded retention:

class DocumentDirectory:
    def __init__(
        self,
        path: Path,
        manifest_name: str,
        *,
        manifest_max_bytes: int,
        manifest_max_depth: int = CANONICAL_JSON_MAX_CONTAINER_DEPTH,
    ) -> None: ...

    @classmethod
    def allocate(
        cls,
        root: str | Path,
        *,
        prefix: str,
        manifest_name: str,
        manifest_max_bytes: int,
        manifest_max_depth: int = CANONICAL_JSON_MAX_CONTAINER_DEPTH,
    ) -> DocumentDirectory: ...

    def publish(self, manifest: Jsonable) -> None: ...
    def open_sidecar(
        self,
        name: str,
        *,
        head_cap: int | None = None,
        tail_cap: int | None = None,
    ) -> SidecarWriter: ...

    def read_manifest(self) -> Jsonable: ...

    def verify_sidecar(
        self,
        name: str,
        *,
        expected_digest: str,
        expected_head_length: int,
        expected_tail_length: int,
    ) -> None: ...
@dataclass(frozen=True, slots=True)
class SidecarSummary:
    head_length: int
    tail_length: int
    produced: int
    dropped: int
    digest: str

class SidecarWriter:
    def write(self, chunk: bytes) -> None: ...
    def finalize(self) -> SidecarSummary: ...

Filesystem and failure semantics

Canonical document publication creates a reserved unique temporary file with private permissions for each call, writes its complete canonical bytes, flushes and closes it, replaces the target in the same directory, and flushes and closes the directory. The case-insensitive .dr-store-document- prefix is reserved for these temporary files and cannot be used by document targets or Document Directory sidecars. Concurrent supported publishers use independent temporary files, and the last successful replacement is authoritative. Publication does not provide locks, compare-and-set, multi-file transactions, or ordering with Sidecar writes.

All-or-nothing visibility depends on the underlying filesystem honoring atomic same-directory replacement; network, synchronized, or other filesystems whose rename semantics are not established are outside current evidence. A final directory flush or close failure raises even though replacement may already be visible and does not roll the document back. Publication and allocation make no power-loss durability promise. Document Directory allocation uses a timestamp and UUID4, but a generated-name collision raises AllocationError rather than being retried, and allocation does not flush the caller-owned root directory.

Canonical files, Document Directories, and persistent SQLite storage capture a lexical absolute path at construction, so later working-directory changes do not redirect their operations. This does not freeze symlink targets. Directory durability is performed only where a pinned descriptor is already owned by the operation.

Canonical document reads open the named directory and regular direct child with required no-follow, directory-relative flags, then stream from the child descriptor they inspected. They read only to the configured byte bound plus the single byte needed to detect overflow, enforce the configured nesting-depth bound, and require one complete UTF-8 strict JSON value whose bytes are exactly canonical. Final-component symlinks and non-regular files are rejected, a replacement after open cannot redirect that read to a different inode, and platforms without the required descriptor operations fail closed.

Outside the reserved publication namespace, name validation prevents lexical traversal syntax. Sidecar creation and writes follow existing final-component symlinks and therefore require trusted, caller-controlled directory contents. Sidecar writer coordination remains the caller's concern. Sidecar finalization flushes the Sidecar descriptor before returning its summary, but it does not flush the Sidecar's directory entry or impose ordering on document publication. Sidecar verification also refuses final-component symlinks for both the Document Directory and named child, requires a regular direct child, and reads from the descriptor it inspected.

A failed Sidecar write raises AllocationError and may leave its descriptor open and its accounting state advanced. The writer is unusable by contract and must be abandoned; retrying it or finalizing it has no supported outcome.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dr_store-0.1.5.tar.gz (23.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dr_store-0.1.5-py3-none-any.whl (32.3 kB view details)

Uploaded Python 3

File details

Details for the file dr_store-0.1.5.tar.gz.

File metadata

  • Download URL: dr_store-0.1.5.tar.gz
  • Upload date:
  • Size: 23.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for dr_store-0.1.5.tar.gz
Algorithm Hash digest
SHA256 7405adb32a3cc776d56cb7ad508f3326903649b0bc02c4111d7679dca43e1d7b
MD5 df9dfb1621033587febd8cbf840add54
BLAKE2b-256 e234c0a439f8ad4463e96ed3c889a459314068d8f1731860813fbcd551d8cd8c

See more details on using hashes here.

Provenance

The following attestation bundles were made for dr_store-0.1.5.tar.gz:

Publisher: release.yml on danielle-rothermel/dr-store

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file dr_store-0.1.5-py3-none-any.whl.

File metadata

  • Download URL: dr_store-0.1.5-py3-none-any.whl
  • Upload date:
  • Size: 32.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for dr_store-0.1.5-py3-none-any.whl
Algorithm Hash digest
SHA256 5f4f3b429372f99b60a536bfbfc4dbde2487e39cb4fdcff7cfa6062b421a5b73
MD5 c54a842792de7e9e25209d2f596ab357
BLAKE2b-256 02e11a03c5f6d2c8851561cc931a130b16cfd528e72f4fc660ae8cb0ee0de5ed

See more details on using hashes here.

Provenance

The following attestation bundles were made for dr_store-0.1.5-py3-none-any.whl:

Publisher: release.yml on danielle-rothermel/dr-store

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page