Lightweight LMDB-backed object store for Python
Project description
lmdb-object-store
Lightweight thread-safe Python object store on top of LMDB. Provides a dict-like API (with buffering and atomic multi-put), automatic map-size growth, and fast zero-copy reads without having to handle LMDB's lower-level details.
Features
- Dict-like interface:
store[key] = obj,obj = store.get(key),del store[key],key in store. - Atomic multi-put:
put_many(items)writes large amounts of data atomically in a single transaction. - Write buffering: small writes are batched in memory;
flush()persists them (auto-flush on size threshold). - Auto map-size growth: retries on
MapFullErrorby growing the LMDB map (2× or +64MiB), up to an optional cap. - Fast reads: Uses LMDB zero-copy reads to minimize copies;
get_many()efficiently combines buffered and DB reads. - Flexible keys: bytes/bytearray/memoryview always supported;
strkeys allowed whenkey_encoding='utf-8'option is used.
Installation
pip install lmdb-object-store
Requirements
- Python 3.10+
- The
lmdbPython package (wheels available for major platforms)
Quick start
from lmdb_object_store import LmdbObjectStore
# Create or open an LMDB-backed object store
with LmdbObjectStore(
"path/to/db",
batch_size=1000, # flush buffer when it reaches this many entries
autoflush_on_read=True, # flush pending writes before reads
key_encoding="utf-8", # allow str keys (encoded with UTF-8)
# Any lmdb.open(...) kwargs may be passed here, e.g.:
map_size=128 * 1024 * 1024, # 128 MiB initial map size
subdir=True, # create directory layout
readonly=False,
# max_map_size is recognized (cap for auto-resize):
max_map_size=4 * 1024 * 1024 * 1024, # 4 GiB cap
) as store:
# Put / get like a dict (values are pickled)
store["user:42"] = {"name": "Ada", "plan": "pro"}
print(store.get("user:42")) # {'name': 'Ada', 'plan': 'pro'}
# Existence checks
if "user:42" in store: # __contains__ has NO side effects (no flush)
assert store.exists("user:42", flush=False) is True
# Delete
del store["user:42"]
# Batch write, atomically (single transaction)
items = {f"k{i}": {"i": i} for i in range(10_000)}
store.put_many(items)
# Fetch many at once
found, not_found = store.get_many(["k1", "kX"], decode_keys=True)
# found -> {'k1': {'i': 1}}, not_found -> ['kX']
# Clean close with strict error policy if needed:
store = LmdbObjectStore("path/to/db", key_encoding="utf-8")
try:
# ... work with store ...
store.close(strict=True) # re-raise if final flush fails (after cleanup)
finally:
# idempotent
try: store.close()
except Exception: pass
API Overview
Constructor
LmdbObjectStore(
db_path: str,
batch_size: int = 1000,
*,
autoflush_on_read: bool = True,
key_encoding: str | None = None,
key_errors: str = "strict",
str_normalize: str | None = None,
**lmdb_kwargs,
)
- db_path: LMDB environment path (passed to
lmdb.open). - batch_size: pending buffer size threshold to auto-flush.
- autoflush_on_read: if
True, flushes buffer before reads (get,get_many,exists(flush=None)). - key_encoding: enable
strkeys (e.g."utf-8"). IfNone, only bytes-like keys are allowed. - key_errors: error strategy for encoding
strkeys ("strict","ignore","replace", ...). - str_normalize: Unicode normalization for
strkeys (e.g.,"NFC","NFKC"). - lmdb_kwargs: forwarded to
lmdb.open(...)(e.g.,map_size,subdir,readonly, etc). Special:max_map_size(cap for automatic map growth) is also recognized.
Put / Get
store.put(key, obj) # buffer write
obj = store.get(key, default=None) # read; from buffer if present, else DB
store.flush() # persist the write buffer
- Values are serialized with
pickle(highest protocol). get()uses zero-copy buffers internally; unpickling happens once per value.
Atomic multi-put
store.put_many(items: Mapping[Any, Any] | Iterable[tuple[Any, Any]])
- Writes all items in one LMDB write transaction.
- If a
MapFullErroroccurs, the store will grow the map (2× or +64MiB) up tomax_map_sizeand retry from the beginning. - Note:
put_many()first flushes any pending buffered writes; the atomic transaction only includes theitemspassed to this call.
Get many
found, not_found = store.get_many(
keys: Sequence[Any],
*,
decode_keys: bool = False,
decode_not_found: bool | None = None, # None → follow decode_keys
)
-
Returns a tuple:
found:{key: value}for keys found (key type isbytesby default, orstrifdecode_keys=True).not_found: list of input keys not found (decoded tostrifdecode_not_found=True).
-
Efficiently merges results from the write buffer and DB;
autoflush_on_readapplies unless overridden via other APIs.
Existence & containment
store.exists(key, *, flush: bool | None = None) -> bool
key in store # __contains__ → NO flush, purely checks current state
flush=None(default) followsautoflush_on_read.flush=Falseguarantees no implicit flush (useful for side-effect-free checks).
Deletion
del store[key] # KeyError if not present
store.delete(key) # schedules deletion via buffer (dict-like semantics)
Lifecycle
with LmdbObjectStore(...) as store:
...
# or
store.close(strict: bool = False)
close(strict=True)re-raises the last flush error after closing the environment; otherwise it logs and completes.
Concurrency Model
-
Internally uses
RLock+Conditionand a simple reader count to coordinate:- Multiple concurrent readers are allowed.
- Writers hold the lock (buffer mutation + flush/commit).
close()sets a “closing” flag and waits until the active reader count reaches zero.
-
Designed for thread-safety within a single process. While LMDB itself supports multi-process access, this wrapper's locking is process-local; if you need multi-process writes, coordinate at a higher level.
Map Size & Auto-Resize
-
Initial size is given by
map_size(forwarded tolmdb.open). -
When a write/commit hits
MapFullError:- The store grows to
max(current*2, current+64MiB), capped atmax_map_sizeif provided, - and retries the operation.
- The store grows to
-
If the cap is reached and still insufficient, the error is propagated.
Performance Tips
- Batch writes: Keep
batch_sizelarge enough for your workload; callflush()at logical boundaries. - Use
put_many()for bulk inserts—single transaction with fewer fsyncs. - Avoid unnecessary decodes: If you don't need
strkeys on output, leavedecode_*parameters off. - Key type: If possible, pass bytes keys directly (saves encoding overhead).
Error Handling
- Unpickling failures are raised as a
RuntimeError(with key context) when reading from buffer/DB. - Missing keys:
__getitem__anddelraiseKeyError;get()returnsdefault. - Final flush failure on close: re-raised if
strict=True; otherwise logged.
Security Note
This library uses Python pickle for value serialization. Never unpickle data from untrusted sources.
Configuration Reference
-
batch_size: int– buffer size threshold to auto-flush (default: 1000). -
autoflush_on_read: bool– flush before reads (default:True). -
key_encoding: Optional[str]– enablestrkeys (e.g.,"utf-8"). IfNone, only bytes-like keys are accepted. -
key_errors: str– encoding error handling ("strict","ignore","replace"). -
str_normalize: Optional[str]– Unicode normalization forstrkeys ("NFC","NFKC", ...). -
lmdb_kwargs– forwarded tolmdb.open(...):map_size,subdir,readonly,lock, ...max_map_size(recognized by this wrapper to cap auto-growth).
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file lmdb_object_store-0.1.0.tar.gz.
File metadata
- Download URL: lmdb_object_store-0.1.0.tar.gz
- Upload date:
- Size: 30.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
309a2b91547787a86484af6a34aa05ac86281952f38fa162e941dee87f863c45
|
|
| MD5 |
4fe6067073a4f8b93597a2527acb00d3
|
|
| BLAKE2b-256 |
2c36d2c6f74ac422a9c8dd4bfcb21aaa0e400241bb7881177cbebd358122b436
|
File details
Details for the file lmdb_object_store-0.1.0-py3-none-any.whl.
File metadata
- Download URL: lmdb_object_store-0.1.0-py3-none-any.whl
- Upload date:
- Size: 12.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
327b61ee6f3247a12e272e19190493f556f90882b6ac077e927da0c7be0e6362
|
|
| MD5 |
0e27bb87b0da283582488d1452a2c254
|
|
| BLAKE2b-256 |
9acf8963b7e4c682ce1e12245292e9411751ba237030db4e35223889113f8aac
|