mmap_ninja_dataframe
Memory-mapped dataframe abstraction based on mmap_ninja
Run tests:
uvx --with-editable . --with joblib --with zstandard --with dnn_cool_synthetic_dataset --with opencv-contrib-python --with transformers pytest
PropertyResult
PropertyResult is a sparse column of computed results, stored as two parallel mmaps under a directory:
index— a numpy mmap of the indices the results were computed forresults— an mmap-ninja mmap of the corresponding results ("numpy","string", or"ragged")
results[j] is the result for index[j]. Because each result carries its own index, results can be recorded for any subset of indices, in any order.
Create and populate
Construct with an out_dir and the mmap_type for the results, then fill it with append / extend:
from mmap_ninja_dataframe import PropertyResult
pr = PropertyResult("my_store/sentiment", mmap_type="string")
pr.extend([0, 2], ["positive", "neutral"]) # results for indices 0 and 2
pr.append(1, "negative") # a single (index, result) pair
mmap_type selects the results mmap:
"string"→StringsMmap(text labels, summaries, …)"numpy"→ numpy mmap (scores, token counts, fixed-shape embeddings, …)"ragged"→RaggedMmap(variable-length sequences)
Results can be recorded partially and out of order:
pr = PropertyResult("my_store/embedding", mmap_type="numpy")
pr.extend([5], [np.random.rand(768).astype(np.float32)]) # index 5 first
pr.extend([0, 1], [emb0, emb1]) # earlier indices later
Recording an index that already has a result doesn't overwrite it in place — it appends a new (index, result) pair. Lookups always return the most recently recorded result for that index:
pr = PropertyResult("my_store/sentiment", mmap_type="string")
pr.append(2, "neutral")
pr.append(2, "positive")
pr[2] # "positive" -- the latest write wins
Look up by index
pr[2] # result recorded for index 2 (KeyError if absent)
pr.get(2) # same, but returns default (None) if absent
pr.get(2, "n/a") # with an explicit default
2 in pr # whether index 2 has a result
pr[idx] = value is equivalent to pr.append(idx, value); pr[indices] = values (a sequence of indices) is equivalent to pr.extend(indices, values).
Inspect
len(pr) # number of recorded (index, result) pairs -- not the number of *distinct* indices
pr.indices() # the indices, in insertion order (may contain repeats -- see "latest wins" above)
pr.mmap_type # "string" / "numpy" / "ragged"
pr.name # basename of out_dir, e.g. "sentiment"
pr.out_dir # the directory it's persisted under
pr.index # raw numpy index mmap (None until the first write)
pr.results # raw results mmap (None until the first write)
Reopen
mmap_type is required on every construction call, including when reopening — it is then validated against what's already on disk, and a mismatch raises ValueError:
pr = PropertyResult("my_store/sentiment", mmap_type="string")
# ValueError if "my_store/sentiment" already holds results of a different mmap_type.
TextPropertiesMmap
TextPropertiesMmap is a deduplicated store of texts annotated with named, independently-computed properties (sentiment, embeddings, summaries, …). Each unique text gets a stable integer index; each property is a PropertyResult keyed by those indices, supplied explicitly by the caller.
How text lookup works
Texts are stored as a StringsMmap (store.text). Alongside it, each text's sha256 content hash is stored in a fixed-width numpy memmap (store.content_hash, dtype "<U64"), used to detect duplicates and to resolve a text back to its index.
Content hashes are never loaded into memory as a Python dict. Instead, a second numpy memmap (content_hash_sorter) holds the permutation that sorts content_hash. Looking up a hash is then a binary search — np.searchsorted(content_hash, target, sorter=sorter) for a single text, or one vectorized np.searchsorted call for a batch — instead of a linear scan or an in-memory dict.
text/content_hash are kept in sync immediately after every append (cheap — just reopening the mmaps). The sorter is the expensive part (O(n log n)), so it's handled separately: rather than recomputing it on every reload, it's rebuilt lazily the next time a hash-based lookup actually needs it (or eagerly, if you ask for that — see rebuild_sorter below). Opening a store you're only going to write to, without any lookups, never pays that cost at all.
Every lookup checks that the sorter's length still matches content_hash's. This should never fail through normal use of a single store instance — it's a safety net for cases like two TextPropertiesMmap instances open on the same out_dir, where one appends texts the other doesn't know about. A mismatch raises RuntimeError telling you to call rebuild_sorter() (or reopen the store) rather than silently returning wrong results.
Create and add texts
from mmap_ninja_dataframe import TextPropertiesMmap, PropertyResult
store = TextPropertiesMmap.from_texts(
"my_store",
texts=["The quick brown fox.", "Hello, world!"],
properties=[PropertyResult("my_store/sentiment", mmap_type="string")],
)
# Add more texts, deduplicating by content hash. Returns the index for each
# input text: texts already in the store resolve to their existing index,
# duplicates within the same call resolve to the same new index, and only
# genuinely new texts get appended. `extend` is an alias for `update`.
indices = store.update(["The quick brown fox.", "A new sentence."])
# indices == [0, 2] -- "The quick brown fox." already existed at index 0
update/extend take a rebuild_sorter flag (default True) that controls only the sorter rebuild -- text/content_hash are always refreshed regardless. True rebuilds the sorter immediately after appending; False defers that O(n log n) cost until it's actually needed by a hash-based lookup (index_of_text, indices_for_texts, or update's own dedup check on a later call), which rebuilds it lazily from the now-current content_hash. Results are correct either way — only when the rebuild happens changes, which matters when appending in many small batches with no lookups in between:
store.update(["A brand new sentence."], rebuild_sorter=False) # sorter rebuild deferred
store.index_of_text("A brand new sentence.") # still resolves correctly -- triggers the lazy rebuild
store.rebuild_sorter() # or: force the rebuild eagerly yourself
from_texts builds a fresh store from a list of texts. To reopen an existing one, construct TextPropertiesMmap directly with the same out_dir — properties are not auto-discovered from disk, so pass the same properties list again:
store = TextPropertiesMmap(
"my_store",
properties=[PropertyResult("my_store/sentiment", mmap_type="string")],
)
Record and read property results
The store behaves like a named collection of properties. Resolve texts to indices, then write to the property directly:
store["sentiment"].extend(indices, ["neutral", "positive"])
store["sentiment"] # the PropertyResult
store.get_property("sentiment") # same thing
store.get_property_names() # ["sentiment"]
store.add_property(PropertyResult("my_store/summary", mmap_type="string"))
store["summary"] = PropertyResult("my_store/summary", mmap_type="string") # equivalent to add_property
store.delete_property("summary") # unregisters it and deletes its directory
Property names "text", "content_hash", and "content_hash_sorter" are reserved (used internally) and raise ValueError if registered.
Look up texts
store.index_of_text("Hello, world!") # 1, or None if not present
store.indices_for_texts(["Hello, world!"]) # [1] -- vectorized; raises KeyError naming
# the first text not found in the store
store.get_text_properties("Hello, world!")
# {"unprocessed": ["sentiment"]} -- no result recorded yet for this text/property
store.get_properties_for_texts(["The quick brown fox."])
# {"text": [...], "content_hash": array([...], dtype='<U64'), "idx": [0], "sentiment": ["neutral"]}
# Raises ValueError if any requested text is missing a result for any registered property.
Check progress
store.get_unprocessed_indices_for_property("sentiment") # numpy array of indices with no result yet
store.get_unprocessed_counts() # {"sentiment": 1} -- only not-yet-complete properties
Inspect
len(store) # number of distinct texts
store.text # StringsMmap of the texts
store.content_hash # numpy memmap of sha256 hexdigests, dtype "<U64"
Release files for mmap-ninja-dataframe 0.9.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mmap_ninja_dataframe-0.9.1.tar.gz | 22.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mmap_ninja_dataframe-0.9.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 39.9 kB
Release files / mmap_ninja_dataframe-0.9.1.tar.gz
| Download URL | mmap_ninja_dataframe-0.9.1.tar.gz |
|---|---|
| Size | 22.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1452f21ed140ebf997b77d69034e92e4fca24a7da0539634c7f5bf01bd7146cf
|
|
BLAKE2b-256 checksum How to use checksums |
306591fbe4b33d051b1dc5d194bb4c4ca7f0ef74f08ebb8799e820eb3905487b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.9
|
Release files / mmap_ninja_dataframe-0.9.1-py3-none-any.whl
| Download URL | mmap_ninja_dataframe-0.9.1-py3-none-any.whl |
|---|---|
| Size | 17.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c7e91f493b81ce5b9ae7e4e720d3ec17baf8ce30f711aee60fbc427195a3a85e
|
|
BLAKE2b-256 checksum How to use checksums |
b05f47d0302fcbf5ce8b46882a724034a5271ba9d7199b18a1a3e5be8160a09a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.9
|