maskflow-llamaindex
MaskFlow for LlamaIndex: keep PII out of your RAG pipeline. Three pieces:
MaskflowNodePostprocessor— a drop-in forllama_index.core.postprocessor.PIINodePostprocessorthat masks PII in retrieved nodes before the response synthesizer. No LLM call, no HuggingFace model; MaskFlow's local engine, so Indian identifiers (Aadhaar, PAN, GSTIN, UPI, IFSC, ABHA, Indian names / addresses) are covered alongside the generic PII.MaskflowIngestionTransform— aTransformComponentthat masks node text at ingestion, so raw PII is never embedded or written to the vector store.unmask_response/MaskflowQueryEngine— restore the originals in the synthesized answer from the per-node maps.
MIT, no gates, no telemetry.
Install
pip install maskflow-llamaindex
Pulls llama-index-core. The first detection run downloads a small spaCy
model for the name/address recognizers; pass patterns_only=True to skip
it.
Query-time masking (index already built)
from maskflow_llamaindex import MaskflowNodePostprocessor, unmask_response
query_engine = index.as_query_engine(node_postprocessors=[MaskflowNodePostprocessor()])
response = query_engine.query("What is Ramesh's PAN?")
# the synthesizer LLM saw "<PAN_1>"; restore the real value for the caller:
answer = unmask_response(str(response), response.source_nodes)
Or wrap the engine so you never forget the unmask step:
from maskflow_llamaindex import MaskflowQueryEngine
engine = MaskflowQueryEngine(
index.as_query_engine(node_postprocessors=[MaskflowNodePostprocessor()])
)
print(engine.query("What is Ramesh's PAN?")) # already restored, streaming too
By default one MaskFlow session is shared across every node in a call, so
<PERSON_NAME_1> is the same person in every retrieved chunk.
PIINodePostprocessor numbers each node independently.
Ingestion-time masking (PII never reaches the store)
from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter
from maskflow_llamaindex import MaskflowIngestionTransform
pipeline = IngestionPipeline(
transformations=[
SentenceSplitter(),
MaskflowIngestionTransform(), # default strategy: redact
embed_model,
]
)
nodes = pipeline.run(documents=docs)
The default strategy="redact" ([REDACTED_PAN]) is not reversible —
there is no mapping, so nothing sensitive is stored. Other strategies:
surrogate (a plausible fake value) and replace (<PAN_1> tokens).
store_mapping=True writes a reverse map into node metadata; that then
lands in the vector store, so it warns.
Migrating from PIINodePostprocessor
# from llama_index.core.postprocessor import PIINodePostprocessor
from maskflow_llamaindex import MaskflowNodePostprocessor as PIINodePostprocessor
mask_pii(text) -> (str, dict), _postprocess_nodes, the
__pii_node_info__ metadata key, and the embed/LLM metadata exclusions are
all the same. MaskflowNodePostprocessor does not take an llm= argument
(it needs none); it adds strategy, min_confidence, patterns_only,
consistent_across_nodes, and mask_query.
PII safety
Query-time maps live only for the query, travelling with
response.source_nodes; nothing is logged. The ingestion transform's
default is non-reversible, so it stores no map at all. See docs/llamaindex.md
in the MaskFlow repo for the design notes.
Metadata
Release files for maskflow-llamaindex 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| maskflow_llamaindex-0.1.1.tar.gz | 11.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| maskflow_llamaindex-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 22.6 kB
Release files / maskflow_llamaindex-0.1.1.tar.gz
| Download URL | maskflow_llamaindex-0.1.1.tar.gz |
|---|---|
| Size | 11.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c994148a133fb2402e57ba555583a086e751e1a7faad01515fda68f6d6b23c5e
|
|
BLAKE2b-256 checksum How to use checksums |
5bed74f5c9b611efeaa3938f08c9293f7e47a2962163df0bb762bb644c2544e9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.
Transparency logRelease files / maskflow_llamaindex-0.1.1-py3-none-any.whl
| Download URL | maskflow_llamaindex-0.1.1-py3-none-any.whl |
|---|---|
| Size | 10.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c437596bfe903e5ded8d842f6797ed74cc925aa5f776d608b3da192e92bd289a
|
|
BLAKE2b-256 checksum How to use checksums |
cb75a6611f74969b4b7bfcc85ebaad1110047f861ecedd780fe518a32c2e2d32
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.
Transparency log