Skip to main content

llama-index-postprocessor-vecr

Retention-guaranteed node compression for LlamaIndex RAG pipelines. This partner package wraps vecr-compress — the only LLM context compressor that makes a deterministic retention contract: order IDs, dates, URLs, emails, and code references are pinned by an auditable regex whitelist before any token-budget packing runs. Filler prose is dropped, high-signal sentences are ranked by question-aware Jaccard scoring, and your retrieved nodes arrive at the synthesizer with all structured facts intact.

Install

pip install llama-index-postprocessor-vecr

This installs vecr-compress and llama-index-core automatically.

30-second example

from llama_index.core.schema import NodeWithScore, TextNode
from llama_index.postprocessor.vecr import VecrNodePostprocessor

nodes = [
    NodeWithScore(
        node=TextNode(
            id_="node-1",
            text=(
                "The refund for order ORD-99172 was approved on 2026-03-15. "
                "The total amount was $1,499.00. "
                "Please see https://refunds.example.com/ORD-99172 for details."
            ),
        ),
        score=0.92,
    ),
    NodeWithScore(
        node=TextNode(
            id_="node-2",
            text="Hi! I hope this message finds you well. Have a great day.",
        ),
        score=0.31,
    ),
]

processor = VecrNodePostprocessor(budget_tokens=60)
kept = processor.postprocess_nodes(nodes, query_str="refund status for ORD-99172")

for n in kept:
    print(n.node.id_, "->", n.node.text[:80])
# node-1 -> The refund for order ORD-99172 was approved ...
# (node-2 dropped: pure filler, no retention match)

The node containing ORD-99172, 2026-03-15, $1,499.00, and the URL is kept because each of those tokens fires a retention rule. The filler-only node is dropped even at an aggressive 60-token budget.

Query-aware scoring

Pass query_str to bias retained sentences toward the user's question:

kept = processor.postprocess_nodes(
    nodes,
    query_str="what is the refund amount?",
)

Internally, this runs Jaccard overlap between each sentence and the query before the knapsack budget packing step.

Compression telemetry

result = processor.compress_with_report(nodes, query_str="refund amount")
print(f"{result.original_tokens} -> {result.compressed_tokens} tokens ({result.ratio:.1%})")
print(f"Pinned facts: {len(result.retained_matches)}")

Custom retention rules

import re
from vecr_compress import RetentionRule, DEFAULT_RULES

custom_rules = DEFAULT_RULES.with_extra([
    RetentionRule(name="sku", pattern=re.compile(r"\bSKU-[A-Z0-9]{6}\b")),
])
processor = VecrNodePostprocessor(budget_tokens=1500, retention_rules=custom_rules)

Links

License

Apache 2.0 — see LICENSE.

Metadata

Release files for llama-index-postprocessor-vecr 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llama-index-postprocessor-vecr 0.1.0
File Size Uploaded
llama_index_postprocessor_vecr-0.1.0.tar.gz 7.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llama-index-postprocessor-vecr 0.1.0
File Interpreter ABI Platform
llama_index_postprocessor_vecr-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 14.9 kB

Release files / llama_index_postprocessor_vecr-0.1.0.tar.gz

Download URL llama_index_postprocessor_vecr-0.1.0.tar.gz
Size 7.7 kB
Tags Source
SHA-256 checksum
How to use checksums
98378bf1b3e0d4864c16dc5276dcdd23fbbd5af65d1d437ef7273292da99cfad
BLAKE2b-256 checksum
How to use checksums
548e1f75eca397448277f6122598f5f81b777f4e4d72e4186338b08797c6cffa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.5

Release files / llama_index_postprocessor_vecr-0.1.0-py3-none-any.whl

Download URL llama_index_postprocessor_vecr-0.1.0-py3-none-any.whl
Size 7.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
65cb531a093ea54c738536c97195df25e30ae8758fd8a39d8e134ab547284ac5
BLAKE2b-256 checksum
How to use checksums
bcf7806987ec47462e2e24dfdb36729ac77c8ccb5d5c8d3c0cd2b5f70fc096cb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.5

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page