Skip to main content

context-dedup

PyPI version License: MPL 2.0

Detect and remove redundant context before sending it to an LLM.

context-dedup is a small, deterministic Python library for finding exact and near-duplicate chunks in RAG results, agent pipelines, and assembled prompts. It uses lexical word n-gram overlap: no embeddings, no LLM calls, and no runtime dependencies.

Installation

pip install context-dedup

Usage

from context_dedup import deduplicate, inspect_context, similarity

chunks = [
    "The Eiffel Tower is located in Paris.",
    "The Eiffel Tower is located in Paris. It was completed in 1889.",
    "Madrid is the capital of Spain.",
]

print(similarity(chunks[0], chunks[1]))
print(inspect_context(chunks))
print(deduplicate(chunks))

Why context-dedup?

Retrieval and prompt assembly can repeat the same facts across overlapping passages, wasting context-window space and attention. context-dedup provides a transparent heuristic between raw retrieval and an LLM call, without adding a model, service, database, or provider SDK.

It is intentionally focused:

  • No embeddings or transformer models
  • No LLM calls or external APIs
  • No runtime dependencies or infrastructure

Inspect context

inspect_context returns an ordinary, serializable dictionary with pair scores, connected duplicate groups, recommended representatives, removable indices, and aggregate redundancy estimates:

report = inspect_context(chunks)

print(report["redundant_pairs"])
print(report["groups"])
print(report["estimated_redundant_words"])

Thresholds are configurable. Their defaults are practical heuristics, not statistically universal values:

report = inspect_context(
    chunks,
    similarity_threshold=0.75,
    containment_threshold=0.85,
    n=3,
)

Deduplicate context

The default strategy keeps the longest chunk in each duplicate group, preserving more information. Use first to keep the earliest chunk instead:

clean_chunks = deduplicate(chunks)
first_chunks = deduplicate(chunks, strategy="first")

Objects with metadata are supported through key; returned objects are the originals:

retrieved = [
    {"text": "Refunds are available within 30 days.", "source": "policy.pdf", "page": 4},
    {"text": "Refunds are available within 30 days. Contact support to begin.", "source": "faq.pdf", "page": 2},
]

clean = deduplicate(retrieved, key=lambda item: item["text"])

Algorithm

Text is lowercased, stripped, and normalized to single spaces. The library builds sets of word trigrams by default, then calculates Jaccard similarity and containment in both directions. A pair is redundant when either configured threshold is reached. Connected components combine transitive pairs into groups, and selection is deterministic.

Limitations

This is a lexical heuristic, not a semantic-equivalence detector. Paraphrases with different wording may not match, while repeated wording can match despite different meaning. Punctuation remains part of word tokens. Version 0.1.0 compares every pair in O(n²) time, which is appropriate for contexts with tens or hundreds of chunks but not large document collections.

Use cases

  • Remove overlap from RAG retrieval results before prompt construction
  • Inspect context redundancy in LLM and agent pipelines
  • Deduplicate passages assembled from multiple document sources
  • Reduce repeated prompt content without provider-specific infrastructure

Features

  • Deterministic word n-gram Jaccard similarity and directional containment
  • Transitive duplicate groups and serializable inspection reports
  • longest and first representative strategies
  • Metadata-preserving key support
  • Configurable thresholds and n-gram size
  • Standard-library implementation with no runtime dependencies

Issues

Report issues in the GitHub issue tracker.

Author

Eduardo J. Barrios — edujbarrios@outlook.com

License

Mozilla Public License 2.0

Release files for context-dedup 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for context-dedup 0.1.0
File Size Uploaded
context_dedup-0.1.0.tar.gz 13.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for context-dedup 0.1.0
File Interpreter ABI Platform
context_dedup-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 24.9 kB

Release files / context_dedup-0.1.0.tar.gz

Download URL context_dedup-0.1.0.tar.gz
Size 13.0 kB
Tags Source
SHA-256 checksum
How to use checksums
27cc4e86e8dc97ae71a911c6ea86f62890f14df6df3c03398eb6e1c4c5de73f3
BLAKE2b-256 checksum
How to use checksums
4770e0a59900fc9d62039477bed7c9bfea8947195352b303f03acfcca88eddb8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.7

Release files / context_dedup-0.1.0-py3-none-any.whl

Download URL context_dedup-0.1.0-py3-none-any.whl
Size 11.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b1d92c4bebe1994fea6525aff9b883681cd447228957de175d36c77bc2eeb8d2
BLAKE2b-256 checksum
How to use checksums
d73a301bd22bea9a9d8a6b22a702fda097aa912eb36d875f1f58a72a399fa7b3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.7

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page