Skip to main content
Yanked

This release has been yanked by its maintainers, and will be ignored by installers, except when explicitly specified.
Consider using release 0.1.2 instead.
Reason given by maintainers: Superseded by 0.1.1; test fixture wording cleanup

ChromeRAG

License: MIT Python 3.11+ GitHub Pages

HTML → RAG-ready Markdown that strips site-template chrome (nav, footer, CTAs, cookie banners) while keeping documentation, pricing tables, and article body text.

Built for enterprise RAG ingest, LLM chunking, vector indexing, and boilerplate / noise removal from scraped HTML — not for pixel-perfect web archiving.

Owns Does not own
HTML string / file → clean Markdown Crawling, Playwright, rate limits
Optional site-chrome learn → extract (STCE) Full browser rendering of empty JS shells (warns; caller must render first)
Schema.org → YAML front-matter Vector DB / embeddings
Precision / coverage priority knobs Hosted SaaS API

Paper: ChromeRAG: Ingest-Time Elimination of Site Template Noise for Enterprise Web RAG Release: v0.1.0
Author: Bhargava Chary Peddapudi


Why ChromeRAG (vs MarkItDown / Trafilatura)?

Tool Best at Gap for corporate web RAG
MarkItDown Office/PDF/HTML → Markdown for LLMs Keeps a lot of page chrome; not tuned to strip SaaS nav/footer
Trafilatura News/article main-content Weaker on docs hubs, pricing matrices, marketing shells
Readability Article extraction Often drops tables / side content needed for RAG
ChromeRAG Ingest-time chrome elimination + schema + tables Focused HTML→RAG Markdown (fetch stays in your crawler)

Benchmarks (public corpus)

Deterministic metrics (primary claims) — independent of any single extractor:

Metric Meaning Better
Recall (content_recall) Fraction of main/article text anchors kept Higher
Noise ret (noise_retention) Fraction of nav/footer chrome anchors kept Lower
Fbal (f_balanced) Balance of high recall + low noise Higher

Baselines compared: ChromeRAG (balanced / coverage / precision), Trafilatura, Readability, MarkItDown, markdownify, html2text, BeautifulSoup text.

Corpus: 373 URLs listed in poc/corpus_urls.json across docs, pricing, marketing, wiki, hub, news, article, cloud. Latest run fetched 277 HTML pages; 238 were scoreable (DOM content-anchor coverage ≥ 0.05; same fixed cohort for every tool); 39 thin pages are excluded from leaderboard means so empty JS shells are not silently averaged into SOTA claims.

Latest leaderboard (238 scoreable pages):

Method Recall ↑ Noise ↓ Fbal ↑
chromerag_coverage 0.701 0.005 0.800
chromerag (balanced) 0.678 0.005 0.783
trafilatura 0.662 0.019 0.752
markitdown 0.702 0.261 0.695
readability 0.456 0.016 0.535

Full tables + per-category breakdown: Results · docs/data/corpus_comparison_summary.md.

Optional LLM judge (secondary): stratified sample scored 1–5 on content keep / noise strip / structure when OPENAI_API_KEY is set (python -m poc.run_llm_eval). Does not replace Recall/Noise/Fbal.


Quick start

pip install chromerag
# Optional: ONNX MiniLM density pruning
# pip install "chromerag[dvdf]"

From source

git clone https://github.com/pedapudibhargav/ChromeRAG.git
cd ChromeRAG
python3 -m venv .venv
source .venv/bin/activate   # Windows: .venv\Scripts\activate

# Core library + CLI
pip install -e .

# Optional: comparison baselines + tests
pip install -e ".[dev,baselines]"

CLI

# Single page
chromerag extract path/to/page.html -o out.md --priority balanced --json-meta

# Priorities: precision | balanced | coverage
chromerag extract page.html -o clean.md --priority coverage

# Thin / JS-shell HTML prints WARNING on stderr (caller must Playwright-render first)
# chromerag extract spa.html -o out.md --fail-on-thin   # exit 3 if thin

# Learn site chrome across a folder, then batch-extract
chromerag learn data/raw -o data/chrome_models/site.json --min-pages 3
chromerag batch data/raw -o data/outputs --chrome-model data/chrome_models/site.json

Python

from chromerag import ChromeRAG, PipelineConfig, ContentPriority

html = open("page.html", encoding="utf-8").read()
result = ChromeRAG(
    config=PipelineConfig.from_priority(ContentPriority.BALANCED, enable_dvdf=False)
).extract(html, url="https://example.com/docs")

print(result.markdown)       # RAG-ready Markdown (+ YAML front-matter when Schema.org present)
print(result.front_matter)   # dict
print(result.tokens_estimate)
print(result.warnings)       # e.g. JS shell → render with Playwright first

Reproduce the evaluation

# 1) Fetch / refresh the public corpus (uses poc/corpus_urls.json)
python -m poc.run_corpus_comparison

# 2) Export tables + JSON into docs/ for GitHub Pages
python -m poc.export_site_results

# 3) Revalidate published numbers + classify thin pages
python scripts/revalidate_corpus.py

# 4) Optional LLM subset judge (needs OPENAI_API_KEY)
python -m poc.run_llm_eval --sample 40

# 5) Unit tests + regenerate docs/tests.html
./scripts/build_docs.sh

Skip re-fetch if HTML is already under data/raw/:

python -m poc.run_corpus_comparison --no-fetch

Docs site (GitHub Pages)

Static files live in docs/ (relative links only).

CI (.github/workflows/pages.yml) on every push to main:

  1. runs pytest
  2. rebuilds docs/tests.html + exports corpus results
  3. deploys docs/ → gh-pages branch

Enable once: Settings → Pages → Deploy from a branch → gh-pages / (root).
Site: https://pedapudibhargav.github.io/ChromeRAG/


Project layout

src/chromerag/     # library (HTML in → Markdown out)
poc/               # fetch, baselines, corpus comparison, LLM eval (not required at runtime)
docs/              # GitHub Pages site + published metrics
tests/             # unit tests
papers/softwarex/  # SoftwareX manuscript draft

Publishing (SoftwareX)

SoftwareX APC (journal OA fee) is paid only after acceptance, not at submission.

License

MIT — see LICENSE / LICENSE.txt (SoftwareX naming) · also Licence.txt.

Release files for chromerag 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for chromerag 0.1.0
File Size Uploaded
chromerag-0.1.0.tar.gz 32.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for chromerag 0.1.0
File Interpreter ABI Platform
chromerag-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 66.4 kB

Release files / chromerag-0.1.0.tar.gz

Download URL chromerag-0.1.0.tar.gz
Size 32.9 kB
Tags Source
SHA-256 checksum
How to use checksums
cb4b2e7bc6311385fecfb8d7623b569f235307fd2d8b18a301d3e588b7242154
BLAKE2b-256 checksum
How to use checksums
e9528a0449a4171b50d61d37622bafaa1e45f78a99e9b503ec53a53c395c3e39
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / chromerag-0.1.0-py3-none-any.whl

Download URL chromerag-0.1.0-py3-none-any.whl
Size 33.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9578b7d6c81bb2f847315b872751928011ce5912a91f58f42c88b82ef24eb183
BLAKE2b-256 checksum
How to use checksums
6a77c4ee61972f1834609e4c37e1ec99294b26aae5e31aec864abb91f6240ca2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

0.1.2

2 release files

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page