Skip to main content

ragpreflight

PyPI version Python 3.10+ License: MIT Downloads DOI

The pre-ingestion RAG audit tool. Catch document failures before you embed them — not after your chatbot starts hallucinating.

pip install ragpreflight
ragpreflight scan my_document.pdf
╭──────────────────────────────────────────────────────────────────────╮
│ ragpreflight — Document Readiness Report                             │
│ File: quarterly_report.pdf                                           │
│ Score: 44/100  █████████░░░░░░░░░░░                                  │
│ Format: PDF  ·  Size: 2.1 MB  ·  Pages: 34  ·  Extractable: 0%      │
╰──────────────────────────────────────────────────────────────────────╯

  CRITICAL  ocr      [pages 3, 7, 11] l / I / 1 confusion — 23 matches
                     Scanner misread lowercase l, capital I, and digit 1 as each other.
                     Examples: "cIinical" → "clinical", "1evel" → "level"
                     Keyword search and semantic retrieval both fail on corrupted tokens.
                     → Re-run OCR at higher DPI, or apply post-correction (pyspellchecker).

  CRITICAL  ocr      [pages 5, 9] mid-word spaces — 14 matches
                     Spaces inserted inside words during scan: "pati ent" → "patient"
                     Each broken word becomes two meaningless tokens in the index.
                     → Apply OCR post-correction before ingestion.

  CRITICAL  ocr      [pages 2, 14–17] rn / m split — 8 matches
                     The letter m was split into rn: "inforrnation" → "information"
                     Common in low-DPI scans of serif fonts.
                     → Re-scan at 300 DPI minimum or run post-OCR cleanup.

  CRITICAL  content  0% of pages have extractable text. This is a scanned PDF.
                     → Apply OCR (Tesseract, AWS Textract, Google Document AI) before ingestion.

  WARNING   structure 8 table(s) detected — will chunk as garbled text without special handling.
                     → Use a table-aware extractor (pdfplumber, Camelot, LlamaParse).

  WARNING   content  Possible PII: 12 email addresses found on pages 4, 18, 22.
                     → Review and redact before ingestion into a shared RAG corpus.

  INFO      metadata No title, author, or creation date in document metadata.
                     → Add metadata to improve retrieval ranking and attribution.

7 issue(s) found  ·  Score: 44  ·  4 critical

RAG Failure Taxonomy  (doi:10.18653/v1/2026.trustnlp-main.27)
  OCR artifacts  → F3  Document Quality Failure       [direct]
  Scanned PDF    → F3  Document Quality Failure       [direct]
  Tables         → F7  Structure-Unaware Chunking     [risk signal]
  PII            → F23 PII / Compliance Leak          [direct]
  No metadata    → F11 Low Recall / Ranking Failure   [risk signal]

The gap nobody talks about

Every RAG evaluation tool — RAGAS, DeepEval, TruLens, RAGChecker — runs after you build your system. They need a live retriever, real queries, LLM-generated outputs, and an LLM judge to score them. By that point, bad documents are already embedded. You're debugging a production system, not preventing the problem.

ragpreflight runs before you embed anything. Give it a folder of files. It tells you which ones will cause failures and why — in seconds, with no API keys, no internet connection, no running model.

Your documents  →  [ragpreflight]  →  fix issues  →  embed  →  RAG system
                         ↑
                  This is the gap.
            Nothing else runs here.

It is grounded in peer-reviewed research: 33 failure modes across 7 pipeline stages from Garani 2026 — the first systematic taxonomy of RAG failure modes published at TrustNLP 2026 (ACL). ragpreflight is the first open-source tool that links every detected issue to a named failure mode from that published taxonomy.


How it compares

ragpreflight RAGAS DeepEval TruLens RAGChecker Unstructured
When it runs pre-ingestion post-gen post-gen post-gen post-gen ingestion
Needs a live RAG system ❌ ✅ ✅ ✅ ✅ ❌
Needs an LLM to run ❌ ✅ ✅ ✅ ✅ ⚠️ optional
Needs labeled queries / golden sets ❌ ⚠️ some metrics ⚠️ some metrics ⚠️ some metrics ✅ ❌
Fully offline ✅ ❌ ❌ ❌ ❌ ⚠️ OSS only
OCR artifact detection ✅ ❌ ❌ ❌ ❌ ❌
Chunking boundary quality ✅ ❌ ❌ ❌ ❌ ❌
Near-duplicate detection ✅ ❌ ❌ ❌ ❌ ❌
PII in source documents ✅ ❌ ⚠️ ❌ ❌ ❌
Staleness detection ✅ ❌ ❌ ❌ ❌ ❌
Taxonomy-grounded (peer-reviewed) ✅ 33 modes ❌ ❌ ❌ ❌ ❌
Multi-format (PDF/DOCX/CSV/HTML/MD…) ✅ ❌ ❌ ❌ ❌ ✅
Readiness score 0–100 ✅ ❌ ❌ ❌ ❌ ❌
HTML audit report ✅ ❌ ⚠️ cloud dashboard ✅ ❌ ❌
CI/CD exit code gating ✅ ⚠️ ✅ ❌ ❌ ❌
SARIF output (GitHub Code Scanning) ✅ ❌ ❌ ❌ ❌ ❌

These tools are complementary, not competing. ragpreflight cleans and validates your corpus before ingestion. RAGAS/DeepEval/TruLens evaluate your live RAG system after deployment. Use both.


Quickstart

Scan a single document

ragpreflight scan quarterly_report.pdf
ragpreflight scan contract.pdf --profile strict   # higher thresholds
ragpreflight scan report.pdf --json               # machine-readable
ragpreflight scan report.pdf --format html --output report.html

Audit your entire knowledge base

ragpreflight audit ./knowledge_base/
ragpreflight audit ./knowledge_base/ --format html --output audit.html
╭──────────────────────────────────────────────────────────────╮
│ ragpreflight — Corpus Readiness Report                       │
│ Directory: ./knowledge_base/                                 │
│ Documents: 17  ·  Avg Score: 77.9  ███████████████░░░░░      │
│ Duplicate Groups: 0                                          │
╰──────────────────────────────────────────────────────────────╯

 Score  File                   Format    Issues  Critical
──────────────────────────────────────────────────────────
     0  empty_file.txt         TXT            1         1   ← fix first
    44  scanned_old.pdf        PDF            5         1   ← needs OCR
    85  contract_draft.pdf     PDF            2         0
    90  data_export.csv        CSV            0         0
    92  technical_spec.pdf     PDF            2         0
    94  meeting_notes.docx     DOCX           2         0

Gate your ingestion pipeline (CI/CD)

# Fail the pipeline if any document scores below 60
ragpreflight score my_doc.pdf          # prints: 92

score=$(ragpreflight score my_doc.pdf)
[ "$score" -ge 60 ] || exit 1

SARIF output for GitHub Code Scanning

ragpreflight scan my_doc.pdf --format sarif --output results.sarif

Upload results.sarif to GitHub Code Scanning and document issues appear as pull request annotations.

Explore the taxonomy

ragpreflight taxonomy list                    # all 33 failure modes
ragpreflight taxonomy list --stage retrieval  # filter by pipeline stage
ragpreflight taxonomy show F7                 # full detail for one mode
ragpreflight coverage                         # what this tool can/cannot detect

Python API

from ragpreflight import scan_document, audit_corpus

# Single document
report = scan_document("my_doc.pdf")
print(report.score)         # 83
print(report.issues)        # list[Issue] — each with category, severity, suggestion

for issue in report.issues:
    print(issue.severity.value, issue.category.value, issue.message)
    if issue.taxonomy_refs:
        # Each issue linked to Garani 2026 failure modes
        for ref in issue.taxonomy_refs:
            print(f"  → {ref.mode_id} ({ref.relationship})")

# Corpus
corpus = audit_corpus("./knowledge_base/")
print(corpus.average_score)        # 77.9
print(corpus.duplicate_groups)     # [[file1, file2], ...]

# Programmatic fail gate
critical = [i for d in corpus.documents for i in d.issues if i.severity.value == "critical"]
if critical:
    raise ValueError(f"{len(critical)} critical issues — fix before ingesting")

All outputs are typed dataclasses, not dicts. Full type hints. Zero global state.


Supported file formats

Format What's checked
PDF (text-based) OCR artifacts, structure, tables, metadata, PII, content density
PDF (scanned) Flags 0% extractability, recommends OCR — score penalised
DOCX Structure, encoding, tables, metadata
TXT / Markdown Encoding, OCR patterns, header hierarchy
CSV / TSV Column consistency, encoding, structural integrity
HTML Boilerplate ratio, structure
XLSX Sheet structure, encoding
PPTX Slide content, structure
IPYNB Cell content, code/text ratio
SRT Transcript length, artifact detection

Quality profiles

ragpreflight scan doc.pdf --profile permissive  # chatbots, internal tools
ragpreflight scan doc.pdf --profile standard    # enterprise search (default)
ragpreflight scan doc.pdf --profile strict      # medical, legal, financial
Profile Min score OCR tolerance Similarity threshold
permissive 40 < 10% 0.95
standard 60 < 5% 0.90
strict 80 < 2% 0.85

Optional extras

pip install ragpreflight           # core — fully offline, no API keys
pip install ragpreflight[full]     # adds: semantic chunk coherence (sentence-transformers)
                                   #       near-duplicate detection (datasketch MinHash)
pip install ragpreflight[llm]      # adds: LLM-powered query generation for retrieval sim

The taxonomy: 33 failure modes, 7 pipeline stages

ragpreflight is the first open-source tool grounded in a peer-reviewed RAG failure taxonomy. Each issue it raises is linked to one of 33 named failure modes across 7 pipeline stages:

Stage Modes ragpreflight coverage
Ingestion F1–F4 F1 proxy · F3 direct · F4 risk signal
Representation F5–F6 — (runtime required)
Retrieval F7–F12 F7 direct · F11 proxy
Generation F13–F17 — (requires LLM outputs)
Evaluation F18–F19 —
Deployment F20–F25 F23 risk signal
Agentic Orchestration F26–F33 — (requires agent traces)

"2 direct + 4 proxy/risk + 27 runtime/unsupported" is honest strength, not a weakness. A tool claiming 33/33 detection is lying.

ragpreflight coverage   # see exactly what is and isn't detected

For runtime coverage (hallucination, faithfulness, latency): DeepEval · Ragas · RAGChecker · Phoenix


External validation — olmOCR-bench

ragpreflight scores were validated against olmOCR-bench (Poznanski et al., 2025), a public benchmark of 1,403 PDFs across 7 difficulty categories with known OCR accuracy from state-of-the-art models.

175 PDFs (25 per category, stratified random seed=42) were audited with no ground truth labels, no OCR outputs, and no LLM calls — purely static analysis.

Category ragpreflight mean score olmOCR best accuracy Match
old_scans (historical LoC scans) 45 44.5% ✅ Near-exact
old_scans_math 59.7 75.1% ✅ Hard, scored hard
long_tiny_text 73.0 81.7% ✅ Medium difficulty
multi_column 82.4 79.4% ✅ Layout complexity
headers_footers 84.7 93.4% ✅ Easy, scored easy
arxiv_math 85.2 75.6% ✅ Typeset, extractable
table_tests 86.4 70.2% ⚠ ragpreflight detects text quality; table structure reconstruction requires runtime

Key finding: The hardest category for OCR models (old_scans, 44.5% accuracy) is also the lowest scored by ragpreflight (mean 45). The ordering is consistent across all categories. Static pre-ingestion analysis predicts OCR and RAG difficulty without running a model.

→ View the full HTML audit report — 175 PDFs with per-document scores, issue breakdowns, and Garani 2026 taxonomy links.


Config file

Create .ragpreflight.toml in your project root or ~:

[ragpreflight]
profile = "standard"
max_file_size_mb = 100
output_format = "terminal"
chunk_size = 512
chunk_overlap = 50

Research

This tool implements the taxonomy from:

@inproceedings{garani-2026-systematic,
  title     = {A Systematic Taxonomy of Failure Modes in Retrieval-Augmented Generation Systems},
  author    = {Garani, Anupama},
  booktitle = {Proceedings of the 6th Workshop on Trustworthy Natural Language Processing (TrustNLP 2026)},
  year      = {2026},
  publisher = {Association for Computational Linguistics},
  doi       = {10.18653/v1/2026.trustnlp-main.27},
  url       = {https://aclanthology.org/2026.trustnlp-main.27/}
}

If you use ragpreflight in research, please cite the paper above.


Contributing

See CONTRIBUTING.md. Issues and PRs welcome.


License

MIT — see LICENSE.


Built on the Garani 2026 taxonomy · ACL Anthology · doi:10.18653/v1/2026.trustnlp-main.27

Metadata

Release files for ragpreflight 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ragpreflight 0.1.1
File Size Uploaded
ragpreflight-0.1.1.tar.gz 186.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ragpreflight 0.1.1
File Interpreter ABI Platform
ragpreflight-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 266.0 kB

Release files / ragpreflight-0.1.1.tar.gz

Download URL ragpreflight-0.1.1.tar.gz
Size 186.7 kB
Tags Source
SHA-256 checksum
How to use checksums
ce64c1067320a5eb286c4c7f146d6e1b6c81bc833e8d0f388fb79c34eb241f1d
BLAKE2b-256 checksum
How to use checksums
8df4217a16d5ea0c48269b6e5675431417e360039588b3ce25cc1821395bcbdf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.4

Release files / ragpreflight-0.1.1-py3-none-any.whl

Download URL ragpreflight-0.1.1-py3-none-any.whl
Size 79.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a75b0854210dd59f57fc8bde5efd5206efd5cbe0efbb32cf7d379e3b17e8efba
BLAKE2b-256 checksum
How to use checksums
73ecd402b8ca72ca5ccc38f391e7898fa8f77adb58b8a2a662b4d56e668c2e82
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.4

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page