ragpreflight
The pre-ingestion RAG audit tool. Catch document failures before you embed them — not after your chatbot starts hallucinating.
pip install ragpreflight
ragpreflight scan my_document.pdf
╭──────────────────────────────────────────────────────────────────────╮
│ ragpreflight — Document Readiness Report │
│ File: quarterly_report.pdf │
│ Score: 44/100 █████████░░░░░░░░░░░ │
│ Format: PDF · Size: 2.1 MB · Pages: 34 · Extractable: 0% │
╰──────────────────────────────────────────────────────────────────────╯
CRITICAL ocr [pages 3, 7, 11] l / I / 1 confusion — 23 matches
Scanner misread lowercase l, capital I, and digit 1 as each other.
Examples: "cIinical" → "clinical", "1evel" → "level"
Keyword search and semantic retrieval both fail on corrupted tokens.
→ Re-run OCR at higher DPI, or apply post-correction (pyspellchecker).
CRITICAL ocr [pages 5, 9] mid-word spaces — 14 matches
Spaces inserted inside words during scan: "pati ent" → "patient"
Each broken word becomes two meaningless tokens in the index.
→ Apply OCR post-correction before ingestion.
CRITICAL ocr [pages 2, 14–17] rn / m split — 8 matches
The letter m was split into rn: "inforrnation" → "information"
Common in low-DPI scans of serif fonts.
→ Re-scan at 300 DPI minimum or run post-OCR cleanup.
CRITICAL content 0% of pages have extractable text. This is a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) before ingestion.
WARNING structure 8 table(s) detected — will chunk as garbled text without special handling.
→ Use a table-aware extractor (pdfplumber, Camelot, LlamaParse).
WARNING content Possible PII: 12 email addresses found on pages 4, 18, 22.
→ Review and redact before ingestion into a shared RAG corpus.
INFO metadata No title, author, or creation date in document metadata.
→ Add metadata to improve retrieval ranking and attribution.
7 issue(s) found · Score: 44 · 4 critical
RAG Failure Taxonomy (doi:10.18653/v1/2026.trustnlp-main.27)
OCR artifacts → F3 Document Quality Failure [direct]
Scanned PDF → F3 Document Quality Failure [direct]
Tables → F7 Structure-Unaware Chunking [risk signal]
PII → F23 PII / Compliance Leak [direct]
No metadata → F11 Low Recall / Ranking Failure [risk signal]
The gap nobody talks about
Every RAG evaluation tool — RAGAS, DeepEval, TruLens, RAGChecker — runs after you build your system. They need a live retriever, real queries, LLM-generated outputs, and an LLM judge to score them. By that point, bad documents are already embedded. You're debugging a production system, not preventing the problem.
ragpreflight runs before you embed anything. Give it a folder of files. It tells you which ones will cause failures and why — in seconds, with no API keys, no internet connection, no running model.
Your documents → [ragpreflight] → fix issues → embed → RAG system
↑
This is the gap.
Nothing else runs here.
It is grounded in peer-reviewed research: 33 failure modes across 7 pipeline stages from Garani 2026 — the first systematic taxonomy of RAG failure modes published at TrustNLP 2026 (ACL). ragpreflight is the first open-source tool that links every detected issue to a named failure mode from that published taxonomy.
How it compares
| ragpreflight | RAGAS | DeepEval | TruLens | RAGChecker | Unstructured | |
|---|---|---|---|---|---|---|
| When it runs | pre-ingestion | post-gen | post-gen | post-gen | post-gen | ingestion |
| Needs a live RAG system | ❌ | ✅ | ✅ | ✅ | ✅ | ❌ |
| Needs an LLM to run | ❌ | ✅ | ✅ | ✅ | ✅ | ⚠️ optional |
| Needs labeled queries / golden sets | ❌ | ⚠️ some metrics | ⚠️ some metrics | ⚠️ some metrics | ✅ | ❌ |
| Fully offline | ✅ | ❌ | ❌ | ❌ | ❌ | ⚠️ OSS only |
| OCR artifact detection | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ |
| Chunking boundary quality | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ |
| Near-duplicate detection | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ |
| PII in source documents | ✅ | ❌ | ⚠️ | ❌ | ❌ | ❌ |
| Staleness detection | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ |
| Taxonomy-grounded (peer-reviewed) | ✅ 33 modes | ❌ | ❌ | ❌ | ❌ | ❌ |
| Multi-format (PDF/DOCX/CSV/HTML/MD…) | ✅ | ❌ | ❌ | ❌ | ❌ | ✅ |
| Readiness score 0–100 | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ |
| HTML audit report | ✅ | ❌ | ⚠️ cloud dashboard | ✅ | ❌ | ❌ |
| CI/CD exit code gating | ✅ | ⚠️ | ✅ | ❌ | ❌ | ❌ |
| SARIF output (GitHub Code Scanning) | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ |
These tools are complementary, not competing. ragpreflight cleans and validates your corpus before ingestion. RAGAS/DeepEval/TruLens evaluate your live RAG system after deployment. Use both.
Quickstart
Scan a single document
ragpreflight scan quarterly_report.pdf
ragpreflight scan contract.pdf --profile strict # higher thresholds
ragpreflight scan report.pdf --json # machine-readable
ragpreflight scan report.pdf --format html --output report.html
Audit your entire knowledge base
ragpreflight audit ./knowledge_base/
ragpreflight audit ./knowledge_base/ --format html --output audit.html
╭──────────────────────────────────────────────────────────────╮
│ ragpreflight — Corpus Readiness Report │
│ Directory: ./knowledge_base/ │
│ Documents: 17 · Avg Score: 77.9 ███████████████░░░░░ │
│ Duplicate Groups: 0 │
╰──────────────────────────────────────────────────────────────╯
Score File Format Issues Critical
──────────────────────────────────────────────────────────
0 empty_file.txt TXT 1 1 ← fix first
44 scanned_old.pdf PDF 5 1 ← needs OCR
85 contract_draft.pdf PDF 2 0
90 data_export.csv CSV 0 0
92 technical_spec.pdf PDF 2 0
94 meeting_notes.docx DOCX 2 0
Gate your ingestion pipeline (CI/CD)
# Fail the pipeline if any document scores below 60
ragpreflight score my_doc.pdf # prints: 92
score=$(ragpreflight score my_doc.pdf)
[ "$score" -ge 60 ] || exit 1
SARIF output for GitHub Code Scanning
ragpreflight scan my_doc.pdf --format sarif --output results.sarif
Upload results.sarif to GitHub Code Scanning and document issues appear as pull request annotations.
Explore the taxonomy
ragpreflight taxonomy list # all 33 failure modes
ragpreflight taxonomy list --stage retrieval # filter by pipeline stage
ragpreflight taxonomy show F7 # full detail for one mode
ragpreflight coverage # what this tool can/cannot detect
Python API
from ragpreflight import scan_document, audit_corpus
# Single document
report = scan_document("my_doc.pdf")
print(report.score) # 83
print(report.issues) # list[Issue] — each with category, severity, suggestion
for issue in report.issues:
print(issue.severity.value, issue.category.value, issue.message)
if issue.taxonomy_refs:
# Each issue linked to Garani 2026 failure modes
for ref in issue.taxonomy_refs:
print(f" → {ref.mode_id} ({ref.relationship})")
# Corpus
corpus = audit_corpus("./knowledge_base/")
print(corpus.average_score) # 77.9
print(corpus.duplicate_groups) # [[file1, file2], ...]
# Programmatic fail gate
critical = [i for d in corpus.documents for i in d.issues if i.severity.value == "critical"]
if critical:
raise ValueError(f"{len(critical)} critical issues — fix before ingesting")
All outputs are typed dataclasses, not dicts. Full type hints. Zero global state.
Supported file formats
| Format | What's checked |
|---|---|
| PDF (text-based) | OCR artifacts, structure, tables, metadata, PII, content density |
| PDF (scanned) | Flags 0% extractability, recommends OCR — score penalised |
| DOCX | Structure, encoding, tables, metadata |
| TXT / Markdown | Encoding, OCR patterns, header hierarchy |
| CSV / TSV | Column consistency, encoding, structural integrity |
| HTML | Boilerplate ratio, structure |
| XLSX | Sheet structure, encoding |
| PPTX | Slide content, structure |
| IPYNB | Cell content, code/text ratio |
| SRT | Transcript length, artifact detection |
Quality profiles
ragpreflight scan doc.pdf --profile permissive # chatbots, internal tools
ragpreflight scan doc.pdf --profile standard # enterprise search (default)
ragpreflight scan doc.pdf --profile strict # medical, legal, financial
| Profile | Min score | OCR tolerance | Similarity threshold |
|---|---|---|---|
permissive |
40 | < 10% | 0.95 |
standard |
60 | < 5% | 0.90 |
strict |
80 | < 2% | 0.85 |
Optional extras
pip install ragpreflight # core — fully offline, no API keys
pip install ragpreflight[full] # adds: semantic chunk coherence (sentence-transformers)
# near-duplicate detection (datasketch MinHash)
pip install ragpreflight[llm] # adds: LLM-powered query generation for retrieval sim
The taxonomy: 33 failure modes, 7 pipeline stages
ragpreflight is the first open-source tool grounded in a peer-reviewed RAG failure taxonomy. Each issue it raises is linked to one of 33 named failure modes across 7 pipeline stages:
| Stage | Modes | ragpreflight coverage |
|---|---|---|
| Ingestion | F1–F4 | F1 proxy · F3 direct · F4 risk signal |
| Representation | F5–F6 | — (runtime required) |
| Retrieval | F7–F12 | F7 direct · F11 proxy |
| Generation | F13–F17 | — (requires LLM outputs) |
| Evaluation | F18–F19 | — |
| Deployment | F20–F25 | F23 risk signal |
| Agentic Orchestration | F26–F33 | — (requires agent traces) |
"2 direct + 4 proxy/risk + 27 runtime/unsupported" is honest strength, not a weakness. A tool claiming 33/33 detection is lying.
ragpreflight coverage # see exactly what is and isn't detected
For runtime coverage (hallucination, faithfulness, latency): DeepEval · Ragas · RAGChecker · Phoenix
External validation — olmOCR-bench
ragpreflight scores were validated against olmOCR-bench (Poznanski et al., 2025), a public benchmark of 1,403 PDFs across 7 difficulty categories with known OCR accuracy from state-of-the-art models.
175 PDFs (25 per category, stratified random seed=42) were audited with no ground truth labels, no OCR outputs, and no LLM calls — purely static analysis.
| Category | ragpreflight mean score | olmOCR best accuracy | Match |
|---|---|---|---|
old_scans (historical LoC scans) |
45 | 44.5% | ✅ Near-exact |
old_scans_math |
59.7 | 75.1% | ✅ Hard, scored hard |
long_tiny_text |
73.0 | 81.7% | ✅ Medium difficulty |
multi_column |
82.4 | 79.4% | ✅ Layout complexity |
headers_footers |
84.7 | 93.4% | ✅ Easy, scored easy |
arxiv_math |
85.2 | 75.6% | ✅ Typeset, extractable |
table_tests |
86.4 | 70.2% | ⚠ ragpreflight detects text quality; table structure reconstruction requires runtime |
Key finding: The hardest category for OCR models (old_scans, 44.5% accuracy) is also the lowest scored by ragpreflight (mean 45). The ordering is consistent across all categories. Static pre-ingestion analysis predicts OCR and RAG difficulty without running a model.
→ View the full HTML audit report — 175 PDFs with per-document scores, issue breakdowns, and Garani 2026 taxonomy links.
Config file
Create .ragpreflight.toml in your project root or ~:
[ragpreflight]
profile = "standard"
max_file_size_mb = 100
output_format = "terminal"
chunk_size = 512
chunk_overlap = 50
Research
This tool implements the taxonomy from:
@inproceedings{garani-2026-systematic,
title = {A Systematic Taxonomy of Failure Modes in Retrieval-Augmented Generation Systems},
author = {Garani, Anupama},
booktitle = {Proceedings of the 6th Workshop on Trustworthy Natural Language Processing (TrustNLP 2026)},
year = {2026},
publisher = {Association for Computational Linguistics},
doi = {10.18653/v1/2026.trustnlp-main.27},
url = {https://aclanthology.org/2026.trustnlp-main.27/}
}
If you use ragpreflight in research, please cite the paper above.
Contributing
See CONTRIBUTING.md. Issues and PRs welcome.
License
MIT — see LICENSE.
Built on the Garani 2026 taxonomy · ACL Anthology · doi:10.18653/v1/2026.trustnlp-main.27
Metadata
Release files for ragpreflight 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ragpreflight-0.1.1.tar.gz | 186.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ragpreflight-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 266.0 kB
Release files / ragpreflight-0.1.1.tar.gz
| Download URL | ragpreflight-0.1.1.tar.gz |
|---|---|
| Size | 186.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ce64c1067320a5eb286c4c7f146d6e1b6c81bc833e8d0f388fb79c34eb241f1d
|
|
BLAKE2b-256 checksum How to use checksums |
8df4217a16d5ea0c48269b6e5675431417e360039588b3ce25cc1821395bcbdf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.4
|
Release files / ragpreflight-0.1.1-py3-none-any.whl
| Download URL | ragpreflight-0.1.1-py3-none-any.whl |
|---|---|
| Size | 79.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a75b0854210dd59f57fc8bde5efd5206efd5cbe0efbb32cf7d379e3b17e8efba
|
|
BLAKE2b-256 checksum How to use checksums |
73ecd402b8ca72ca5ccc38f391e7898fa8f77adb58b8a2a662b4d56e668c2e82
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.4
|