🩺 rag-doctor
Diagnose why your RAG pipeline returned the wrong answer — in under 2 seconds.
No database. No API keys. No cloud calls. Just pass your documents and get a root cause.
Quick Start · How It Works · Examples · CLI · Docs
The Problem
RAG pipelines fail silently. You get a wrong answer and have no idea if it's a chunking problem, a retrieval miss, a position bias, or a hallucination. Existing evaluation tools give you a score — they don't tell you why.
rag-doctor tells you why.
══════════════════════════════════════════════════════════════
RAG-DOCTOR ✗ ISSUES FOUND
══════════════════════════════════════════════════════════════
Root Cause : context_position_bias (RC-2)
Severity : HIGH
Finding : Best document at position 1/2 — in danger zone (risk: 1.00)
──────────────────────────────────────────────────────────────
Fix: Enable a reranker to push the most relevant document to position 0.
Config Patch: {"retrieval.reranker": true}
══════════════════════════════════════════════════════════════
Quick Start
pip install rag-doctor
from rag_doctor import Doctor
from rag_doctor.connectors.mock import MockConnector
connector = MockConnector(corpus=[
{"id": "doc1", "content": "For liver disease patients maximum acetaminophen dose is 2000mg per day."},
{"id": "doc2", "content": "Standard adult dose: up to 4000mg per day."},
])
docs = connector.retrieve("acetaminophen dose liver disease", top_k=3)
answer = "The maximum daily dose is 4000mg."
report = Doctor.default().diagnose(
query = "What is the max acetaminophen dose for liver disease?",
answer = answer,
docs = docs,
expected = "For liver disease patients max dose is 2000mg per day.",
)
print(report.to_text())
How It Works
rag-doctor runs a deterministic six-tool agent loop. Each tool targets a specific failure mode:
| Tool | Root Cause | What It Catches |
|---|---|---|
RetrievalAuditor |
RC-1 retrieval_miss |
Correct document not in top-k results |
PositionTester |
RC-2 context_position_bias |
Correct doc retrieved but ignored in middle position |
ChunkAnalyzer |
RC-3 chunk_fragmentation |
Mid-sentence truncation, incoherent chunks |
HallucinationTracer |
RC-4 hallucination |
Answer claims not grounded in retrieved documents |
QueryRewriter |
RC-5 query_mismatch |
Query vocabulary doesn't match document vocabulary |
ChunkOptimizer |
RC-3 sub-tool | Grid-searches best chunk_size and strategy |
No Database. No LLM. No API Keys.
rag-doctor builds an ephemeral VectorStore in memory from whatever documents you pass. Your production database is never touched.
Embedding backends (auto-selected, no config needed):
| Priority | Backend | Install |
|---|---|---|
| 1 | sentence-transformers |
pip install sentence-transformers |
| 2 | Ollama nomic-embed-text |
ollama pull nomic-embed-text |
| 3 | TF-IDF (stdlib + numpy) | nothing — built in |
| 4 | Char n-gram fallback | nothing — always available |
Three Ways to Use It
Mode A — Debug from Logs (no re-query needed)
from rag_doctor import Doctor
from rag_doctor.connectors.base import Document
docs = [
Document(content=row["text"], score=row["score"], position=i)
for i, row in enumerate(db_rows)
]
report = Doctor.default().diagnose(
query="What is the refund policy?",
answer="30 days for all customers.",
docs=docs,
expected="Enterprise customers get 90-day refunds.",
)
print(report.root_cause) # retrieval_miss
print(report.fix_suggestion) # Increase top_k or check corpus coverage.
Mode B — Corpus-Level Evaluation (CI / pytest)
connector = MockConnector(corpus=YOUR_CORPUS)
docs = connector.retrieve(query, top_k=5)
report = Doctor.default(connector).diagnose(query=query, answer=answer, docs=docs, expected=expected)
assert report.severity in ("low", "medium"), report.to_text()
Mode C — Connect Your Production Stack
class ChromaConnector(PipelineConnector):
def retrieve(self, query, top_k=5):
results = self.collection.query(query_texts=[query], n_results=top_k)
return [Document(content=d, score=s, position=i) for i,(d,s) in enumerate(...)]
report = Doctor.default(ChromaConnector(my_collection)).diagnose(...)
CLI Reference
# Single query
rag-doctor diagnose \
--query "What is the termination notice?" \
--answer "30 days." \
--expected "Enterprise requires 90 days written notice."
# Batch from JSONL
rag-doctor batch --input examples/batch_example.jsonl --fail-on-severity high
# JSON output for CI
rag-doctor diagnose --query "..." --answer "..." --output json | jq .root_cause
Local Setup (Mac)
git clone https://github.com/your-org/rag-doctor
cd rag-doctor
chmod +x scripts/test_local_mac.sh
./scripts/test_local_mac.sh
Documentation
| Doc | Description |
|---|---|
| docs/user-guide.md | Complete user guide — all 5 journeys, all 6 root causes |
| docs/architecture.md | Internal design: embedding chain, VectorStore, agent loop |
| docs/tools-reference.md | API reference for all 6 tools |
| docs/connectors.md | Building custom connectors |
| docs/configuration.md | Thresholds and config options |
| docs/publishing.md | How to release to PyPI |
License
MIT — free for personal and commercial use.
Metadata
Release files for rag-doctor 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| rag_doctor-1.0.0.tar.gz | 66.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| rag_doctor-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 105.4 kB
Release files / rag_doctor-1.0.0.tar.gz
| Download URL | rag_doctor-1.0.0.tar.gz |
|---|---|
| Size | 66.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cbbcf0e88d8926af66930d71a86ed85893407bfc5264f980f9459b65011c36d4
|
|
BLAKE2b-256 checksum How to use checksums |
8674958598c8b4a681e4070180e0db363f9f772ef812f9ce8c69d5c54c8ad3d9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.5
|
Release files / rag_doctor-1.0.0-py3-none-any.whl
| Download URL | rag_doctor-1.0.0-py3-none-any.whl |
|---|---|
| Size | 39.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2751bdf9b216a3f35e2d5fe4af9f598a849b78d4a4a487e4ef6ed25d89f3a58a
|
|
BLAKE2b-256 checksum How to use checksums |
ec163c7fdacc992169c02b0ee82f7e418586271f70ba35b9ec46170d16dc1ddb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.5
|