Skip to main content

RAG Debugger

Alpha. Test a RAG fix against your real retriever before you ship it.

pip install -e .
rag-debugger analyze trace.json
rag-debugger repair trace.json --retriever myapp.search:retrieve
rag-debugger check traces/ --retriever myapp.search:retrieve --k 8

The tool loads one failed trace, picks one small retrieval change, calls your retriever, and keeps the change only when the failed claim becomes supported and no previously supported claim gets worse.

judge_supported is that judge's label. It is not ground truth. On a 200-example slice, Llama, Qwen, and Laya all scored below the 0.635 majority baseline. See RESULTS.md.

Repair

Four changes only: increase_k, restore_original_query, rerank_candidates, merge_retrieved.

rag-debugger repair failure.json --retriever myproject.retrieval:search --judge qwen

--judge is llama, qwen, openai, overlap, local, or module:function. With no flag, OpenAI is used when OPENAI_API_KEY is set. Otherwise the lexical overlap judge runs, so the command works with no key and no GPU.

A run that should be kept looks like this:

RAG Debugger
Judge: overlap

Failed claim
"Paris is the capital of France."

Diagnosis
Likely retrieval failure

Experiment
increase_k: 1 → 6

Before
unsupported

After
supported

Regression check
0 previously supported claims checked
0 regressions

judge_supported=true

Recommendation
ACCEPT EXPERIMENT

A run that should be thrown away looks like this. The failed claim may improve, and a claim that was already supported gets worse:

Recommendation
REJECT EXPERIMENT

Try that case:

rag-debugger repair examples/traces/accept_increase_k.json \
  --retriever examples.demo_retriever:retrieve --judge overlap

rag-debugger repair examples/traces/reject_increase_k.json \
  --retriever examples.demo_retriever:retrieve --judge overlap

The reject trace already supports “The desk lamp uses 40 watts.” Wider k returns Paris and drops the lamp, so the recommendation is REJECT EXPERIMENT.

Check a change across many traces

Before you change k, the chunk size, the embedding model, or the reranker, run the new retriever over your logged traces:

rag-debugger check logged_traces/ --retriever myapp.search:retrieve_v2 --judge overlap
rag-debugger check logged.jsonl --retriever myapp.search:retrieve --k 8 --out report.json
Traces checked: 26
Fixed:     17 / 19 failing traces
Broken:    4 / 7 working traces
Mixed:     1 traces fixed one claim and broke another
Unchanged: 4

Recommendation
REVIEW BEFORE SHIPPING

The logged retrieved_chunks are the "before". The retriever you pass is the "after". Each broken trace is listed by name with the claim that lost its support. --k overrides every trace's logged top_k.

Wikipedia example

python -m examples.wiki_demo logs BM25 top-1 over 2,067 SQuAD Wikipedia passages (150 questions), then checks two changes. The gold line is whether the dataset answer string is in the retrieved text. It is not the sealed study.

Replacing BM25 with all-MiniLM-L6-v2 at the same k:

Fixed Broken
Gold answer string 13 / 33 missed 33 / 117 already answered
Overlap judge 12 / 31 failing traces 31 / 119 working traces

Packing BM25 top-6 into a 90-word budget fixed 6 and broke 93 of the 117 answers BM25 already had. The tool's recommendation on both changes is REVIEW BEFORE SHIPPING.

Your retriever

def retrieve(query: str, k: int):
    return [{"id": "1", "text": "..."}]

LangChain:

from rag_debugger.integrations.langchain import as_retriever, trace

failure = trace("What does the desk lamp use?", chain_result, top_k=1)
failure.save("failure.json")
# rag-debugger repair failure.json --retriever myapp:retrieve
# or, in Python, pass as_retriever(vectorstore.as_retriever()) to run_verified_repair

as_retriever sets k and returns {id, text} chunks.

Trace

{
  "question": "...",
  "answer": "...",
  "retrieved_chunks": [{"id": "r1", "text": "..."}],
  "corpus_chunks": [{"id": "c47", "text": "..."}],
  "metadata": {"retriever": "faiss", "top_k": 4}
}

Frozen results

The sealed holdout and the mixed Hotpot comparison are frozen. Diagnosis did not beat blind increase-k: 66/76 versus 67/76, and 11/35 versus 12/35. On 12 controls, blind increase-k regressed 4 working answers. The guarded path regressed 0. Do not retune those sets.

On the SQuAD development half, selective repair also stays behind increase-k at the 5% regression budget (overlap 354 vs 382, Llama 333 vs 382, Qwen 274 vs 382, Mistral 247 vs 382). The guard matched no-guard on all four. The final test is not scored. Details are in RESULTS.md.

rag-debugger with no subcommand prints help. rag-debugger serve is the older demo server.

Release files for ragfix 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ragfix 0.2.0
File Size Uploaded
ragfix-0.2.0.tar.gz 78.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ragfix 0.2.0
File Interpreter ABI Platform
ragfix-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 171.5 kB

Release files / ragfix-0.2.0.tar.gz

Download URL ragfix-0.2.0.tar.gz
Size 78.0 kB
Tags Source
SHA-256 checksum
How to use checksums
c7047d9b58beea08880be08c7d1dcf50b931759c629a13d7d5e446e585a92f5d
BLAKE2b-256 checksum
How to use checksums
9ea8846eb02e3886e3a5b006e87235269b117a6ef18ab3cb087ec2eac46876c1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / ragfix-0.2.0-py3-none-any.whl

Download URL ragfix-0.2.0-py3-none-any.whl
Size 93.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
66cf82ee8c29b7901d2c6977745ef7c93bb88e1c98c5b73ebd8001a8276c7b0d
BLAKE2b-256 checksum
How to use checksums
5b574bcad46f9e9e11cf3825c94f7f1396d391635b6c05a401d43c77adf0a30c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

0.2.2

2 release files

0.2.1

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page