Skip to main content

RAG Debugger

Alpha. Test a RAG fix against your real retriever before you ship it.

pip install ragfix
ragfix analyze trace.json
ragfix repair trace.json --retriever myapp.search:retrieve
ragfix check traces/ --retriever myapp.search:retrieve --k 8

Package: https://pypi.org/project/ragfix/

The tool loads one failed trace, picks one small retrieval change, calls your retriever, and keeps the change only when the failed claim becomes supported and no previously supported claim gets worse.

judge_supported is that judge's label. It is not ground truth. On a 200-example slice, Llama, Qwen, and Laya all scored below the 0.635 majority baseline. See RESULTS.md.

Repair

Four changes only: increase_k, restore_original_query, rerank_candidates, merge_retrieved.

ragfix repair failure.json --retriever myproject.retrieval:search --judge qwen

--judge is llama, qwen, openai, overlap, local, or module:function. With no flag, OpenAI is used when OPENAI_API_KEY is set. Otherwise the lexical overlap judge runs, so the command works with no key and no GPU.

A run that should be kept looks like this:

RAG Debugger
Judge: overlap

Failed claim
"Paris is the capital of France."

Diagnosis
Likely retrieval failure

Experiment
increase_k: 1 → 6

Before
unsupported

After
supported

Regression check
0 previously supported claims checked
0 regressions

judge_supported=true

Recommendation
ACCEPT EXPERIMENT

A run that should be thrown away looks like this. The failed claim may improve, and a claim that was already supported gets worse:

Recommendation
REJECT EXPERIMENT

Try that case:

ragfix repair examples/traces/accept_increase_k.json \
  --retriever examples.demo_retriever:retrieve --judge overlap

ragfix repair examples/traces/reject_increase_k.json \
  --retriever examples.demo_retriever:retrieve --judge overlap

The reject trace already supports “The desk lamp uses 40 watts.” Wider k returns Paris and drops the lamp, so the recommendation is REJECT EXPERIMENT.

Check a change across many traces

Before you change k, the chunk size, the embedding model, or the reranker, run the new retriever over your logged traces:

ragfix check logged_traces/ --retriever myapp.search:retrieve_v2 --judge overlap
ragfix check logged.jsonl --retriever myapp.search:retrieve --k 8 --out report.json
Traces checked: 26
Fixed:     17 / 19 failing traces
Broken:    4 / 7 working traces
Mixed:     1 traces fixed one claim and broke another
Unchanged: 4

Recommendation
REVIEW BEFORE SHIPPING

The logged retrieved_chunks are the "before". The retriever you pass is the "after". Each broken trace is listed by name with the claim that lost its support. --k overrides every trace's logged top_k.

Wikipedia example

python -m examples.wiki_demo logs BM25 top-1 over 2,067 SQuAD Wikipedia passages (150 questions), then checks two changes. The gold line is whether the dataset answer string is in the retrieved text. It is not the sealed study.

Replacing BM25 with all-MiniLM-L6-v2 at the same k:

Fixed Broken
Gold answer string 13 / 33 missed 33 / 117 already answered
Overlap judge 12 / 31 failing traces 31 / 119 working traces

Packing BM25 top-6 into a 90-word budget fixed 6 and broke 93 of the 117 answers BM25 already had. The tool's recommendation on both changes is REVIEW BEFORE SHIPPING.

Your retriever

def retrieve(query: str, k: int):
    return [{"id": "1", "text": "..."}]

LangChain:

from ragfix.integrations.langchain import as_retriever, trace

failure = trace("What does the desk lamp use?", chain_result, top_k=1)
failure.save("failure.json")
# ragfix repair failure.json --retriever myapp:retrieve
# or, in Python, pass as_retriever(vectorstore.as_retriever()) to run_verified_repair

as_retriever sets k and returns {id, text} chunks.

Trace

{
  "question": "...",
  "answer": "...",
  "retrieved_chunks": [{"id": "r1", "text": "..."}],
  "corpus_chunks": [{"id": "c47", "text": "..."}],
  "metadata": {"retriever": "faiss", "top_k": 4}
}

Frozen results

The sealed holdout and the mixed Hotpot comparison are frozen. Diagnosis did not beat blind increase-k: 66/76 versus 67/76, and 11/35 versus 12/35. On 12 controls, blind increase-k regressed 4 working answers. The guarded path regressed 0. Do not retune those sets.

On the SQuAD development half, selective repair also stays behind increase-k at the 5% regression budget (overlap 354 vs 382, Llama 333 vs 382, Qwen 274 vs 382, Mistral 247 vs 382). The guard matched no-guard on all four. The final test is not scored. Details are in RESULTS.md.

ragfix with no subcommand prints help. ragfix serve is the older demo server.

Release files for ragfix 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ragfix 0.2.1
File Size Uploaded
ragfix-0.2.1.tar.gz 78.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ragfix 0.2.1
File Interpreter ABI Platform
ragfix-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 171.7 kB

Release files / ragfix-0.2.1.tar.gz

Download URL ragfix-0.2.1.tar.gz
Size 78.2 kB
Tags Source
SHA-256 checksum
How to use checksums
f307c2ebe6d5932b97c462eb542a18bf49e8e3e0e419f83bd87236d7569324b4
BLAKE2b-256 checksum
How to use checksums
ea604aea1535c6d10afcf66ec0dc524f47336832dbe5e5a10ff47d4a0e0265d1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / ragfix-0.2.1-py3-none-any.whl

Download URL ragfix-0.2.1-py3-none-any.whl
Size 93.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ea7c963b615d06bb94a360d20e0012191c8dfd8c447a3d5cd09685851d0479a2
BLAKE2b-256 checksum
How to use checksums
e4d7547a78a1e3c537a480d30152768d6ca6a8cb57de019aeec4d9f37ba906e3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

0.2.2

2 release files

This release

0.2.1 This release

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page