Toolkit RAG Quality
Retrieval regression testing in CI. toolkit-rag scores a retriever's ranked results with metrics that match trec_eval exactly, tells you which queries got worse and whether the change is statistically significant, fails the build when it regresses beyond budget, and checks that eval queries and answers have not leaked into the corpus. It is deterministic, makes no model calls and has no runtime dependencies (LangChain and LlamaIndex adapters are optional extras).
It complements, rather than replaces, two good tools: Ragas judges generation quality (faithfulness, answer relevance) with LLMs, and ranx is a research library with many more IR metrics, fusion and multi-system statistics. toolkit-rag is the CI piece: a gate with a stable report format, per-query diffs and a GitHub Action.
Capabilities
| Capability | Status | Notes |
|---|---|---|
Retrieval metrics at k: hit rate, recall, precision, MRR, nDCG, MAP (score) |
Working | Equal to trec_eval: hard-coded pytrec_eval reference values in tests/test_trec_reference.py and tests/test_graded_reference.py; on BEIR SciFact all six metrics match trec_eval to 12 decimals. |
Graded relevance and --relevance-level |
Working | Integer grades; nDCG uses linear gains, as trec_eval does. Checked at relevance levels 1 and 2. |
TREC qrels/run and BEIR qrels import, TREC and JSONL export (score, convert) |
Working | trec_eval tie-breaking for runs; formats auto-detected. |
Regression gate (compare) |
Working | Per-metric budgets for all six metrics; fails when the reports use a different k, relevance level or query set. |
Per-query diff (compare) |
Working | Which queries got worse or better, per metric, with the most regressed listed. |
Paired significance tests (compare --alpha) |
Working | Randomization test (exact up to 16 queries) and paired t-test, checked against SciPy. |
GitHub Action (action.yml) |
Working | Composite action for the regression gate, with a step summary; tested end to end in CI. |
Eval leakage check (leakage) |
Working | Exact containment of eval queries/answers in corpus documents; exits 4 above --max-leaks. |
Near-duplicate documents (overlap --method minhash) |
Working | MinHash-LSH candidates verified with exact Jaccard over word shingles; the miss probability at the threshold is reported. Checked against exhaustive exact Jaccard. Pure Python: a 5,000 x 5,000 document comparison takes about a minute. |
Exact-duplicate overlap between corpora (overlap) |
Working | Exact match after lowercasing and whitespace collapsing, via SHA-256 fingerprints. |
Retriever adapters (run-retriever) |
Working | LangChain and LlamaIndex retrievers (optional extras) or any callable, to a TREC or JSONL run. |
| Report envelope (default JSON output) | Working | Every report is an in-toto Statement v1 in canonical JSON, with SHA-256 digests of the inputs. See Report format. |
Report validation (validate-report) |
Working | Checks an envelope against the v1 rules (including verdict/exit-code consistency) and, for rag.score, that every headline metric is present. |
| Semantic (embedding) similarity for paraphrase leakage | Planned | Not implemented: leakage finds copied text, not reworded text. |
| Generation-side metrics (faithfulness, answer relevance) | Not planned | Use Ragas or a similar LLM-judged tool. |
Install
The package is not published on PyPI yet. Install from source:
git clone https://github.com/AKIVA-AI/toolkit-rag-quality.git
cd toolkit-rag-quality
pip install . # or: pip install -e ".[dev]" for development
toolkit-rag --version
Python 3.10 or newer is required. Optional extras: .[langchain], .[llamaindex].
5-minute example: BEIR SciFact
This compares two configurations of a small BM25 retriever on the BEIR SciFact test set (300 queries, 5,183 abstracts, 2.8 MB download). The retriever is examples/bm25.py, a pure-Python BM25 kept small for the example. The candidate drops document titles from the index, a plausible-looking change. Run from the repository root:
curl -LO https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/scifact.zip
unzip -q scifact.zip
# 1. Produce a run file for each configuration (about 2 seconds each).
toolkit-rag run-retriever --retriever "examples.bm25:baseline()" --queries scifact/queries.jsonl \
--qrels scifact/qrels/test.tsv --k 100 --out baseline.trec --tag bm25-title
toolkit-rag run-retriever --retriever "examples.bm25:candidate()" --queries scifact/queries.jsonl \
--qrels scifact/qrels/test.tsv --k 100 --out candidate.trec --tag bm25-notitle
# 2. Score both at k = 10 against the BEIR qrels.
toolkit-rag score --queries scifact/qrels/test.tsv --retrieved baseline.trec --k 10 --out baseline.json --format table
toolkit-rag score --queries scifact/qrels/test.tsv --retrieved candidate.trec --k 10 --out candidate.json --format table
# 3. Gate the candidate: 2% budget, fail only on significant drops.
toolkit-rag compare --baseline baseline.json --candidate candidate.json --alpha 0.05 --format markdown
# 4. Check whether any SciFact claim is copied into the corpus.
toolkit-rag leakage --corpus scifact/corpus.jsonl --items scifact/queries.jsonl --field text --format table
Step 2 gives the baseline nDCG@10 = 0.6598 (the BEIR paper reports 0.665 for its Elasticsearch BM25 baseline). Step 3 exits 4 and prints:
| metric | baseline | candidate | change % | budget % | p-value | worse / better | result |
|---|---|---|---|---|---|---|---|
| recall_at_k | 0.7768 | 0.7629 | -1.79 | 2.00 | 0.1014 | 7 / 1 | pass |
| precision_at_k | 0.0853 | 0.0833 | -2.34 | 2.00 | 0.0762 | 7 / 1 | pass (not significant) |
| ndcg_at_k | 0.6598 | 0.6427 | -2.59 | 2.00 | 0.0056 | 30 / 16 | FAIL |
| mrr_at_k | 0.6295 | 0.6128 | -2.64 | 2.00 | 0.0186 | 26 / 14 | FAIL |
| map_at_k | 0.6169 | 0.5984 | -2.99 | 2.00 | 0.0072 | 30 / 16 | FAIL |
| hit_rate_at_k | 0.8000 | 0.7867 | -1.67 | 2.00 | 0.2246 | 5 / 1 | pass |
followed by the most regressed queries (query 1194 drops from nDCG@10 = 1.0 to 0). Dropping titles costs 2.6% nDCG@10, and 30 queries got worse against 16 that improved, which the randomization test finds significant (p = 0.006). Precision also dropped past its budget, but not significantly, so --alpha lets it pass. Step 4 exits 4: 18 of the 1,109 SciFact claims (train and test) share at least 60% of their word 3-grams with one abstract, two of them verbatim. That is expected for a fact-checking dataset whose claims were written from those abstracts, and it shows what the check reports; for your own eval set, a leak usually means the question was copied into the knowledge base.
The numbers above were produced with this repository's code in a clean Python 3.12 container.
Usage
Score retrieval results:
toolkit-rag score --queries queries.jsonl --retrieved retrieved.jsonl --k 5 --out report.json
Compare a candidate report to a baseline (CI gating):
toolkit-rag compare --baseline baseline.json --candidate report.json \
--max-recall-regression-pct 2.0 --max-regression-pct 5.0 \
--budget ndcg=1 --budget hit_rate=off --alpha 0.05 --format markdown
What the gate checks:
- Budgets. Each metric's mean may drop by at most its budget, in percent of the baseline.
--max-recall-regression-pct(default 2) covers recall;--max-regression-pctcovers the other five metrics and defaults to the recall budget.--budget METRIC=PCT(repeatable) overrides either for one metric, and--budget METRIC=offreports a metric without gating it. Metric names:recall,precision,ndcg,mrr,map,hit_rate(or the*_at_ksummary keys). - Comparability. The gate fails if the reports use a different
k, a different relevance level, or a different query set. - Per-query diff. For every query and metric, the report lists baseline, candidate and delta (
details.per_query_diff), counts the queries that got worse, better or stayed the same (summary.per_query), and names the most regressed queries by--primary-metric(default nDCG;--top, default 10). - Significance. Each metric gets a two-sided paired test on its per-query scores: a randomization (sign-flip) test by default, or
--test t-test. With 16 queries or fewer, the randomization test enumerates every sign pattern (exact); above that it draws--permutations(default 10,000) patterns from a fixed--seed, so results are reproducible. Both tests are checked against SciPy (scipy.stats.permutation_testandscipy.stats.ttest_rel) intests/test_stats.py. --alpha. Without it, any drop over budget fails. With it, a metric over budget fails only if its drop is also significant (p < alpha), so noise on a small query set does not break the build. With fewer than two queries there is no p-value, and an over-budget drop fails.--alphaneeds per-query rows in both reports.
--format markdown prints a table of metrics (change, budget, p-value, worse/better counts) and the most regressed queries, suitable for a CI step summary.
Produce a run file from your own retriever (LangChain, LlamaIndex, or any Python callable):
pip install ".[langchain]" # or ".[llamaindex]"; a plain callable needs neither
toolkit-rag run-retriever --retriever my_pipeline:build_retriever() \
--queries queries.jsonl --qrels qrels.trec --k 100 --out run.trec --tag my-retriever
--retriever module:attrnames a retriever object;module:factory()calls a zero-argument factory. The current directory is on the import path. This imports and runs that code, so only point it at code you trust.- LangChain retrievers are called with
invoke(query), LlamaIndex retrievers withretrieve(query), and a callable withf(query)(returning doc ids). The framework is detected from the class, or set with--framework. - Document ids come from
Document.id(LangChain) ornode.node_id(LlamaIndex), or from a metadata field with--id-key doc_id. A result without an id is an error. - Each ranking is de-duplicated and cut to
--k.--qrelslimits the run to judged queries.--to jsonlwrites a JSONL run instead of TREC. - The adapters are tested against real
langchain-coreandllama-index-corein CI (tests/test_adapters.py).
Find duplicate documents between two corpora (for example, to check test/train contamination):
toolkit-rag overlap --a corpus_a.jsonl --b corpus_b.jsonl --out overlap.json # exact
toolkit-rag overlap --a corpus_a.jsonl --b corpus_b.jsonl --method minhash --threshold 0.8 # near-duplicate
--method exact(default) matches documents whose text is identical after lowercasing and collapsing whitespace.--method minhashmatches documents whose word 5-gram shingles (--shingle) have a Jaccard similarity of at least--threshold(default 0.8, the setting used for training-data deduplication by Lee et al., 2022, "Deduplicating Training Data Makes Language Models Better"). Text is normalized (Unicode NFKC, lowercase, punctuation dropped) first. MinHash-LSH only proposes candidate pairs; every reported pair is verified with the exact Jaccard similarity, so there are no false positives. A pair can be missed: the LSH banding is chosen so a pair exactly at the threshold is found with probability at least 99.5%, and the report states the actual miss probability (miss_probability_at_threshold). The pairs are indetails.pairs.
Check whether eval queries or answers appear in the corpus:
toolkit-rag leakage --corpus corpus.jsonl --items eval.jsonl --field query --field answer --out leakage.json
leakage measures containment: the share of an eval text's word 3-grams that also occur in one corpus document, computed exactly with an inverted index. Jaccard similarity is the wrong measure here: a 12-word question copied verbatim into a 300-word document has a Jaccard similarity of about 0.03 but a containment of 1.0. An item leaks when any of its fields reaches --threshold (default 0.6). What that default catches, from tests/test_leakage.py, for a 12-word question:
| Eval text vs the corpus document | Containment | Flagged at 0.6 |
|---|---|---|
| Verbatim copy, or a case/punctuation variant | 1.0 | yes |
| One word substituted | 0.7 | yes |
| Two words substituted, far apart | 0.5 | no |
| Unrelated question | 0.0 | no |
Lower the threshold to catch paraphrase-level copies (and expect more false positives on templated text); raise it to flag only near-verbatim copies. Paraphrases with different wording are not detected: that needs embedding similarity, which is not implemented. leakage exits 4 when more items leak than --max-leaks (default 0). Items and corpus rows may use BEIR field names (_id, title, text).
A corpus or item file with more rows than --max-records (default 50,000) is rejected, not truncated.
Check a report's shape:
toolkit-rag validate-report --report report.json
Every command prints its report as JSON by default; --format table or --format markdown prints a human-readable summary instead. --out always writes the JSON report. Global flags: --verbose and --log-format text|json.
GitHub Action
action.yml at the repository root is a composite action that runs the regression gate and writes the Markdown summary to the job's step summary. Produce the two score reports in earlier steps (for example, the baseline from main and the candidate from the pull request), then:
- uses: AKIVA-AI/toolkit-rag-quality@<commit-sha> # pin a commit until a release tag exists
with:
baseline: reports/baseline.json
candidate: reports/candidate.json
max-recall-regression-pct: "2.0"
budgets: |
ndcg=1
hit_rate=off
alpha: "0.05" # optional: fail only on significant drops
| Input | Default | Meaning |
|---|---|---|
baseline, candidate |
required | Score reports from toolkit-rag score --out |
max-recall-regression-pct |
2.0 |
Recall budget, in percent of the baseline |
max-regression-pct |
recall budget | Budget for the other metrics |
budgets |
none | One METRIC=PCT or METRIC=off per line |
alpha |
none | Significance level for --alpha |
test |
permutation |
permutation or t-test |
report |
rag-compare.json |
Where the compare report is written |
fail-on-regression |
true |
Fail the step when the gate does not pass |
python-version |
3.12 |
Empty string: use the runner's Python |
Outputs: verdict (pass, fail or error), exit-code and report. The step fails when the gate does not pass. GitHub drops a composite action's outputs when the action fails, so to read verdict in a later step set fail-on-regression: "false" and decide there. The action installs the package from the action's own checkout, so the gate always matches the pinned version. CI runs the action end to end on fixtures (the action job in .github/workflows/ci.yml).
With few queries a real drop may not reach significance: in that CI job, 4 of 8 queries losing their only relevant document gives p = 0.125, so alpha: 0.05 lets it pass. Use alpha to absorb noise on large query sets, not to gate small ones.
Metric definitions
All metrics use the cutoff k and follow trec_eval (P_k, recall_k, ndcg_cut_k, map_cut_k, success_k, and recip_rank over the top k). Judgments can be binary or graded (integer grades). A document is relevant when its grade is at least the relevance level (--relevance-level, default 1, like trec_eval -l).
- precision@k = relevant documents in the top k, divided by k (not by the number of documents returned).
- recall@k = relevant documents in the top k, divided by the number of relevant documents.
- nDCG@k uses the grade as a linear gain (grade 1 for
relevant_ids) and the discount1/log2(rank + 1). The ideal DCG sorts all positive grades in the judgments and keeps the top k. As in trec_eval, nDCG ignores the relevance level, and zero or negative grades give no gain. - MAP@k is the sum of precision at each relevant rank within k, divided by the number of relevant documents.
- MRR@k is the reciprocal rank of the first relevant document within k, or 0.
- hit rate@k is 1 if any relevant document appears in the top k.
These definitions are checked against values computed with trec_eval (through pytrec_eval): binary cases in tests/test_trec_reference.py, graded cases at relevance levels 1 and 2 in tests/test_graded_reference.py.
Input handling:
- Repeated ids in a retrieved list are removed before the cutoff; the first occurrence wins.
- A query with no judgments at all (for example an empty
relevant_ids) is unjudged. It is excluded from the averages, counted inunjudged_queries, and logged. - A judged query with no document at or above the relevance level is scored (0 on everything except possibly nDCG) and averaged in, as trec_eval does. It is counted in
queries_without_relevant. - A judged query with no retrieved row scores 0 on every metric and is counted in
queries_without_results. - Run entries for queries that have no judgments are ignored and counted in
unscored_run_queries. - Rows without an
idare counted (skipped_query_rows,skipped_retrieved_rows) and logged. A repeated query id, a row with bothrelevant_idsandrelevance, or a non-integer grade is an error. kand--relevance-levelmust be positive integers. Scoring with no judged queries is an error.
Data formats
score --queries takes relevance judgments and score --retrieved takes ranked results. The format is detected from the first line, or set with --queries-format jsonl|trec|beir and --retrieved-format jsonl|trec.
Queries JSONL (one object per line), binary or graded:
{"id":"q1","query":"...","relevant_ids":["doc-1","doc-9"]}
{"id":"q2","relevance":{"doc-3":2,"doc-4":1,"doc-7":0}}
Retrieved JSONL, ranked best first:
{"id":"q1","retrieved_ids":["doc-9","doc-2","doc-1"]}
TREC qrels (qid iter docid grade) and TREC runs (qid Q0 docid rank score tag). As in trec_eval, a run is ranked by score (highest first) with ties broken by doc id in descending order; the rank column is ignored. A document listed twice for one query is an error.
BEIR qrels: the tab-separated qrels/<split>.tsv file of a BEIR dataset (query-id, corpus-id, score header).
Convert between formats:
toolkit-rag convert qrels --in scifact/qrels/test.tsv --to trec --out test.qrels
toolkit-rag convert run --in run.trec --to jsonl --out run.jsonl
toolkit-rag convert run --in run.jsonl --to trec --tag my-retriever --out run.trec
A JSONL run carries ranks only, so a TREC export writes the score len(ranking) - rank + 1: strictly decreasing, so trec_eval reproduces the list order.
Corpora JSONL:
{"id":"doc-1","text":"..."}
Report format
score, compare, overlap and leakage write their result as a report envelope: an in-toto Statement v1, the attestation format used by SLSA and Sigstore. The full convention is in docs/report-envelope.md and the JSON Schema is schemas/report-envelope.v1.json.
{
"_type": "https://in-toto.io/Statement/v1",
"subject": [{"name": "run.jsonl", "digest": {"sha256": "..."}}],
"predicateType": "https://github.com/AKIVA-AI/toolkit-rag-quality/report/v1",
"predicate": {
"tool": {"name": "toolkit-rag-quality", "version": "..."},
"kind": "rag.score",
"created_at": "2026-09-26T18:00:00Z",
"verdict": "pass",
"exit_code": 0,
"inputs": [{"name": "queries.jsonl", "digest": {"sha256": "..."}}, {"name": "run.jsonl", "digest": {"sha256": "..."}}],
"summary": {"k": 5, "queries": 300, "recall_at_k": 0.71, "...": "..."},
"details": {"per_query": [{"id": "q1", "recall": 1.0, "...": "..."}]}
}
}
- The file is canonical JSON (sorted keys, no insignificant whitespace, trailing newline), so its SHA-256 is stable. Set
SOURCE_DATE_EPOCHto pincreated_atand get byte-identical reports from the same inputs. verdictfollows the exit code:pass(0),fail(4),error(2 or 3). When an input is invalid and--outis given, anerrorreport is written; when an input file is missing, no report is written.--legacy-jsonemits the pre-1.0 shape (schema_version,summary,per_query) for one more minor version.comparereads both shapes.
kind |
subject |
predicate.summary |
predicate.details |
|---|---|---|---|
rag.score |
the run file | k, relevance_level, queries, unjudged_queries, queries_without_results, queries_without_relevant, unscored_run_queries, skipped_query_rows, skipped_retrieved_rows, and hit_rate_at_k, recall_at_k, precision_at_k, mrr_at_k, ndcg_at_k, map_at_k |
per_query: one row per judged query with recall, precision, mrr, ndcg, ap, hit |
rag.compare |
the candidate report | passed, reason, failed_metrics, alpha, test, queries_compared, per_query.<metric> (worse, better, unchanged), most_regressed, and metrics.<name> with baseline, candidate, regression_pct, max_regression_pct, gated, over_budget, p_value, significant, passed |
per_query_diff: per query and metric, baseline, candidate, delta |
rag.overlap (exact) |
both corpora | a_docs, b_docs, overlap_docs, overlap_rate, skipped-row counts, match |
empty |
rag.overlap (minhash) |
both corpora | method, a_docs, b_docs, pairs, a_docs_with_match, overlap_rate, threshold, shingle, num_perm, bands, rows, candidates_checked, miss_probability_at_threshold |
pairs: a_id, b_id, jaccard |
rag.leakage |
the eval items | items, texts_checked, corpus_docs, leaked_items, leaked_texts, threshold, shingle, fields, max_leaks, skipped counts |
leaks: item_id, field, doc_id, containment, shared_shingles, item_shingles |
| any, on error | the input files that exist | error: the message |
empty |
To sign a report, use the optional toolkit-ml-provenance CLI (Ed25519 key, or Sigstore keyless with its sigstore extra):
toolkit-mlsbom sign-file report.json # then: toolkit-mlsbom verify-file report.json
Exit codes
| Code | Meaning |
|---|---|
0 |
Success, or compare passed |
2 |
CLI or input error (bad arguments, unreadable file, invalid input) |
3 |
Unexpected error |
4 |
compare failed its budget or leakage found more leaks than allowed (verdict: fail), or validate-report found an invalid report |
Development
pip install -e ".[dev]" # add ,langchain,llamaindex to run the adapter tests
pytest -q
ruff check .
pyright src/
See CONTRIBUTING.md and CHANGELOG.md.
Contributing and security
Contributions are welcome: see CONTRIBUTING.md and the Code of Conduct. Please report security problems privately, as described in SECURITY.md.
Releasing
Releases are cut by pushing a vX.Y.Z tag. CI runs the tests, builds the
sdist and wheel, checks them, attaches them to a GitHub Release and publishes
them to PyPI with Trusted Publishing. RELEASING.md describes
the process and how to verify a release.
License
Apache License 2.0. See LICENSE and NOTICE.
Versions 0.2.0 and earlier were released under the MIT License and remain available under it.
Metadata
Release files for toolkit-rag-quality 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| toolkit_rag_quality-1.0.0.tar.gz | 80.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| toolkit_rag_quality-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 132.3 kB
Release files / toolkit_rag_quality-1.0.0.tar.gz
| Download URL | toolkit_rag_quality-1.0.0.tar.gz |
|---|---|
| Size | 80.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5df7093c80f92ff6ffa098b0744db10ee76529522d3dacbdcf5145185afd9d04
|
|
BLAKE2b-256 checksum How to use checksums |
ccf940f5c1a8b8d0872bfd22904744197e1ff61562ebf874cb123865ade56442
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / toolkit_rag_quality-1.0.0-py3-none-any.whl
| Download URL | toolkit_rag_quality-1.0.0-py3-none-any.whl |
|---|---|
| Size | 52.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8c6d94e3c27475fa6c435adb1661b9e76c483d9557f67753d5741426ce723b87
|
|
BLAKE2b-256 checksum How to use checksums |
18917bac636a5dd492edd8da932f7ea75334cd3a84ef8bc4ce09186f04107042
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log