GASP
Grounding-Aware Sensitivity by Perturbation — a span-level detector of ungrounded content in retrieval-augmented generation (RAG).
GASP scores each answer sentence by its grounding sensitivity: the change in the sentence's likelihood when the retrieved context is perturbed. A grounded sentence loses much of its likelihood when its supporting passage is removed; an unsupported sentence barely reacts. GASP needs only a probabilistic scorer, no trained verifier and no labeled data, and it returns, for each sentence, the chunk that best supports it.
Install
pip install gasp-rag # core
pip install gasp-rag[torch] # with PyTorch and transformers, needed to run a scorer
Quickstart
from gasp import GASP
detector = GASP("Qwen/Qwen2.5-1.5B-Instruct", k_chunks=5)
context = "..." # the retrieved passages, as one string
answer = "..." # the generated answer to check
question = "..." # the query (optional)
result = detector.detect(context=context, answer=answer, query=question)
for s in result:
print(f"{s.sensitivity:+.2f} {s.text}")
if s.supporting_chunk:
print(f" supported by: {s.supporting_chunk[:80]}...")
Higher sensitivity means the sentence depends more on the retrieved evidence and is more likely grounded. Lower sensitivity means it barely reacts to removing evidence and is more likely unsupported. To flag sentences, pass a threshold:
result = detector.detect(context, answer, threshold=0.5)
for s in result.flagged():
print("likely unsupported:", s.text)
Thresholds are corpus dependent and are best calibrated on held-out data; the continuous
sensitivity score is the primary output.
Options
Everything is configurable on the detector:
GASP(
model_id, # any Hugging Face causal LM, small or large, CPU or GPU
k_chunks=5, # number of context chunks
threshold=None, # flag sentences below this sensitivity
economical=False, # two-pass variant: faster, no attribution
sensitivity_feature="max_drop", # or "gap", "mean_drop", "top2_drop", "max_jsd"
max_ctx_tokens=1800, # context truncation
max_ans_tokens=256, # answer truncation
device=None, # "cpu" or "cuda"
dtype=None, # "float16", "bfloat16", "float32"
)
Change the scorer at any time by constructing a new detector with a different model_id.
Many answers and files
Score a list of answers, or read a .jsonl/.csv file:
items = [
{"context": ctx1, "answer": ans1, "query": q1},
{"context": ctx2, "answer": ans2},
]
detections = detector.detect_batch(items, threshold=0.5)
# results as plain dicts, ready for pandas or JSON
import pandas as pd
rows = [r for d in detections for r in d.to_records()]
df = pd.DataFrame(rows)
print(detections[0].summary()) # {'n_sentences': ..., 'n_flagged': ..., 'mean_sensitivity': ...}
Command line
No code needed to score a file of RAG outputs:
# input.jsonl: one object per line with "context", "answer", and optional "query"
gasp detect --input outputs.jsonl --model Qwen/Qwen2.5-1.5B-Instruct \
--k-chunks 5 --threshold 0.5 --output results.jsonl
# fast two-pass variant on CPU
gasp detect --input outputs.jsonl --model Qwen/Qwen2.5-0.5B-Instruct \
--economical --device cpu --output results.jsonl
Metrics
If you have reference labels (1 for an unsupported span, 0 for a grounded one), score
the detector with the full metric suite, ROC-AUC, PR-AUC, and point metrics at a threshold:
from gasp import evaluate
m = evaluate(labels, sensitivities_negated, threshold=None) # {'roc_auc': ..., 'pr_auc': ...}
Or from the command line, on a results file that carries a label column:
gasp eval --input labeled_results.csv --label-col label --score-col sensitivity --threshold 0.5
How it works
For each answer sentence GASP re-scores the fixed answer under three conditions, the full context, no context, and each context chunk removed in turn, and reads the log-likelihood drops and Jensen-Shannon divergences at the sentence's tokens. The largest per-chunk drop is the sentence's grounding sensitivity, and the chunk that produced it is returned as the candidate supporting passage. Every method sees the same character-span chunks and sentences, so the segmentation is defined once and never re-tokenized.
API
GASP(model_id, k_chunks=5, threshold=None, economical=False, sensitivity_feature="max_drop", ...)— the detector.GASP.detect(context, answer, query="", threshold=None) -> Detection— score one answer.GASP.detect_batch(items, threshold=None) -> list[Detection]— score many answers.Detection— iterable ofSentenceResult, with.flagged(),.to_records(),.summary().SentenceResult—index,text,sensitivity,supporting_chunk,supporting_chunk_index,features,flagged,.to_dict().evaluate,roc_auc,pr_auc,threshold_metrics— evaluation metrics.read_items,write_records— read a.jsonl/.csvof items, write result records.Scorer— the lower-level scorer, if you want the raw per-sentence features.Case,sentence_spans,chunk_spans— the canonical segmentation, usable without a model.
Citation
If you use GASP, please cite:
@article{bouke2026gasp,
title = {Grounding-Aware Sensitivity by Perturbation for span-level hallucination
detection in retrieval-augmented generation},
author = {Bouke, Mohamed Aly},
year = {2026},
note = {Preprint}
}
License
MIT. See LICENSE.
Metadata
Release files for gasp-rag 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| gasp_rag-0.2.0.tar.gz | 15.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| gasp_rag-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 31.6 kB
Release files / gasp_rag-0.2.0.tar.gz
| Download URL | gasp_rag-0.2.0.tar.gz |
|---|---|
| Size | 15.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2574f1c084485fc9bfe67c1a95e868d7acdb72212e88cec8161a84dae987d4ed
|
|
BLAKE2b-256 checksum How to use checksums |
bcb0879b8f7f89b51b1c739d2b909103c53d2abb3fd45c0e35b7d26cfa67d00c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.0
|
Release files / gasp_rag-0.2.0-py3-none-any.whl
| Download URL | gasp_rag-0.2.0-py3-none-any.whl |
|---|---|
| Size | 16.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6a206e90a1d7969b9911959aa4bd290339a20e18156191928855db6a7f4586cb
|
|
BLAKE2b-256 checksum How to use checksums |
1028e3714db743342476d1ce4659c2f9b30e713bd10460c15c623b7845ff189f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.0
|