Skip to main content

GASP

PyPI Python License: MIT

Grounding-Aware Sensitivity by Perturbation — a span-level detector of ungrounded content in retrieval-augmented generation (RAG).

GASP scores each answer sentence by its grounding sensitivity: the change in the sentence's likelihood when the retrieved context is perturbed. A grounded sentence loses much of its likelihood when its supporting passage is removed; an unsupported sentence barely reacts. GASP needs only a probabilistic scorer, no trained verifier and no labeled data, and it returns, for each sentence, the chunk that best supports it.

Install

pip install gasp-rag          # core
pip install gasp-rag[torch]   # with PyTorch and transformers, needed to run a scorer

Quickstart

from gasp import GASP

detector = GASP("Qwen/Qwen2.5-1.5B-Instruct", k_chunks=5)

context = "..."          # the retrieved passages, as one string
answer = "..."           # the generated answer to check
question = "..."         # the query (optional)

result = detector.detect(context=context, answer=answer, query=question)

for s in result:
    print(f"{s.sensitivity:+.2f}  {s.text}")
    if s.supporting_chunk:
        print(f"        supported by: {s.supporting_chunk[:80]}...")

Higher sensitivity means the sentence depends more on the retrieved evidence and is more likely grounded. Lower sensitivity means it barely reacts to removing evidence and is more likely unsupported. To flag sentences, pass a threshold:

result = detector.detect(context, answer, threshold=0.5)
for s in result.flagged():
    print("likely unsupported:", s.text)

Thresholds are corpus dependent and are best calibrated on held-out data; the continuous sensitivity score is the primary output.

Options

Everything is configurable on the detector:

GASP(
    model_id,                       # any Hugging Face causal LM, small or large, CPU or GPU
    k_chunks=5,                     # number of context chunks
    threshold=None,                 # flag sentences below this sensitivity
    economical=False,               # two-pass variant: faster, no attribution
    sensitivity_feature="max_drop", # or "gap", "mean_drop", "top2_drop", "max_jsd"
    max_ctx_tokens=1800,            # context truncation
    max_ans_tokens=256,             # answer truncation
    device=None,                    # "cpu" or "cuda"
    dtype=None,                     # "float16", "bfloat16", "float32"
)

Change the scorer at any time by constructing a new detector with a different model_id.

Many answers and files

Score a list of answers, or read a .jsonl/.csv file:

items = [
    {"context": ctx1, "answer": ans1, "query": q1},
    {"context": ctx2, "answer": ans2},
]
detections = detector.detect_batch(items, threshold=0.5)

# results as plain dicts, ready for pandas or JSON
import pandas as pd
rows = [r for d in detections for r in d.to_records()]
df = pd.DataFrame(rows)
print(detections[0].summary())   # {'n_sentences': ..., 'n_flagged': ..., 'mean_sensitivity': ...}

Command line

No code needed to score a file of RAG outputs:

# input.jsonl: one object per line with "context", "answer", and optional "query"
gasp detect --input outputs.jsonl --model Qwen/Qwen2.5-1.5B-Instruct \
            --k-chunks 5 --threshold 0.5 --output results.jsonl

# fast two-pass variant on CPU
gasp detect --input outputs.jsonl --model Qwen/Qwen2.5-0.5B-Instruct \
            --economical --device cpu --output results.jsonl

Metrics

If you have reference labels (1 for an unsupported span, 0 for a grounded one), score the detector with the full metric suite, ROC-AUC, PR-AUC, and point metrics at a threshold:

from gasp import evaluate
m = evaluate(labels, sensitivities_negated, threshold=None)   # {'roc_auc': ..., 'pr_auc': ...}

Or from the command line, on a results file that carries a label column:

gasp eval --input labeled_results.csv --label-col label --score-col sensitivity --threshold 0.5

How it works

For each answer sentence GASP re-scores the fixed answer under three conditions, the full context, no context, and each context chunk removed in turn, and reads the log-likelihood drops and Jensen-Shannon divergences at the sentence's tokens. The largest per-chunk drop is the sentence's grounding sensitivity, and the chunk that produced it is returned as the candidate supporting passage. Every method sees the same character-span chunks and sentences, so the segmentation is defined once and never re-tokenized.

API

  • GASP(model_id, k_chunks=5, threshold=None, economical=False, sensitivity_feature="max_drop", ...) — the detector.
  • GASP.detect(context, answer, query="", threshold=None) -> Detection — score one answer.
  • GASP.detect_batch(items, threshold=None) -> list[Detection] — score many answers.
  • Detection — iterable of SentenceResult, with .flagged(), .to_records(), .summary().
  • SentenceResult — index, text, sensitivity, supporting_chunk, supporting_chunk_index, features, flagged, .to_dict().
  • evaluate, roc_auc, pr_auc, threshold_metrics — evaluation metrics.
  • read_items, write_records — read a .jsonl/.csv of items, write result records.
  • Scorer — the lower-level scorer, if you want the raw per-sentence features.
  • Case, sentence_spans, chunk_spans — the canonical segmentation, usable without a model.

Reproducing the evaluation

The experiment pipeline, the corrected source-level results, the figures, and the human study live in the project repository at github.com/drbouke/GASP (see pipeline/ and results/). In short, GASP beats entailment and attribution baselines and matches the per-chunk trained verifiers, while a full-context fact-checker and an LLM judge rank spans more accurately at higher compute. Adding GASP to a verifier helps the weaker entailment, attribution, and per-chunk verifiers but not the strongest full-context fact-checker or the LLM judge, so it is best used as a cheap, training-free standalone detector with built-in attribution and as a complement to weaker verifiers.

Citation

If you use GASP, please cite:

@article{bouke2026gasp,
  title   = {Grounding-Aware Sensitivity by Perturbation for span-level hallucination
             detection in retrieval-augmented generation},
  author  = {Bouke, Mohamed Aly},
  year    = {2026},
  note    = {Preprint}
}

License

MIT. See LICENSE.

Metadata

Release files for gasp-rag 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gasp-rag 0.2.1
File Size Uploaded
gasp_rag-0.2.1.tar.gz 16.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gasp-rag 0.2.1
File Interpreter ABI Platform
gasp_rag-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 32.4 kB

Release files / gasp_rag-0.2.1.tar.gz

Download URL gasp_rag-0.2.1.tar.gz
Size 16.1 kB
Tags Source
SHA-256 checksum
How to use checksums
9e331d86498c02e6fcc28b6c42b7ce4962d126a8f8fb91c5241a396267fdaee0
BLAKE2b-256 checksum
How to use checksums
45c37a33e257c72a480c0ece83fdab2b7ae7efccf53a829af0d561b97214b014
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release files / gasp_rag-0.2.1-py3-none-any.whl

Download URL gasp_rag-0.2.1-py3-none-any.whl
Size 16.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
551115f92983ef3e6cd364f9f3962804bb60cef48fce491a452958bfde16f02f
BLAKE2b-256 checksum
How to use checksums
9c9da17f4edf83f1712643900e05936c1db2db23e63359a023f48d770439c62e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page