Skip to main content

GASP

PyPI Python License: MIT

Grounding-Aware Sensitivity by Perturbation — a span-level detector of ungrounded content in retrieval-augmented generation (RAG).

GASP scores each answer sentence by its grounding sensitivity: the change in the sentence's likelihood when the retrieved context is perturbed. A grounded sentence loses much of its likelihood when its supporting passage is removed; an unsupported sentence barely reacts. GASP needs only a probabilistic scorer, no trained verifier and no labeled data, and it returns, for each sentence, the chunk that best supports it.

Install

pip install gasp-rag          # core
pip install gasp-rag[torch]   # with PyTorch and transformers, needed to run a scorer

Quickstart

from gasp import GASP

detector = GASP("Qwen/Qwen2.5-1.5B-Instruct", k_chunks=5)

context = "..."          # the retrieved passages, as one string
answer = "..."           # the generated answer to check
question = "..."         # the query (optional)

result = detector.detect(context=context, answer=answer, query=question)

for s in result:
    print(f"{s.sensitivity:+.2f}  {s.text}")
    if s.supporting_chunk:
        print(f"        supported by: {s.supporting_chunk[:80]}...")

Higher sensitivity means the sentence depends more on the retrieved evidence and is more likely grounded. Lower sensitivity means it barely reacts to removing evidence and is more likely unsupported. To flag sentences, pass a threshold:

result = detector.detect(context, answer, threshold=0.5)
for s in result.flagged():
    print("likely unsupported:", s.text)

Thresholds are corpus dependent and are best calibrated on held-out data; the continuous sensitivity score is the primary output.

Options

Everything is configurable on the detector:

GASP(
    model_id,                       # any Hugging Face causal LM, small or large, CPU or GPU
    k_chunks=5,                     # number of context chunks
    threshold=None,                 # flag sentences below this sensitivity
    economical=False,               # two-pass variant: faster, no attribution
    sensitivity_feature="max_drop", # or "gap", "mean_drop", "top2_drop", "max_jsd"
    max_ctx_tokens=1800,            # context truncation
    max_ans_tokens=256,             # answer truncation
    device=None,                    # "cpu" or "cuda"
    dtype=None,                     # "float16", "bfloat16", "float32"
)

Change the scorer at any time by constructing a new detector with a different model_id.

Many answers and files

Score a list of answers, or read a .jsonl/.csv file:

items = [
    {"context": ctx1, "answer": ans1, "query": q1},
    {"context": ctx2, "answer": ans2},
]
detections = detector.detect_batch(items, threshold=0.5)

# results as plain dicts, ready for pandas or JSON
import pandas as pd
rows = [r for d in detections for r in d.to_records()]
df = pd.DataFrame(rows)
print(detections[0].summary())   # {'n_sentences': ..., 'n_flagged': ..., 'mean_sensitivity': ...}

Command line

No code needed to score a file of RAG outputs:

# input.jsonl: one object per line with "context", "answer", and optional "query"
gasp detect --input outputs.jsonl --model Qwen/Qwen2.5-1.5B-Instruct \
            --k-chunks 5 --threshold 0.5 --output results.jsonl

# fast two-pass variant on CPU
gasp detect --input outputs.jsonl --model Qwen/Qwen2.5-0.5B-Instruct \
            --economical --device cpu --output results.jsonl

Metrics

If you have reference labels (1 for an unsupported span, 0 for a grounded one), score the detector with the full metric suite, ROC-AUC, PR-AUC, and point metrics at a threshold:

from gasp import evaluate
m = evaluate(labels, sensitivities_negated, threshold=None)   # {'roc_auc': ..., 'pr_auc': ...}

Or from the command line, on a results file that carries a label column:

gasp eval --input labeled_results.csv --label-col label --score-col sensitivity --threshold 0.5

How it works

For each answer sentence GASP re-scores the fixed answer under three conditions, the full context, no context, and each context chunk removed in turn, and reads the log-likelihood drops and Jensen-Shannon divergences at the sentence's tokens. The largest per-chunk drop is the sentence's grounding sensitivity, and the chunk that produced it is returned as the candidate supporting passage. Every method sees the same character-span chunks and sentences, so the segmentation is defined once and never re-tokenized.

API

  • GASP(model_id, k_chunks=5, threshold=None, economical=False, sensitivity_feature="max_drop", ...) — the detector.
  • GASP.detect(context, answer, query="", threshold=None) -> Detection — score one answer.
  • GASP.detect_batch(items, threshold=None) -> list[Detection] — score many answers.
  • Detection — iterable of SentenceResult, with .flagged(), .to_records(), .summary().
  • SentenceResult — index, text, sensitivity, supporting_chunk, supporting_chunk_index, features, flagged, .to_dict().
  • evaluate, roc_auc, pr_auc, threshold_metrics — evaluation metrics.
  • read_items, write_records — read a .jsonl/.csv of items, write result records.
  • Scorer — the lower-level scorer, if you want the raw per-sentence features.
  • Case, sentence_spans, chunk_spans — the canonical segmentation, usable without a model.

Citation

If you use GASP, please cite:

@article{bouke2026gasp,
  title   = {Grounding-Aware Sensitivity by Perturbation for span-level hallucination
             detection in retrieval-augmented generation},
  author  = {Bouke, Mohamed Aly},
  year    = {2026},
  note    = {Preprint}
}

License

MIT. See LICENSE.

Metadata

Release files for gasp-rag 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gasp-rag 0.2.0
File Size Uploaded
gasp_rag-0.2.0.tar.gz 15.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gasp-rag 0.2.0
File Interpreter ABI Platform
gasp_rag-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 31.6 kB

Release files / gasp_rag-0.2.0.tar.gz

Download URL gasp_rag-0.2.0.tar.gz
Size 15.6 kB
Tags Source
SHA-256 checksum
How to use checksums
2574f1c084485fc9bfe67c1a95e868d7acdb72212e88cec8161a84dae987d4ed
BLAKE2b-256 checksum
How to use checksums
bcb0879b8f7f89b51b1c739d2b909103c53d2abb3fd45c0e35b7d26cfa67d00c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release files / gasp_rag-0.2.0-py3-none-any.whl

Download URL gasp_rag-0.2.0-py3-none-any.whl
Size 16.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6a206e90a1d7969b9911959aa4bd290339a20e18156191928855db6a7f4586cb
BLAKE2b-256 checksum
How to use checksums
1028e3714db743342476d1ce4659c2f9b30e713bd10460c15c623b7845ff189f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release history Release notifications | RSS feed

0.2.1

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page