Detect retrieval regressions when chunking, indexing, ranking, or retrieval configurations change, using stable source-document evidence instead of fragile chunk IDs.
Why SPANCHOR · Quick start · Metrics · Framework integrations · CI · CLI · Documentation
Why SPANCHOR?
Retrieval evaluation is often tied to chunk IDs or chunk boundaries. Change the chunk size and the labels no longer line up, so you cannot tell whether retrieval got better or worse.
Traditional evaluation
Question → expected chunk ID → retrieved chunk ID → pass/fail
Re-chunking invalidates the gold labels.
SPANCHOR
Question → expected source span → retrieved text → source-span overlap
→ per-query metrics → regression decision
The canonical source document is the stable reference point. Documents are normalized (NFC Unicode, \n line endings) and hash-validated, so an anchor that no longer matches its source fails loudly instead of silently drifting.
With SPANCHOR you can change your chunker, index, retriever, or reranker and still measure whether the same source evidence is being retrieved, query by query.
Who is SPANCHOR for?
SPANCHOR is useful if you:
- build RAG systems and change chunking or retrieval strategies
- evaluate retrievers or rerankers
- maintain retrieval benchmarks
- want retrieval regression tests in CI
- use LangChain, LlamaIndex, or a custom retrieval stack
- need source-evidence-level evaluation rather than fragile chunk-ID comparisons
How it works
- Anchor expected evidence to character spans in canonical source documents.
- Map retrieved chunk text back to those source spans (or pass spans directly).
- Score each query: Recall@K, Precision@K, Hit@K, FullEvidence@K, IoU.
- Compare a baseline run against a candidate run under a per-metric policy.
- Gate CI on the result.
flowchart LR
A[Source documents] --> B[Canonicalization]
B --> C[Source anchors]
D[Your retriever] --> E[Retrieved results]
E --> M[Chunk-to-source mapping]
M --> F[SPANCHOR evaluation]
C --> F
F --> G[Per-query metrics]
G --> K[Baseline vs candidate comparison]
K --> L{Regression?}
L -->|No| N[Exit 0]
L -->|Yes| O[Exit 1]
SPANCHOR sits after retrieval: RAG pipeline → retrieval → SPANCHOR → regression gate → CI decision.
Key features
- Source-anchored evaluation: gold labels point to spans in canonical documents
- Character-level span metrics: Recall@K, Precision@K, Hit@K, FullEvidence@K, IoU
- Per-query regression detection with configurable per-metric thresholds
- Baseline vs candidate comparison with markdown reports
- Adapters: generic dicts, LangChain, LlamaIndex
- CLI for validate, evaluate, compare, locate, check-corpus, and anchor add
- CI exit codes for pass, regression, usage error, and unmapped-chunk-rate failures
- Local-first and deterministic: no network access or API keys required
- Typed Python API: checked with
mypy --strict
Installation
pip install spanchor # core library and CLI
pip install "spanchor[langchain]" # + LangChain adapter dependencies
pip install "spanchor[llamaindex]" # + LlamaIndex adapter dependencies
The LangChain and LlamaIndex integrations are optional extras and are not installed by default. Requires Python 3.11 or newer (tested on 3.11, 3.12, 3.13).
60-second quick start
from spanchor import Anchor, Document, Query, compare, evaluate
from spanchor.adapters.generic import dicts_to_retrieval_results
from spanchor.canonical.normalize import compute_hash
# 1. Canonical source document
text = "CloudSync is a cloud storage service. It syncs files across devices."
doc = Document.from_text("doc1", text)
# 2. Anchor the expected evidence to a source span
evidence = "CloudSync is a cloud storage service."
start = text.index(evidence)
anchor = Anchor("doc1", start, start + len(evidence), compute_hash(evidence))
# 3. Define the query
query = Query("q1", "What is CloudSync?", (anchor,))
# 4. Convert retrieval results (chunk text from any retriever)
baseline_results = {"q1": dicts_to_retrieval_results([{"text": evidence, "score": 0.9}])}
candidate_results = {"q1": dicts_to_retrieval_results([{"text": "It syncs files across devices.", "score": 0.9}])}
# 5. Evaluate both configurations
documents = {"doc1": doc}
baseline = evaluate(documents, [query], baseline_results, k=5)
candidate = evaluate(documents, [query], candidate_results, k=5)
print(baseline.aggregate_metrics["mean_recall@5"]) # 1.0
# 6. Compare per query against a policy (max allowed drop per metric)
result = compare(baseline, candidate, policy={"recall@5": 0.05})
# 7. Detect the regression
print(result.has_regression) # True
print(result.per_query_status) # {'q1': 'REGRESSION'}
Results given as chunk text (no document_id, no offsets) are mapped to source spans automatically. If you already know the source span, pass document_id, start, and end instead.
A gold set with only a handful of queries triggers a low-confidence warning; real gold sets should be larger.
Regression example
A validated benchmark from this repository: 100 gold queries over 4 technical documents (examples/real_world_validation/).
| Baseline | Candidate | |
|---|---|---|
| Chunk size | 250 | 200 |
| Overlap | 0 | 50 |
| Recall@5 | 0.405 | 0.249 |
Regressions: 19 queries exceeded the configured threshold
Status: FAIL
Exit code: 1
This is a measured result from the included validation run, not a hypothetical. It is specific to this corpus and these configurations; it demonstrates the regression-testing mechanism and is not a general claim about chunk size or overlap.
Metrics
| Metric | Meaning |
|---|---|
| Recall@K | Fraction of gold-evidence characters covered by the top-K retrieved spans |
| Precision@K | Fraction of retrieved characters that overlap gold evidence |
| Hit@K | Whether at least one gold span is covered by at least 50% (default min_overlap) |
| FullEvidence@K | Whether every gold span is covered by at least 50% |
| IoU | Intersection-over-union between retrieved and gold characters (diagnostic) |
Aggregate metrics (for example mean_recall@5) are for reporting and trending.
Per-query metrics are what the regression gate uses. Each query is classified as IMPROVED, UNCHANGED, or REGRESSION against the policy, and any single REGRESSION fails the comparison. An aggregate improvement does not cancel an individual query regression.
Framework integrations
All adapters return SPANCHOR RetrievalResult lists, ready for evaluate().
Generic: dict_to_retrieval_result() and dicts_to_retrieval_results() accept common field aliases (text/chunk/content/body, score/similarity, and others).
from spanchor.adapters.generic import dicts_to_retrieval_results
results = dicts_to_retrieval_results([{"text": "retrieved chunk", "score": 0.95}])
LangChain (pip install "spanchor[langchain]")
from spanchor.adapters.langchain import documents_to_retrieval_results
docs = retriever.invoke("your query") # list of LangChain Documents
results = documents_to_retrieval_results(docs) # score_key / document_id_key are configurable
LlamaIndex (pip install "spanchor[llamaindex]")
from spanchor.adapters.llamaindex import nodes_to_retrieval_results
nodes = retriever.retrieve("your query") # list of NodeWithScore
results = nodes_to_retrieval_results(nodes)
Use SPANCHOR as a CI regression gate
Baseline retrieval run
↓
Candidate retrieval run
↓
spanchor compare
↓
Policy thresholds
↓
PASS / FAIL
spanchor evaluate docs/ gold.jsonl baseline_results.json -o baseline.json
spanchor evaluate docs/ gold.jsonl candidate_results.json -o candidate.json
spanchor compare baseline.json candidate.json --policy policy.json --report comparison.md
policy.json sets the maximum allowed per-query drop for each metric:
{"recall@5": 0.05, "precision@5": 0.05, "hit@5": 0.1}
| Exit code | Meaning |
|---|---|
0 |
Success, or no regression (compare) |
1 |
Regression detected (compare) |
2 |
Schema, usage, or validation error |
3 |
Unmapped chunk rate exceeded (evaluate) |
CLI
| Command | Purpose |
|---|---|
spanchor validate DOCS_DIR GOLD_PATH |
Validate every anchor against its source document |
spanchor evaluate DOCS_DIR GOLD_PATH RESULTS_PATH |
Compute metrics (--k, --output, --report, --max-unmapped-rate, --redact) |
spanchor compare BASELINE CANDIDATE |
Per-query regression comparison (--policy, --report, --max-recall-drop, --max-precision-drop, --redact) |
spanchor locate TEXT DOCS_DIR |
Find text in documents and print anchor offsets and hash |
spanchor check-corpus DOCS_DIR GOLD_PATH |
Check anchor hashes, bounds, and text against the corpus |
spanchor anchor add |
Add a validated anchor to a gold JSONL file |
Run spanchor <command> --help for full options.
Real-world validation and adoption demo
examples/adoption_demo/ runs a complete, self-checking scenario: source documents, gold queries, baseline and candidate retrieval, a regression scenario that fails, and a passing scenario that succeeds.
cd examples/adoption_demo
python run_demo.py
The underlying benchmark and its scripts live in examples/real_world_validation/.
Documentation
| Resource | Purpose |
|---|---|
| Adoption demo | Integrating SPANCHOR into a retrieval pipeline |
| Real-world validation | 100-query benchmark and regression scenario |
| Minimal demo | Small CLI walkthrough with a gold set |
| Changelog | Release history |
Scope
SPANCHOR is an evaluation layer. It does not:
- retrieve documents
- generate embeddings
- chunk documents
- generate answers
- host a vector database
- parse PDFs or run OCR
- provide an LLM
- replace a RAG framework
It is designed to sit next to whichever of these you already use.
SPANCHOR is an early-stage open-source project; APIs and integrations may evolve.
Development
git clone https://github.com/Mukeshram-07/spanchor.git
cd spanchor
pip install -e ".[dev]"
pytest
mypy --strict src/spanchor
ruff check src/spanchor --select E,W,F,I --ignore B008,E501
ruff format src/spanchor tests --check
pip install build && python -m build
Testing
v0.2.0 is validated by 754 automated tests. Measured line and branch coverage is currently about 79%; this is a snapshot, not a guarantee.
pytest
Roadmap
Current release: v0.2.0
Implemented:
- Source-anchored evaluation and per-query regression detection
- Generic, LangChain, and LlamaIndex adapters
- CLI with CI exit codes
- Adoption demo and real-world validation example
Planned / exploratory (not committed to any release)
- Schema versioning improvements
- Richer policy configuration
- Performance work
- Additional framework adapters
Contributing
Contributions are welcome. Please open a pull request that includes tests for new behavior and passes pytest, mypy --strict, and ruff.
License
Apache-2.0. See LICENSE.
Metadata
Release files for spanchor 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| spanchor-0.3.0.tar.gz | 1.3 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| spanchor-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.3 MB
Release files / spanchor-0.3.0.tar.gz
| Download URL | spanchor-0.3.0.tar.gz |
|---|---|
| Size | 1.3 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1714febef9378ba1a81a35aa4ad1eaab2e9ef20d493fd325e3cae1a0bc6fd536
|
|
BLAKE2b-256 checksum How to use checksums |
46fd269657c4e78f86857b0f6e219b3d90b953ce4185dc5dc806a77e953759f5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency logRelease files / spanchor-0.3.0-py3-none-any.whl
| Download URL | spanchor-0.3.0-py3-none-any.whl |
|---|---|
| Size | 79.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
326383c60c2de3af7b3f029f54a2b3e6152afaf182619c187bf026fb99ef2e7f
|
|
BLAKE2b-256 checksum How to use checksums |
7608e6b01b480fb063142cf367c3c8de09526b3b1b34b5f5d985d59be8091fb5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency log