Spanchor
Regression-testing for RAG retrieval pipelines using stable source-document anchors.
Overview
Spanchor is a Python library and CLI tool that enables deterministic, regression-testing for RAG retrieval systems. Instead of relying on fragile chunk IDs, spanchor anchors gold labels to character spans in canonical source documents, then evaluates any retriever output against those stable anchors.
Core Principle: Separate SOURCE EVIDENCE from RETRIEVAL REPRESENTATION.
Key Features
- 🎯 Stable Anchors: Gold labels point to character spans in canonical documents (NFC normalized, hash-verified)
- 📊 Character-Level Metrics: Recall@K, Precision@K, Hit@K, FullEvidence@K, IoU
- 🔄 Chunk Mapping: Automatically maps chunk-text retrieval results back to source spans
- 🚨 Regression Detection: Compare baseline vs candidate runs with configurable thresholds
- ✅ CI-Ready: Exit codes and markdown reports for easy integration
- 🔒 Local-First: No network calls, no telemetry, deterministic operation
Installation
# Using uv (recommended)
uv pip install -e .
# With development dependencies
uv pip install -e ".[dev]"
Quick Start
from spanchor import Document, Anchor, evaluate
# 1. Create canonical documents
doc = Document.from_text("doc1", "Hello world! This is a test.")
# 2. Create anchors (in practice, use CLI helpers)
anchor = Anchor(
document_id="doc1",
start=0,
end=12,
expected_text_hash=doc.sha256,
)
# 3. Evaluate retrieval results
# (See full documentation for evaluation API)
CLI Commands
# Validate gold set against documents
spanchor validate docs/ gold.jsonl
# Evaluate retrieval results
spanchor evaluate docs/ gold.jsonl results.json --report report.md
# Compare baseline vs candidate
spanchor compare baseline.json candidate.json --report comparison.md
# Find text in documents for anchoring
spanchor locate "search text" docs/
# Check corpus health
spanchor check-corpus docs/ gold.jsonl
Requirements
- Python 3.11, 3.12, or 3.13
- Runtime:
typer,rich - Dev:
pytest,hypothesis,mypy,ruff
License
Apache-2.0
Non-Goals
Spanchor is NOT:
- A RAG framework
- A vector database
- An embedding manager
- A document parser/OCR tool
- An LLM provider integration
- An answer quality evaluator
It's a focused tool for retrieval regression testing only.
Metadata
Release files for spanchor 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| spanchor-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Release files / spanchor-0.1.0-py3-none-any.whl
| Download URL | spanchor-0.1.0-py3-none-any.whl |
|---|---|
| Size | 68.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5c384702aa74d85af1586a431073234821b88244d6cbe8a8204ab910322c60f0
|
|
BLAKE2b-256 checksum How to use checksums |
73f15463d95fc8c4873b0d54882fc4ced5e661eebdc73bad12732a92214d9440
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|