tesserakit-rag
Compile a document corpus plus a set of queries into a validated retrieval eval dataset.
tessera-rag reads a directory holding a corpus/ of documents and a queries file, builds a canonical RagCase dataset (each query with its gold retrieval target documents and optional expected answer), verifies every document reference, and emits a dataset plus reports.
Scope (v0.1)
This pack builds and validates the dataset. It does not run retrieval (no embeddings, no vector store, no scoring). Like the api pack (no HTTP execution) and evals (no LLM calls), execution is a runtime concern deferred to a later version. v0.1 is the offline "is this retrieval eval set well-formed and internally consistent" pass.
Input shape
my_rag_eval/
corpus/
refunds.md
billing/disputes.md
queries.jsonl (or queries.yaml)
Document ids are the corpus-relative path without suffix: refunds, billing/disputes.
Each query (one JSON object per line in queries.jsonl, or a YAML list):
{"id": "q1", "query": "Can I get a refund after 45 days?", "expected_answer": "No, the window is 30 days.", "relevant_docs": ["refunds"], "tags": ["billing"]}
relevant_docs is the gold set the retriever should surface. expected_answer is optional; queries without one are flagged for human review.
Compile a RAG eval pack
tessera rag compile --input examples/rag/ --output ./out/rag_pack
Artifacts written:
dataset.jsonl canonical RagCase rows (query, expected, relevant doc ids)
corpus_index.jsonl RagDocument rows (id, title, counts, sha256)
validation_report.md reference + hygiene findings
coverage_report.md answer/target coverage, orphan docs, avg targets per query
retrieval_targets.md per-query gold document set (with titles)
Validation rules
parse_error— a queries line/file could not be parsedempty_corpus— no documents found undercorpus/empty_document,duplicate_doc_idmissing_query_text,duplicate_querydangling_doc_reference— a query references a document not in the corpusquery_without_relevant_docs— no retrieval target to score againstquery_without_expected_answer— needs human revieworphan_document— a corpus document no query references
Metadata
Release files for tesserakit-rag 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tesserakit_rag-0.4.0.tar.gz | 8.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tesserakit_rag-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 18.0 kB
Release files / tesserakit_rag-0.4.0.tar.gz
| Download URL | tesserakit_rag-0.4.0.tar.gz |
|---|---|
| Size | 8.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d8b5d75688ea388773625b6fd93be95c9a150e21d2882235341b6e9982431cf9
|
|
BLAKE2b-256 checksum How to use checksums |
902b38c41d188678486226e2779688539688154c3f3d538ae39fd6a13c75a2e9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.11
|
Release files / tesserakit_rag-0.4.0-py3-none-any.whl
| Download URL | tesserakit_rag-0.4.0-py3-none-any.whl |
|---|---|
| Size | 10.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
553008a0d876155e78c31496105d2eb68c5f6d042b8b0e9dd8a3c98382927ceb
|
|
BLAKE2b-256 checksum How to use checksums |
c79184fb2409bb55b9c4132b0b48e1523c8fdda4cbb0571237e92cdb85fa48ad
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.11
|