Skip to main content
The lsr-benchmark banner image

lsr-benchmark

CI Maintenance Code coverage
Release PyPi Downloads Commit activity

CLI • Python API • Citation

The lsr-benchmark aims to support holistic evaluations of the learned sparse retrieval paradigm to contrast efficiency and effectiveness across diverse retrieval scenarios. Please see the corresponding paper for an overview of the methodology.

Task

The learned sparse retrieval paradigm conducts retrieval in three steps:

  1. Documents are segmented into passages so that the passages can be processed by pre-trained transformers.
  2. Documents and queries are embedded into a sparse learned embedding.
  3. Retrieval systems create an index of the document embeddings to return a ranking for each embedded query.

You can submit solutions to step 2 (i.e., models that embed documents and queries into sparse embeddings) and/or solutions to step 3 (i.e., retrieval systems). The idea is then to validate all combinations of embeddings with all retrieval systems to identify which solutions work well for which use case, taking different notions of efficiency/effectiveness trade-offs into consideration. The passage segmentation for step 1 is open source (i.e., created via lsr-benchmark segment-corpus <IR-DATASETS-ID>) but fixed for this task.

Installation

You can install the lsr-benchmark via:

pip3 install lsr-benchmark

If you want the latest features, you can install from the main branch:

pip3 install git+https://github.com/reneuir/lsr-benchmark.git

Supported Corpora and Embeddings

Please run lsr-benchmark overview for an up-to-date overview over all datasets and all embeddings. Alternatively, online overview in TIRA provides an overview.

Retrieval Suites

Predefined suites select the datasets, embeddings, and retrieval engines for a benchmark run:

lsr-benchmark retrieval --suite reneuir-2026/full --out my-reneuir-2026-results

Suites are maintained in lsr_benchmark/retrieval_suites.py. A suite cannot be combined with positional retrieval engines, --dataset, or --embedding.

Running Tests

We have a suite of unit tests that you can run via:

# first install the local version of the lsr-benchmark
pip3 install -e .[dev,test]
# then run the unit tests
pytest .

Documentation and Tutorials

We have a set of tutorials available.

The lsr-benchmark --help command serves as entrypoint to the documentation.

Instructions to add new datasets are available in the data directory.

  • ToDo: Write how to add new datasets, embeddings, retrieval, evaluation
    • short video

Data

The formats for data inputs and outputs aim to support slicing and dicing diverse query and document distributions while enabling caching, allowing for GreenIR research.

You can slice and dice the document texts and document embeddings via the API. The document texts for private corpora are only available within the TIRA sandbox whereas the document embeddings are publicly available for all corpora (as one can not re-construct the original documents from sparse embeddings).

dataset = lsr_benchmark.load('<IR-DATASETS-ID>')

# process the document embeddings:
for doc in dataset.docs_iter(embedding='<EMBEDDING-MODEL>', passage_aggregation="first-passage"):
    doc # namedtuple<doc_id, embedding>

# process the document embeddings for all segments:
for doc in dataset.docs_iter(embedding='<EMBEDDING-MODEL>'):
    doc # namedtuple<doc_id, segments.embedding>

# process the document texts:
for doc in dataset.docs_iter(embedding=None):
    doc # namedtuple<doc_id, segments.text>

# process the document texts via segmented versions in ir_datasets
lsr_benchmark.register_to_ir_datasets()
for segmented_doc in ir_datasets.load(f"lsr-benchmark/{dataset}/segmented")
    doc # namedtuple<doc_id, segment>

Format of Document Texts

Inspired by the processing of MS MARCO v2.1, each document consists of a doc_id and a list of text segments that are short enough to be processed by pre-trained transformers. For instance, a document that consists of 4 passages (e.g., "text-of-passage-1 text-of-passage-2 text-of-passage-3 text-of-passage-4") would be represented as:

  • doc_id: 12fd3396-e4d7-4c0f-b468-5a82402b5336
  • segments:
    • {"start": 1, "end": 2, "text": "text-of-passage-1 text-of-passage-2"}
    • {"start": 2, "end": 3, "text": "text-of-passage-2 text-of-passage-3"}
    • {"start": 3, "end": 4, "text": "text-of-passage-3 text-of-passage-4"}

Format of Document Embeddings

Each document consists of a doc_id and a list of text segments that are short enough to be processed by pre-trained transformers. For instance, a document that consists of 4 passages would be represented as:

  • doc_id: 12fd3396-e4d7-4c0f-b468-5a82402b5336
  • segments:
    • {"start": 1, "end": 2, "embedding": {"term-1": 0.123, "term-2": 0.912}}
    • {"start": 2, "end": 3, "embedding": {"term-1": 0.421, "term-3": 0.743}}
    • {"start": 3, "end": 4, "embedding": {"term-2": 0.108, "term-4": 0.043}}

Evaluation

The online overview in TIRA provides an overview of aggregated evaluations. Alternatively, all data and further custom evaluations are available in the step-04-evaluation directory of this repository.

Our evaluation methodology encourages the development of diverse and novel measures for lsr models that take efficiency and effectiveness into consideration. We assume that a suitable interpretation of efficiency for a target task highly depends on the application and its context. Therefore, we aim to measure as many efficiency-oriented aspects as possible in a standardized way with the tirex-tracker to ensure that different efficiency/effectiveness interpretations can be evaluated post-hoc. This methodology and related aspects were developed as part of the ReNeuIR workshop series held at SIGIR 2022, 2023, 2024, and 2025.

Metadata

Release files for lsr-benchmark 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lsr-benchmark 0.1.1
File Size Uploaded
lsr_benchmark-0.1.1.tar.gz 9.2 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for lsr-benchmark 0.1.1
File Interpreter ABI Platform
lsr_benchmark-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 18.6 MB

Release files / lsr_benchmark-0.1.1.tar.gz

Download URL lsr_benchmark-0.1.1.tar.gz
Size 9.2 MB
Tags Source
SHA-256 checksum
How to use checksums
cdc691d3c6e277cfdca699caaef6e2330fff271a8f4ead9ba05e08ae61f0b9d5
BLAKE2b-256 checksum
How to use checksums
b784fdb056960b3631d6784f3599e53c9e5aa2e4eeea3e2392f162a0729f1918
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 19, 2026.

Transparency log

Release files / lsr_benchmark-0.1.1-py3-none-any.whl

Download URL lsr_benchmark-0.1.1-py3-none-any.whl
Size 9.4 MB
Tags Python 3
SHA-256 checksum
How to use checksums
1dc92d4b3c073bb99222a96759ce6653c13910809676a0df7f9a42488b3eaa1c
BLAKE2b-256 checksum
How to use checksums
026ccff479fd255027ee9b14ee919b16a202cef579e358eeb04a860125939afa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 19, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page