Skip to main content

🍬 SuiteEval

Python License PyTerrier

Tools for running IR evaluation suites with PyTerrier.
SuiteEval helps you define, run, and aggregate evaluations across datasets while managing temporary indices and memory footprint.

📘 Overview

SuiteEval provides:

  • Declaration of pipelines (BM25, dense, re-ranking chains).
  • Execution of evaluation suites (e.g., BEIR-style benchmarks).
  • DatasetContext utilities for temporary paths and text loading.
  • DataFrame outputs for downstream analysis.

Workflow:

  1. Implement pipelines(context) that yields one or more PyTerrier pipelines (optionally named).
  2. Pass it to a suite (e.g., BEIR).
  3. Analyse the returned DataFrame.

🚀 Getting Started

Install from PyPI

pip install suiteeval

Install from source

git clone https://github.com/Parry-Parry/suiteeval.git
cd suiteeval
pip install -e .

⚙️ Defining Pipelines

Write a callable that accepts a DatasetContext and returns or yields pipelines.

  • Return a list/tuple of pipelines or (pipeline, name) pairs; or
  • Yield pipelines to keep only one large model resident in memory.

DatasetContext provides:

  • context.path — temporary working directory for indices/artifacts.
  • context.get_corpus_iter() — iterator suitable for indexing.
  • context.text_loader() — attaches document text for re-ranking.

Example

from suiteeval import NanoBEIR
from pyterrier_pisa import PisaIndex
from pyterrier_dr import ElectraScorer
from pyterrier_t5 import MonoT5ReRanker

def pipelines(context):
    index = PisaIndex(context.path + "/index.pisa")
    index.index(context.get_corpus_iter())

    bm25 = index.bm25(num_results=10)
    yield bm25 >> context.text_loader() >> MonoT5ReRanker(), "BM25 >> monoT5"
    yield bm25 >> context.text_loader() >> ElectraScorer(), "BM25 >> monoELECTRA"

results = BEIR(pipelines)

This would product a table as follows:

name nDCG@10 dataset
0 BM25 >> monoT5 0.26704 nano-beir/arguana
1 BM25 >> monoELECTRA 0.311608 nano-beir/arguana
2 BM25 >> monoT5 0.35844 nano-beir/climate-fever
3 BM25 >> monoELECTRA 0.369699 nano-beir/climate-fever
4 BM25 >> monoT5 0.647339 nano-beir/dbpedia-entity
5 BM25 >> monoELECTRA 0.647961 nano-beir/dbpedia-entity
6 BM25 >> monoT5 0.895196 nano-beir/fever
7 BM25 >> monoELECTRA 0.896831 nano-beir/fever
8 BM25 >> monoT5 0.482808 nano-beir/fiqa
9 BM25 >> monoELECTRA 0.478316 nano-beir/fiqa
10 BM25 >> monoT5 0.881539 nano-beir/hotpotqa
11 BM25 >> monoELECTRA 0.878735 nano-beir/hotpotqa
12 BM25 >> monoT5 0.58611 nano-beir/msmarco
13 BM25 >> monoELECTRA 0.576923 nano-beir/msmarco
14 BM25 >> monoT5 0.33767 nano-beir/nfcorpus
15 BM25 >> monoELECTRA 0.333013 nano-beir/nfcorpus
16 BM25 >> monoT5 0.721034 nano-beir/nq
17 BM25 >> monoELECTRA 0.696591 nano-beir/nq
18 BM25 >> monoT5 0.900461 nano-beir/quora
19 BM25 >> monoELECTRA 0.877153 nano-beir/quora
20 BM25 >> monoT5 0.359567 nano-beir/scidocs
21 BM25 >> monoELECTRA 0.351578 nano-beir/scidocs
22 BM25 >> monoT5 0.761235 nano-beir/scifact
23 BM25 >> monoELECTRA 0.741878 nano-beir/scifact
24 BM25 >> monoT5 0.709486 nano-beir/webis-touche2020
25 BM25 >> monoELECTRA 0.704123 nano-beir/webis-touche2020
26 BM25 >> monoELECTRA 0.56563 Overall
27 BM25 >> monoT5 0.564356 Overall

🧪 Running Suites

Entry points (e.g., BEIR) accept your pipeline factory and return a DataFrame:

results = BEIR(pipelines)  # per-dataset metrics and system names (if provided)

📦 Reproducibility & Resource Management

  • Temporary indices live under context.path and are cleaned up.
  • Prefer yielding pipelines when using large models.
  • Name systems via (pipeline, "<name>") for clear result tables and logs.

Persistent Index Storage

By default, indices are stored in temporary directories. To persist indices across runs, use the index_dir parameter:

# Indices will be stored in ./indices/<corpus-name>/
# Run files will be stored in ./results/<dataset-name>/
results = BEIR(
    pipelines,
    save_dir="./results",   # Where to save run files (per-dataset)
    index_dir="./indices"   # Where to store indices (per-corpus)
)

Key differences:

  • save_dir creates per-dataset subdirectories (e.g., ./results/beir-arguana/)
  • index_dir creates per-corpus subdirectories (e.g., ./indices/beir-arguana/)
  • Multiple datasets sharing a corpus will reuse the same index directory

Automatic Result Caching

When using save_dir, SuiteEval automatically skips inference for pipelines that already have saved run files. If a {pipeline_name}.res.gz file exists for all datasets in a corpus, the suite loads results from disk instead of re-running the pipeline.

# First run: executes inference and saves results
results = BEIR(pipelines, save_dir="./results")

# Second run: automatically loads from ./results/{dataset}/{name}.res.gz
results = BEIR(pipelines, save_dir="./results")

To force re-running inference, use save_mode="overwrite":

# Always re-run inference, even if files exist
results = BEIR(pipelines, save_dir="./results", save_mode="overwrite")

🔧 Customising a Suite

A suite run is a fixed sequence of small, overridable steps. To change one part of the computation, override the hook that covers it rather than reimplementing __call__:

__call__
  └─ resolve_config          split call arguments into a RunConfig
  └─ run
     └─ iter_corpus_groups   group datasets by shared corpus
        └─ select_members    pick datasets to evaluate in this group
        └─ build_context     build the shared DatasetContext (indexes once)
        └─ iter_pipeline_batches
           └─ wrap_pipeline          decorate each pipeline
           └─ prepare_topics_qrels   fetch topics and qrels
           └─ evaluate_batch
              ├─ has_cached_run / run_file_path / load_cached_run
              └─ run_experiment      the pt.Experiment call
                 └─ measures_for     pick metrics for a dataset
           └─ annotate_results       tag rows with their dataset
           └─ release_pipelines      free memory between batches
  └─ postprocess_results     aggregate the concatenated table
     └─ compute_overall_mean

Wrap every pipeline — applies to both sequential and grouped execution:

class MySuite(Suite):
    _datasets = ["my-corpus/test"]

    def wrap_pipeline(self, pipeline, context):
        return pipeline >> MyResultFilter(context.dataset.get_qrels())

Reshape the results table before the Overall rows are added:

def postprocess_results(self, results, config):
    results = merge_sub_collections(results)
    return super().postprocess_results(results, config)

Change where runs are cached:

def run_file_path(self, save_dir, dataset_name, pipeline_name):
    return os.path.join(save_dir, f"{dataset_name}--{pipeline_name}.res.gz")

🛠️ Compatibility

Works with modern PyTerrier and common extensions
(e.g., pyterrier_pisa, pyterrier_dr, pyterrier_t5).
For older environments, ensure standard PyTerrier transformer interfaces.

👥 Authors

🧾 Version History

Version Date Changes
0.1.8 2026-08-30 Overridable evaluation hooks; fix save_dir being dropped after the first corpus
0.1.7 2026-02-16 Tempoary removal of DL23 until qrels are adeed
0.1.6 2026-02-03 Fix duplicate Overall rows, auto-detect all metrics
0.1.5 2026-01-07 Custom index folder support for persistent indices
0.1.4 2025-12-01 Fix save directory handling
0.1.3 2025-12-01 PyTerrier 1.0 compatibility, mixed datasets support
0.1.2 2025-10-29 Documentation improvements and bug fixes
0.1 2025-10-03 Initial release

License

This project is licensed under the MIT License — see the LICENSE file for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

suiteeval-0.1.8.tar.gz (33.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

suiteeval-0.1.8-py3-none-any.whl (38.3 kB view details)

Uploaded Python 3

File details

Details for the file suiteeval-0.1.8.tar.gz.

File metadata

  • Download URL: suiteeval-0.1.8.tar.gz
  • Upload date:
  • Size: 33.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.7

File hashes

Hashes for suiteeval-0.1.8.tar.gz
Algorithm Hash digest
SHA256 f6e25580c400c40e19095bfab5629532daaffea61b58fe377b9d9c548dcf3bb8
MD5 089ea8087747bada182e44d87b96018b
BLAKE2b-256 bd371891f4b5187c02cada9456f5eaca3907623e7f45a814e985c04e81e200b0

See more details on using hashes here.

File details

Details for the file suiteeval-0.1.8-py3-none-any.whl.

File metadata

  • Download URL: suiteeval-0.1.8-py3-none-any.whl
  • Upload date:
  • Size: 38.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.7

File hashes

Hashes for suiteeval-0.1.8-py3-none-any.whl
Algorithm Hash digest
SHA256 12473e60fe13bf44540ef47de2dce86ec30204824e39da3aa5e1c16c129ae276
MD5 57e856b767ba481bbd0a3235a104bb36
BLAKE2b-256 2ed81dd2d48cd48fe632f66e37dd9850139d3b366368b19b3619c4c2db4b540d

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.1.8 This release

2 files

0.1.7

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page