Skip to main content

pyterrier_splade

An example of a SPLADE learned sparse indexing and retrieval using PyTerrier transformers.

Installation

%pip install -q git+https://github.com/cmacdonald/pyt_splade.git

Indexing

Indexing takes place as a pipeline: we apply SPLADE transformation of the documents, which maps raw text into a dictionary of BERT WordPiece tokens and corresponding weights. The underlying indexer, Terrier, is configured to handle arbitrary word counts without further tokenisation (pretokenised=True).

The Terrier indexer is configured to index tokens unchanged.

import pyterrier as pt

import pyterrier_splade
splade = pyterrier_splade.Splade()
indexer = pt.IterDictIndexer('./msmarco_psg', pretokenised=True)

indxr_pipe = splade.doc_encoder() >> indexer
index_ref = indxr_pipe.index(dataset.get_corpus_iter(), batch_size=128)

Retrieval

Similarly, SPLADE encodes the query into BERT WordPieces and corresponding weights. We apply this as a query encoding transformer.

splade_retr = splade.query_encoder() >> pt.terrier.Retriever('./msmarco_psg', wmodel='Tf')

Scoring

SPLADE can also be used as a text scoring function.

first_stage = ... # e.g., BM25, dense retrieval, etc.
splade_scorer = first_stage >> dataset.text_loader() >> splade.scorer()

PISA

For faster retrieval with SPLADE, you can use the fast PISA retrieval backend provided by PyTerrier_PISA:

import pyterrier_splade
splade = pyterrier_splade.Splade()
dataset = pt.get_dataset('irds:msmarco-passage')
index = PisaIndex('./msmarco-passage-splade', stemmer='none')

# indexing
idx_pipeline = splade.doc_encoder() >> index.toks_indexer()
idx_pipeline.index(dataset.get_corpus_iter())

# retrieval

retr_pipeline = splade.query_encoder() >> index.quantized()

Demo

We have a demo of PyTerrier_SPLADE at https://huggingface.co/spaces/terrierteam/splade

Note

Note that this package used to be named pyt_splade. The package is still available under that name (but this may be removed in the future). The new name is pyterrier_splade.

Credits

  • Craig Macdonald
  • Sean MacAvaney
  • Nicola Tonellotto

Metadata

Release files for pyterrier-splade 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyterrier-splade 0.1.0
File Size Uploaded
pyterrier_splade-0.1.0.tar.gz 11.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyterrier-splade 0.1.0
File Interpreter ABI Platform
pyterrier_splade-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 23.3 kB

Release files / pyterrier_splade-0.1.0.tar.gz

Download URL pyterrier_splade-0.1.0.tar.gz
Size 11.0 kB
Tags Source
SHA-256 checksum
How to use checksums
f7d52079939ac9301a78e87edf98db5eedb3541e05fb65acfd28d2e397c397c8
BLAKE2b-256 checksum
How to use checksums
d44d581707d6ed978f98137b4d4e9abf450143d39c7606d08c366c5b9ce618df
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.6

Release files / pyterrier_splade-0.1.0-py3-none-any.whl

Download URL pyterrier_splade-0.1.0-py3-none-any.whl
Size 12.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
61a1a1ff360cc438d5a2872089e7c345a8c4c15562f40326122f154a22d55ced
BLAKE2b-256 checksum
How to use checksums
f8ac2507170b1c7aa5e5f0b1e19d061f5498128dbdf77936b8526bb289c788e3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.6

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page