IMRNNs
Interpretable Modular Retrieval Neural Networks
Efficient dense-retrieval reranking through lightweight embedding modulation.
Paper · PyPI · Model checkpoint · Project website · Author portfolio
IMRNNs adds a small, trainable reranking layer to a frozen dense retriever. It modulates query and document embeddings, improves the ordering of an existing candidate set, and exposes how that modulation changes each retrieval score. The base encoder remains unchanged.
The public release provides the installable Python package, command-line tools, reproducible training and evaluation workflows, and a validated MiniLM–SciFact checkpoint hosted on Hugging Face.
Highlights
- Lightweight adaptation — train compact query and document adapters while keeping the base retriever frozen.
- Drop-in reranking — work with text directly or pass existing NumPy and PyTorch embeddings.
- Interpretable scores — inspect base scores, adapted scores, score deltas, and vocabulary-level concepts.
- Strict evaluation — compare raw and adapted nDCG@10, Recall@10, and MRR@10 over identical candidate sets.
- Reproducible artifacts — retain encoder revisions, data hashes, split metadata, training history, and evaluation results in each checkpoint.
Installation
Install the stable package from PyPI:
python -m pip install imrnns
Install the optional training and BEIR evaluation dependencies:
python -m pip install "imrnns[train,eval]"
IMRNNs supports Python 3.10–3.12 and CPU inference. For local development, see the development guide.
Quick start
from imrnns import IMRNNAdapter
adapter = IMRNNAdapter.from_pretrained(
encoder="minilm",
dataset="scifact",
device="cpu",
)
results = adapter.rerank(
query="What is scientific evidence?",
documents=[
"Evidence is used to support or refute a scientific claim.",
"The stock market closed higher today.",
"A recipe describes how to prepare a meal.",
],
top_k=3,
)
for result in results:
print(
result.rank,
result.base_score,
result.adapted_score,
result.score_delta,
result.text,
)
from_pretrained() downloads the matching adapter checkpoint from
Hugging Face. Each result contains
its rank, original index, optional document ID and text, base-retriever cosine
score, adapted score, and score delta.
Validated results
The included checkpoint was selected on SciFact validation and evaluated once on the complete official 300-query test set. Raw and adapted rankings use the same top-100 MiniLM candidates.
| Metric | Raw MiniLM | IMRNN | Improvement |
|---|---|---|---|
| nDCG@10 | 0.64508 | 0.69166 | +0.04658 |
| Recall@10 | 0.78333 | 0.84333 | +0.06000 |
| MRR@10 | 0.60472 | 0.64800 | +0.04328 |
The same recipe passed all three validation metrics on NFCorpus and ArguAna. Held-out testing passed the strict all-metrics rule on SciFact and ArguAna. See the training study for full metrics and split details.
Rerank existing embeddings
rerank_embeddings() accepts NumPy arrays or PyTorch tensors and does not
invoke a text encoder:
adapter = IMRNNAdapter.from_pretrained(
encoder="minilm",
dataset="scifact",
load_encoder=False,
)
results = adapter.rerank_embeddings(
query_embedding=query_embedding,
document_embeddings=document_embeddings,
document_ids=document_ids,
top_k=10,
)
Explain a retrieval decision
explanation = adapter.explain(
query="What currency is used in Mexico?",
document="The Mexican peso is the currency of Mexico.",
top_tokens=10,
)
print(explanation.top_query_tokens)
print(explanation.top_document_tokens)
print(explanation.score_delta)
The explanation back-projects the query and document modulation vectors through the Moore–Penrose pseudoinverse of the learned projector and reports aligned encoder-vocabulary concepts.
Training and evaluation
Build a shared embedding cache with dense hard negatives:
imrnns cache \
--encoder minilm \
--dataset scifact \
--cache-dir ./cache/minilm-scifact \
--device cpu
Train an adapter:
imrnns train \
--encoder minilm \
--dataset scifact \
--cache-dir ./cache/minilm-scifact \
--output-dir ./checkpoints
Evaluate it against the frozen base retriever:
imrnns evaluate \
--encoder minilm \
--dataset scifact \
--datasets-dir ./datasets \
--cache-dir ./cache/minilm-scifact \
--checkpoint checkpoints/validated/minilm/imrnns-minilm-scifact.pt
The evaluation JSON contains base_metrics, adapted metrics,
metric_delta, and beats_base_all_metrics. A run passes only when every
adapted metric is greater than its corresponding base metric.
Default training configuration
- same-dimensional, identity-initialized projector;
- 128 hidden units and zero dropout;
- one highest-relevance positive and 63 dense hard negatives from the raw retriever's top 100 results;
- improvement-margin objective with margin
0.05; - Adam with learning rate
1e-4, weight decay1e-5, and batch size 32; - up to 30 epochs with patience 7;
- best-epoch selection by validation mean nDCG@10, Recall@10, and MRR@10, including epoch 0 as a candidate.
Local BEIR-format datasets
IMRNNs also accepts local datasets with the standard BEIR layout:
my_dataset/
├── corpus.jsonl
├── queries.jsonl
└── qrels/
├── train.tsv
└── test.tsv
For datasets with official train and test qrels, 15% of the official training queries are reserved for validation and the test set remains untouched. A dataset with one qrels split uses a deterministic 70/15/15 partition.
Runnable examples are available in the examples directory.
Command-line interface
| Command | Purpose |
|---|---|
imrnns info |
Show package, checkpoint, and training-recipe information |
imrnns download |
Download a released adapter checkpoint |
imrnns list-assets |
List supported encoders and checkpoints |
imrnns cache |
Prepare datasets, embeddings, and hard negatives |
imrnns train |
Train and validate an adapter |
imrnns evaluate |
Compare adapted and base-retriever metrics |
imrnns run |
Execute the cache, train, and evaluate pipeline |
Use imrnns <command> --help for complete options.
Development
git clone https://github.com/YashSaxena21/IMRNNs.git
cd IMRNNs
python -m pip install -e ".[dev]"
pytest
ruff check src tests scripts examples
python -m build
twine check dist/*
Bug reports, feature proposals, and focused pull requests are welcome through GitHub Issues and GitHub Pull Requests.
Citation
If IMRNNs supports your research, please cite the EACL 2026 paper:
@inproceedings{saxena-etal-2026-imrnns,
title = {{IMRNN}s: An Efficient Method for Interpretable Dense Retrieval via Embedding Modulation},
author = {Saxena, Yash and Padia, Ankur and Gunaratna, Kalpa and Gaur, Manas},
booktitle = {Findings of the Association for Computational Linguistics: EACL 2026},
year = {2026},
pages = {6324--6337},
doi = {10.18653/v1/2026.findings-eacl.333},
url = {https://aclanthology.org/2026.findings-eacl.333/}
}
Authors and acknowledgments
IMRNNs is authored by Yash Saxena, Ankur Padia, Kalpa Gunaratna, and Manas Gaur. Visit Yash Saxena's portfolio for additional projects and publications.
This work is associated with the University of Maryland, Baltimore County and the KAI² Lab, and was published in the Findings of EACL 2026.
License
The code, scripts, documentation, and checkpoints are available under the Creative Commons Attribution 4.0 International License. Attribution requirements are documented in ATTRIBUTION.md.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file imrnns-0.2.1.tar.gz.
File metadata
- Download URL: imrnns-0.2.1.tar.gz
- Upload date:
- Size: 1.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d915bd3a7a75ff186c39e2c945d1baa0d79cccb369fe4350f7981032706618be
|
|
| MD5 |
f9074ea934a177a3d7d9ed03f187a8f7
|
|
| BLAKE2b-256 |
4082312d5beaad41711b948ac12a6a61b97a721815cd97b9467f01ab420ef649
|
Provenance
The following attestation bundles were made for imrnns-0.2.1.tar.gz:
Publisher:
publish.yml on YashSaxena21/IMRNNs
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
imrnns-0.2.1.tar.gz -
Subject digest:
d915bd3a7a75ff186c39e2c945d1baa0d79cccb369fe4350f7981032706618be - Sigstore transparency entry: 2492917342
- Sigstore integration time:
-
Permalink:
YashSaxena21/IMRNNs@a5f3b51236b712d7171878dfd56475b1c05c5131 -
Branch / Tag:
refs/tags/v0.2.1 - Owner: https://github.com/YashSaxena21
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a5f3b51236b712d7171878dfd56475b1c05c5131 -
Trigger Event:
push
-
Statement type:
File details
Details for the file imrnns-0.2.1-py3-none-any.whl.
File metadata
- Download URL: imrnns-0.2.1-py3-none-any.whl
- Upload date:
- Size: 42.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ead679baa81263b81ffb27f36bd49a7deb6671f81ca4a26e8f452efd9bbbf451
|
|
| MD5 |
ae632eb93c8afa087ad18b8129d4b77b
|
|
| BLAKE2b-256 |
e5f2c8a833086e663e8aa45f0981c9dcc50279a12efeb86eb745c152019c751c
|
Provenance
The following attestation bundles were made for imrnns-0.2.1-py3-none-any.whl:
Publisher:
publish.yml on YashSaxena21/IMRNNs
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
imrnns-0.2.1-py3-none-any.whl -
Subject digest:
ead679baa81263b81ffb27f36bd49a7deb6671f81ca4a26e8f452efd9bbbf451 - Sigstore transparency entry: 2492917361
- Sigstore integration time:
-
Permalink:
YashSaxena21/IMRNNs@a5f3b51236b712d7171878dfd56475b1c05c5131 -
Branch / Tag:
refs/tags/v0.2.1 - Owner: https://github.com/YashSaxena21
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a5f3b51236b712d7171878dfd56475b1c05c5131 -
Trigger Event:
push
-
Statement type: