IMRNNs
Interpretable Modular Retrieval Neural Networks (IMRNNs) rerank candidates from a frozen dense retriever using lightweight query/document embedding modulation.
EACL 2026 paper · Hugging Face checkpoint · project website
This repository contains one implementation and one validated checkpoint:
MiniLM trained for SciFact at
checkpoints/validated/minilm/imrnns-minilm-scifact.pt.
Install
pip install imrnns
From this checkout:
python -m pip install .
Training and BEIR evaluation dependencies are optional:
python -m pip install ".[train,eval]"
Python 3.10–3.12 and CPU inference are supported.
Quickstart
from imrnns import IMRNNAdapter
adapter = IMRNNAdapter.from_pretrained(
encoder="minilm",
dataset="scifact",
device="cpu",
)
results = adapter.rerank(
query="What is scientific evidence?",
documents=[
"Evidence is used to support or refute a scientific claim.",
"The stock market closed higher today.",
"A recipe describes how to prepare a meal.",
],
top_k=3,
)
for result in results:
print(result.rank, result.base_score, result.adapted_score, result.score_delta, result.text)
Each result includes its rank, original index, optional document ID/text, raw base-retriever cosine score, adapted score, and score delta.
Rerank existing embeddings
rerank_embeddings() does not invoke a text encoder:
adapter = IMRNNAdapter.from_pretrained(
encoder="minilm",
dataset="scifact",
load_encoder=False,
)
results = adapter.rerank_embeddings(
query_embedding=query_embedding,
document_embeddings=document_embeddings,
document_ids=document_ids,
top_k=10,
)
Interpret a retrieval decision
explanation = adapter.explain(
query="What currency is used in Mexico?",
document="The Mexican peso is the currency of Mexico.",
top_tokens=10,
)
print(explanation.top_query_tokens)
print(explanation.top_document_tokens)
print(explanation.score_delta)
The explanation back-projects query/document modulation vectors through the Moore–Penrose pseudoinverse of the learned projector and reports aligned encoder-vocabulary concepts.
Training recipe
Prepare a shared embedding and dense-negative cache:
imrnns cache \
--encoder minilm \
--dataset scifact \
--cache-dir ./cache/minilm-scifact \
--device cpu
Train and evaluate:
imrnns train \
--encoder minilm \
--dataset scifact \
--cache-dir ./cache/minilm-scifact \
--output-dir ./checkpoints
The fixed recipe uses:
- a same-dimensional projector initialized to identity;
- 128 hidden units and zero dropout;
- one highest-relevance positive and 63 dense hard negatives from the raw retriever's top 100 per query;
- improvement-margin loss
mean(max(0, 0.05 + base_margin - adapted_margin)); - Adam with learning rate
1e-4, weight decay1e-5, and batch size 32; - up to 30 epochs with patience 7;
- best-epoch selection by validation mean nDCG@10, Recall@10, and MRR@10, including epoch 0 as a candidate.
Training checkpoints store the exact encoder revision, cache manifest, split hashes, training history, base/adapted test metrics, metric deltas, and strict pass status. A dataset passes only when every adapted metric is greater than the corresponding base metric.
Validated results
The included checkpoint was selected on SciFact validation and evaluated once on the complete official 300-query test set with the same top-100 candidates for the raw and adapted rankings.
| Metric | Raw MiniLM | IMRNN | Delta |
|---|---|---|---|
| nDCG@10 | 0.64508 | 0.69166 | +0.04658 |
| Recall@10 | 0.78333 | 0.84333 | +0.06000 |
| MRR@10 | 0.60472 | 0.64800 | +0.04328 |
The recipe also passed all three validation metrics on NFCorpus and ArguAna. Held-out testing passed the strict all-metrics rule on SciFact and ArguAna; NFCorpus improved nDCG@10 and Recall@10 but missed MRR@10 by 0.00014. Full numbers and split details are in TRAINING_STUDY.md.
Evaluate a checkpoint
imrnns evaluate \
--encoder minilm \
--dataset scifact \
--datasets-dir ./datasets \
--cache-dir ./cache/minilm-scifact \
--checkpoint checkpoints/validated/minilm/imrnns-minilm-scifact.pt
The JSON output includes base_metrics, adapted metrics, metric_delta, and
beats_base_all_metrics.
Local BEIR-format datasets
The loader accepts:
my_dataset/
├── corpus.jsonl
├── queries.jsonl
└── qrels/
├── train.tsv
└── test.tsv
For official train/test datasets, 15% of official train is reserved for validation and official test remains untouched. A dataset with one qrels split uses a deterministic 70/15/15 partition.
CLI
imrnns info
imrnns download --encoder minilm --dataset scifact
imrnns list-assets
imrnns cache ...
imrnns train ...
imrnns evaluate ...
imrnns run ...
Development
python -m pip install -e ".[dev]"
pytest
ruff check src tests scripts examples
python -m build
twine check dist/*
Citation
@inproceedings{saxena-etal-2026-imrnns,
title = {{IMRNN}s: An Efficient Method for Interpretable Dense Retrieval via Embedding Modulation},
author = {Saxena, Yash and Padia, Ankur and Gunaratna, Kalpa and Gaur, Manas},
booktitle = {Findings of the Association for Computational Linguistics: EACL 2026},
year = {2026},
pages = {6324--6337},
doi = {10.18653/v1/2026.findings-eacl.333},
url = {https://aclanthology.org/2026.findings-eacl.333/}
}
License
The code, scripts, documentation, and checkpoints are licensed under the Creative Commons Attribution 4.0 International License. See LICENSE and ATTRIBUTION.md.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file imrnns-0.2.0.tar.gz.
File metadata
- Download URL: imrnns-0.2.0.tar.gz
- Upload date:
- Size: 1.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d8fd00d82f1b7e030a4e1f8a27ada861d209dc8f2a6775e126a77d042faf3600
|
|
| MD5 |
510820b4ffac4092002b76cb2508356f
|
|
| BLAKE2b-256 |
33728aa64eed6bb2c647d7594539d1abfd564ac67e66e93ae764c1843823a8e4
|
Provenance
The following attestation bundles were made for imrnns-0.2.0.tar.gz:
Publisher:
publish.yml on YashSaxena21/IMRNNs
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
imrnns-0.2.0.tar.gz -
Subject digest:
d8fd00d82f1b7e030a4e1f8a27ada861d209dc8f2a6775e126a77d042faf3600 - Sigstore transparency entry: 2492801841
- Sigstore integration time:
-
Permalink:
YashSaxena21/IMRNNs@b9cf6b96c67d8485b942f9bcb7a8787c2ba8d40b -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/YashSaxena21
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@b9cf6b96c67d8485b942f9bcb7a8787c2ba8d40b -
Trigger Event:
push
-
Statement type:
File details
Details for the file imrnns-0.2.0-py3-none-any.whl.
File metadata
- Download URL: imrnns-0.2.0-py3-none-any.whl
- Upload date:
- Size: 41.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
374af589b1b031f2fcfe2c2c75faaa0b8a22f3f80df415d6a3c0aa9cc5d99584
|
|
| MD5 |
397b7322a11c67baeb48a2bfcfae0cf9
|
|
| BLAKE2b-256 |
44176d48301491d27919a2322436b017527a28ef023bc7a792016d00907b6214
|
Provenance
The following attestation bundles were made for imrnns-0.2.0-py3-none-any.whl:
Publisher:
publish.yml on YashSaxena21/IMRNNs
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
imrnns-0.2.0-py3-none-any.whl -
Subject digest:
374af589b1b031f2fcfe2c2c75faaa0b8a22f3f80df415d6a3c0aa9cc5d99584 - Sigstore transparency entry: 2492801874
- Sigstore integration time:
-
Permalink:
YashSaxena21/IMRNNs@b9cf6b96c67d8485b942f9bcb7a8787c2ba8d40b -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/YashSaxena21
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@b9cf6b96c67d8485b942f9bcb7a8787c2ba8d40b -
Trigger Event:
push
-
Statement type: