llama-index-retrievers-whoosh
A pure-Python BM25 (lexical) retriever for LlamaIndex, powered by Whoosh.
LlamaIndex pipelines usually reach for a vector index, but dense retrieval has a
well-known blind spot: it can quietly miss the exact tokens that matter most —
product SKUs, function names, error codes like ERR_2043, gene symbols, ticket
IDs. A lexical BM25 retriever is the classic complement, and Whoosh gives you one
in pure Python: no server, no native wheels, and an index that is just a
folder on disk.
Install
pip install llama-index-retrievers-whoosh
This pulls in llama-index-core and whoosh3 (the maintained Whoosh fork).
Quick start
from llama_index.retrievers.whoosh import WhooshRetriever
retriever = WhooshRetriever.from_texts(
texts=[
"Whoosh is a pure-Python full-text search library.",
"BM25 ranks documents by term rarity and frequency.",
],
ids=["a", "b"],
metadatas=[{"src": "readme"}, {"src": "docs"}],
k=4,
)
nodes = retriever.retrieve("pure python search")
for n in nodes:
print(n.score, n.node.metadata["id"], n.node.text)
Every result is a standard LlamaIndex NodeWithScore, so it drops straight into
any query engine or router.
Persist to disk
Pass a path to build an on-disk index once, then reopen it later:
WhooshRetriever.from_texts(texts=texts, ids=ids, path="./whoosh_index")
retriever = WhooshRetriever.from_index("./whoosh_index", k=8)
The index is just a directory — copy it, commit it, ship it in a container.
Hybrid (lexical + vector) search
Combine this retriever with any vector retriever using LlamaIndex's
QueryFusionRetriever, which does Reciprocal Rank Fusion for you:
from llama_index.core.retrievers import QueryFusionRetriever
fusion = QueryFusionRetriever(
[whoosh_retriever, vector_retriever],
similarity_top_k=8,
num_queries=1, # set >1 to also fuse query rewrites
mode="reciprocal_rerank",
)
nodes = fusion.retrieve("ERR_2043 timeout after upgrade")
Lexical retrieval catches the exact ERR_2043 token; the vector retriever
catches the paraphrases. Fusing them is consistently stronger than either alone.
Why Whoosh?
- Pure Python — no Java, no C extensions, no server to run.
- BM25F ranking out of the box.
- The index is a folder — trivial to build in CI, cache, or ship.
License
BSD-2-Clause, matching Whoosh. See the Whoosh repository for the underlying library and its history.
Metadata
Release files for llama-index-retrievers-whoosh 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llama_index_retrievers_whoosh-0.1.1.tar.gz | 4.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llama_index_retrievers_whoosh-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 9.0 kB
Release files / llama_index_retrievers_whoosh-0.1.1.tar.gz
| Download URL | llama_index_retrievers_whoosh-0.1.1.tar.gz |
|---|---|
| Size | 4.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f096dbc3600249edfa6c7ceb6da22060eeb835b50823c06e5ceaff342284622f
|
|
BLAKE2b-256 checksum How to use checksums |
eb0e4210528738f0237b66fbed6c6e5e5612718bce14a4400773cad14beb54b8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / llama_index_retrievers_whoosh-0.1.1-py3-none-any.whl
| Download URL | llama_index_retrievers_whoosh-0.1.1-py3-none-any.whl |
|---|---|
| Size | 4.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d1d5057d97a59487a68fe23c151260282125984f348fa80eb9bb44da5c59afbf
|
|
BLAKE2b-256 checksum How to use checksums |
de7b7d3c2654379d22a039dbed41f0145d89fef56d1e9e283cdf20fd02b62463
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|