fitzyracing-rank-bm25 (Rank-BM25)
Part of 360 Bench, tested fixes for abandoned PyPI packages.
Rank-BM25: a two line search engine.
This is a fork of rank_bm25 by Dorian Brown and its contributors, published as a drop-in replacement. The upstream package (~8.7M downloads a month, and the engine behind LangChain's
BM25Retriever) has not been released since 0.2.2 (February 2022), and the fixes below have sat in open issues and pull requests since 2023–2024. The import name is stillrank_bm25, so no code changes are needed. All credit for the original library goes to its author and contributors; it remains available under the same Apache-2.0 license (see LICENSE and NOTICE). To cite it, use the upstream CITATION.If a fixed
rank-bm25release appears on PyPI, prefer it and switch back.
What's fixed in this fork
Based on upstream rank-bm25==0.2.2 (tag 0.2.2).
The API is unchanged. Scores are identical to 0.2.2 except in the one case described under
"Half-corpus terms" below.
-
Empty corpus: a clear error instead of
ZeroDivisionError: division by zero.BM25Okapi([])(andBM25L,BM25Plus, also with a tokenizer or an empty iterator) crashed deep inside_initialize, which is hard to understand when rank_bm25 sits a few layers down, for example in a RAG pipeline whose retriever got no documents. It now raisesrank_bm25.EmptyCorpusException: Cannot build a BM25 index from an empty corpus, as the maintainer proposed in #36 and PR #45. The exception subclasses bothValueErrorandZeroDivisionError, so existingexcept ZeroDivisionError:code still catches it. (Unlike #45 it is defined inrank_bm25itself rather than in a new top-levelexceptionsmodule, and it also works for iterable corpora, whichlen(corpus)would not.) -
Half-corpus terms:
BM25Okapino longer gives a term in exactly half the documents an idf of 0.BM25Okapiuses idf = log((N − n + 0.5) / (n + 0.5)), which is exactly 0 when a term is in n = N/2 of the N documents, and negative when it is in more than half. Negative values are replaced by the floorepsilon * average_idf, but 0 was not, so a term in exactly half the documents counted for less than a term in every document, and documents that matched it scored 0 (#39, #43):corpus = ["This text contains keyword1 and Keyword2", "That is a text that contains keyword1 and term1", "Page contains no keywords but contains term1 and term2", "This text contains no keywords"] BM25Okapi([d.split() for d in corpus]).get_scores("This is a question about keyword1 & term1".split()) # 0.2.2: array([0. , 1.52856224, 0. , 0. ]) # fork: array([0.09779242, 1.60992918, 0.04068347, 0.05242379])
The fork gives these terms the same
epsilon * average_idffloor that terms in more than half the documents already get. Exactly which scores change:- only
BM25Okapi(BM25LandBM25Plusare untouched); - only for corpora with an even number of documents N, and only for query terms that occur in exactly N/2 of them;
- only when the floor
epsilon * average_idfis positive (with the defaultepsilon=0.25, when the average idf of the vocabulary is positive, which is the normal case). If it is 0 or negative, which happens in tiny corpora such as 2 documents, these terms keep idf 0, exactly as before, because a negative floor would rank matching documents below non-matching ones; - in those cases, documents containing such a term gain
epsilon * average_idf × (the usual BM25 term-frequency factor)per occurrence in the query; no score ever goes down, andbm25.average_idfis unchanged. Passingepsilon=0restores 0.2.2's results exactly.
A randomized comparison of 20,000 small corpora and queries against 0.2.2 (all three algorithms) found no other differences.
What this fork deliberately does not change: PR #40 switches
BM25Okapito a different idf formula, log((N + 1) / (n + 0.5)). That removes the problem for every corpus size, including 2 documents, but changes everyBM25Okapiscore, so it is not a drop-in fix. In a 2-document corpus a term in one document therefore still scores 0 here; if you need that case, useBM25PlusorBM25L, whose idf is always positive. - only
-
Builds from source again.
setup.pyimported a localversion.pythat wasn't shipped in the sdist, so building from the sdist under PEP 517 failed (#56, PR #57). The fork uses a staticpyproject.toml.
Left out on purpose, to keep the fork a drop-in replacement:
- Upstream
masterhas two commits after 0.2.2 that were never released: #20 removes aq_freqfactor from theBM25Lscore (a real correction, but it changes everyBM25Lscore) and #23 rewrites onelog()inBM25Plus(can change the last digit of scores). Neither is included. - PR #58 (apply the constructor tokenizer
to string queries in
get_top_n, #38) is a behaviour change the maintainer wanted to hold back on, so it is not included. Tokenize your queries yourself, as the README below says.
Install
pip install fitzyracing-rank-bm25
Python 3.8+; the only dependency is numpy, as before.
Switching from rank-bm25
This distribution installs the same rank_bm25 module as the original, so your code keeps
doing from rank_bm25 import BM25Okapi. Uninstall the original first, then install the
fork:
pip uninstall -y rank-bm25
pip install fitzyracing-rank-bm25
The order matters. Both distributions own the same rank_bm25.py, so if you install the fork
first and uninstall rank-bm25 afterwards, pip deletes the shared file and the import breaks.
If that happens, run pip install --force-reinstall --no-deps fitzyracing-rank-bm25.
In requirements.txt / pyproject.toml, replace rank-bm25 (or rank_bm25) with
fitzyracing-rank-bm25.
If you get rank-bm25 through another package
pip cannot replace a dependency with a differently named package. If a dependency requires
rank-bm25, install the fork alongside it and then remove the original's files:
pip install fitzyracing-rank-bm25
pip uninstall -y rank-bm25
pip install --force-reinstall --no-deps fitzyracing-rank-bm25 # restore the file the uninstall removed
Afterwards pip check reports <package> requires rank-bm25, which is not installed; that is
expected. Repeat the steps if a later install pulls rank-bm25 back in.
uv users can do this properly with an override that drops the original:
# pyproject.toml
[project]
dependencies = ["fitzyracing-rank-bm25", "...the package that depends on rank-bm25..."]
[tool.uv]
override-dependencies = ["rank-bm25; sys_platform == 'never'"]
(or uv pip install --override overrides.txt ... with that same line in overrides.txt).
Tests
pip install -e ".[dev]"
pytest
tests/test_fork_fixes.py fails on rank-bm25 0.2.2 and passes on this fork; it also pins
scores that must stay identical to 0.2.2.
About Rank-BM25
A collection of algorithms for querying a set of documents and returning the ones most relevant to the query. The most common use case for these algorithms is, as you might have guessed, to create search engines.
So far the algorithms that have been implemented are:
- Okapi BM25
- BM25L
- BM25+
- BM25-Adpt
- BM25T
These algorithms were taken from this paper, which gives a nice overview of each method, and also benchmarks them against each other. A nice inclusion is that they compare different kinds of preprocessing like stemming vs no-stemming, stopword removal or not, etc. Great read if you're new to the topic.
Usage
For this example we'll be using the BM25Okapi algorithm, but the others are used in pretty much the same way.
Initalizing
First thing to do is create an instance of the BM25 class, which reads in a corpus of text and does some indexing on it:
from rank_bm25 import BM25Okapi
corpus = [
"Hello there good man!",
"It is quite windy in London",
"How is the weather today?"
]
tokenized_corpus = [doc.split(" ") for doc in corpus]
bm25 = BM25Okapi(tokenized_corpus)
# <rank_bm25.BM25Okapi at 0x1047881d0>
Note that this package doesn't do any text preprocessing. If you want to do things like lowercasing, stopword removal, stemming, etc, you need to do it yourself.
The only requirements is that the class receives a list of lists of strings, which are the document tokens.
Ranking of documents
Now that we've created our document indexes, we can give it queries and see which documents are the most relevant:
query = "windy London"
tokenized_query = query.split(" ")
doc_scores = bm25.get_scores(tokenized_query)
# array([0. , 0.93729472, 0. ])
Good to note that we also need to tokenize our query, and apply the same preprocessing steps we did to the documents in order to have an apples-to-apples comparison
Instead of getting the document scores, you can also just retrieve the best documents with
bm25.get_top_n(tokenized_query, corpus, n=1)
# ['It is quite windy in London']
And that's pretty much it!
Metadata
Release files for fitzyracing-rank-bm25 0.2.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| fitzyracing_rank_bm25-0.2.3.tar.gz | 15.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| fitzyracing_rank_bm25-0.2.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 28.6 kB
Release files / fitzyracing_rank_bm25-0.2.3.tar.gz
| Download URL | fitzyracing_rank_bm25-0.2.3.tar.gz |
|---|---|
| Size | 15.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
65dd0931e33716a5b753fd12108562af38a40023b617787ec0af61122e92913d
|
|
BLAKE2b-256 checksum How to use checksums |
9877286e3a1f0bf1b13b1fe7faca6d501ba40ba49e438099e0ce64a7fb792f8a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.5
|
Release files / fitzyracing_rank_bm25-0.2.3-py3-none-any.whl
| Download URL | fitzyracing_rank_bm25-0.2.3-py3-none-any.whl |
|---|---|
| Size | 13.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4f78e6e9e0ab8e1393b97cfe2a5c37449896f52d2f71e76072ddeec4b06c6816
|
|
BLAKE2b-256 checksum How to use checksums |
b5e564921411f0758f504fee4099f973cff18ff9722a48e1d809f4e39f9063b9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.5
|