birt-gd
BIRT implements β³-IRT and β⁴-IRT (Beta Item Response Theory) fit by gradient descent in TensorFlow. Unlike classic IRT, which models binary correct/incorrect responses, Beta-IRT models a continuous response pij ∈ (0, 1) — e.g. the probability that classifier/respondent j assigns to the correct class of item i — which makes it well suited to evaluating and comparing probabilistic classifiers, not just human test-takers.
Table of contents
Background
Given a matrix X of response probabilities pij ∼ Β(αij, βij) — the probability of respondent j correctly classifying item i — the model estimates:
- ability (θi) per respondent
- difficulty (δj) per item
- discrimination (ωj / βj for β⁴-IRT, a single aj for β³-IRT) per item
using:
θi = σ(ti), δj = σ(dj), ωj = softplus(oj), βj = tanh(bj)
E[pij | θi, δj, ωj, βj] = 1 / (1 + (δj/(1-δj))ωjβj × (θi/(1-θi))-ωjβj)
β⁴-IRT (Beta4) fits this with unconstrained gradient descent (link functions remove the bounded-parameter symmetry problem β³-IRT has); set_priors=True (default) initializes abilities/difficulties from the data's own moments instead of random draws, which converges faster and more reliably. β³-IRT (Beta3) is the earlier, single-discrimination-parameter model — kept for comparison/backwards compatibility. See Citation for the papers behind both.
Installation
pip install birt-gd
Requirements
- Python >= 3.10
- tensorflow ^2.18.0
- pandas ^2.2.3
- scikit-learn ^1.6.1
- matplotlib ^3.10.0
- seaborn ^0.13.2
- tqdm ^4.67.1
Usage
from birt import Beta4
import pandas as pd
data = pd.DataFrame({
'a': [0.99, 0.89, 0.87, 0.50],
'b': [0.32, 0.25, 0.45, 0.20],
'c': [0.50, 0.50, 0.50, 0.50],
})
bgd = Beta4(
learning_rate=1,
epochs=10000,
n_respondents=data.shape[1],
n_items=data.shape[0],
n_inits=1000,
random_seed=1,
tol=10**(-5),
set_priors=True,
)
bgd.fit(data.values)
bgd.abilities # array([0.626, 0.416, 0.474], dtype=float32)
bgd.difficulties # array([0.456, 0.478, 0.442, 0.608], dtype=float32)
bgd.discriminations # array([0.992, 1.000, 0.961, 0.792], dtype=float32)
bgd.score # Pseudo-R2, e.g. 0.888
Beta3 shares the same interface (drop set_priors, since β³-IRT has a single discrimination parameter):
from birt import Beta3
b3 = Beta3(learning_rate=1, epochs=10000, n_respondents=data.shape[1], n_items=data.shape[0])
b3.fit(data.values)
Summary
bgd.summary()
ESTIMATES
-----
| Min 1Qt Median 3Qt Max Std.Dev
Ability | 0.00012 0.21369 0.57847 0.69513 0.93050 0.33468
Difficulty | 0.03876 0.27725 0.58860 0.84598 0.96604 0.30748
Discrimination | 0.25266 0.73648 1.04295 1.35130 2.09018 0.47445
pij | 0.00000 0.04613 0.40412 0.81140 0.99958 0.36590
-----
Pseudo-R2 | 0.88788
Plots
bgd.plot(xaxis=..., yaxis=..., ann=True, kwargs={'color': 'red'}) — scatter of any pair among discrimination, difficulty, ability, average_response, average_item.
bgd.boxplot(x=..., y=..., kwargs={...}) — boxplot of ability, difficulty or discrimination.
More end-to-end examples: example/00_example.ipynb.
Development
git clone https://github.com/Manuelfjr/birt-gd
cd birt-gd
poetry install
poetry shell
mc_analysis/ holds the Monte Carlo simulation study used to validate the model; it ships with the repo but not with the PyPI package.
Contributing
Issues and pull requests are welcome at github.com/Manuelfjr/birt-gd. There's no test suite yet, so please describe how a change was verified (e.g. output of the Usage example) in the PR description.
Citation
birt-gd is the reference implementation for the following papers — please cite the one matching the model you use (Beta4 → β⁴-IRT, Beta3 → β³-IRT):
@article{ferreirajunior2023beta4irt,
title = {{$\beta^4$-IRT}: A New {$\beta^3$-IRT} with Enhanced Discrimination Estimation},
author = {Ferreira-Junior, Manuel and Reinaldo, Jessica T. S. and Silva Filho, Telmo M. and Lima Neto, Eufrasio A. and Prudencio, Ricardo B. C.},
journal = {arXiv preprint arXiv:2303.17731},
year = {2023}
}
@inproceedings{chen2019beta3irt,
title = {{$\beta^3$-IRT}: A New Item Response Model and its Applications},
author = {Chen, Yu and Silva Filho, Telmo and Prudencio, Ricardo B. C. and Diethe, Tom and Flach, Peter},
booktitle = {Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS)},
year = {2019}
}
Support
- E-mail: ferreira.jr.ufpb@gmail.com
- Site: manuelfjr.github.io
License
GNU General Public License v3.0 © Manuel Ferreira Junior
Author
Manuel Ferreira Junior |
Contributors
Telmo de Menezes e Silva Filho |
Peter Flach |
Ricardo Prudêncio |
Eufrásio de Andrade Lima Neto |
Metadata
Release files for birt-gd 0.1.50
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| birt_gd-0.1.50.tar.gz | 26.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| birt_gd-0.1.50-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 49.4 kB
Release files / birt_gd-0.1.50.tar.gz
| Download URL | birt_gd-0.1.50.tar.gz |
|---|---|
| Size | 26.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a629d9c126abef9aaebde57bcf2ad4ee2a5ecd0d76c2105cccb9b30c68b38fc1
|
|
BLAKE2b-256 checksum How to use checksums |
b5b5d43888b8709546eded04f6b2b847ee4cfd445fb612286e385eeb15ebc58c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.0
|
Release files / birt_gd-0.1.50-py3-none-any.whl
| Download URL | birt_gd-0.1.50-py3-none-any.whl |
|---|---|
| Size | 22.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e5dcf70acb192777a57228bcfd7750dda578ac4e7ce2c9669b0a90a62e753589
|
|
BLAKE2b-256 checksum How to use checksums |
6b3196b9be562d540f14c22f2da129ad3d2e8f382fe14be1f67fb83cc58ad5f8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.0
|