This release is a pre-release and may not be stable for production use.
EvalSuite
Unified, reproducible evaluation for machine learning and research.
EvalSuite brings classification and regression metrics (with clinical, statistical, segmentation and object-detection evaluation on the roadmap) into one consistent, validated, documented framework.
Status: in development (v0.1.0 in progress). The API may change before 0.1.0.
Installation
pip install evalsuite-python
The package is installed as evalsuite-python and imported as evalsuite:
import evalsuite as es
Why EvalSuite
- One consistent API. Every metric returns a result object that behaves like a number and exports to JSON, pandas, Markdown and LaTeX.
- Explicit conventions. Averaging, label order, the positive class and zero-division behaviour are stated and recorded in every result, never silently assumed.
- Validated. Each metric is tested against scikit-learn where definitions coincide, plus property-based tests and edge cases.
- Documented. Every metric carries its definition, formula, range, input requirements and references,
available programmatically through
metric_info(). - Efficient.
evaluate()validates inputs once and computes the confusion matrix once for all metrics. - Lightweight. Requires only NumPy, SciPy and pandas.
Quick start
import evalsuite as es
y_true = [0, 1, 1, 0, 1, 0]
y_pred = [0, 1, 0, 0, 1, 1]
y_prob = [0.1, 0.9, 0.4, 0.2, 0.8, 0.6]
result = es.evaluate(y_true, y_pred, y_prob=y_prob)
print(result.summary())
result["f1"] # MetricResult(f1=0.666667)
f"{result['mcc']:.3f}" # '0.333'
result.to_latex(caption="Test-set performance")
result.to_dataframe()
es.f1(y_true, y_pred) # individual metrics
es.roc_auc(y_true, y_prob)
es.metric_info("classification.mcc").formula # documentation
es.list_metrics("regression")
Metrics in this release
Classification (binary, multiclass, multilabel; micro/macro/weighted/samples/per-class averaging; sample weights): accuracy, balanced accuracy, precision, recall, specificity, NPV, F1, F-beta, Jaccard, MCC, Cohen's kappa (unweighted, linear, quadratic), Hamming loss, confusion matrix, ROC AUC (binary, one-vs-rest, one-vs-one), average precision, ROC and PR curves, log loss, Brier score, top-k accuracy.
Regression (single and multi-output; sample weights): MAE, MSE, RMSE, R², adjusted R², MAPE, sMAPE, MSLE, RMSLE, median absolute error, explained variance, max error, mean bias error, quantile (pinball) loss, Huber loss, relative absolute error, relative squared error.
Conventions
average="auto"resolves to"binary"for binary targets and"macro"otherwise; the resolved value is stored inresult.params["average"].- Labels are sorted unless you pass
labels=[...]; that order defines per-class outputs and the columns of 2-Dy_prob. - Undefined ratios (zero denominators) return 0 with an
UndefinedMetricWarning; passzero_division=np.nanto propagate NaN, or0/1to choose silently. - Domain violations raise clear errors instead of being patched over (for example MAPE with zero targets).
Development
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest --cov=evalsuite
ruff check . && ruff format --check . && mypy
Links
- PyPI: https://pypi.org/project/evalsuite-python/
- Website and documentation: https://evalsuite-nine.vercel.app
- Website source: https://github.com/mkcs28/evalsuite
License
MIT. See LICENSE.
Metadata
Release files for evalsuite-python 0.1.0a1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| evalsuite_python-0.1.0a1.tar.gz | 38.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| evalsuite_python-0.1.0a1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 74.5 kB
Release files / evalsuite_python-0.1.0a1.tar.gz
| Download URL | evalsuite_python-0.1.0a1.tar.gz |
|---|---|
| Size | 38.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
702bc094640bb5e58008f4af9b55b7070081e3a222e8dc8b956df0d73941c0e5
|
|
BLAKE2b-256 checksum How to use checksums |
3d7df2dd787efdcf781e00e92fb1717bf0abf6c39ca8ab9664f2b67cdf3631df
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency logRelease files / evalsuite_python-0.1.0a1-py3-none-any.whl
| Download URL | evalsuite_python-0.1.0a1-py3-none-any.whl |
|---|---|
| Size | 36.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4596815a32474c4ec2f0e72718a18236172e31fdb1c6be33cf5e64454eb9a1db
|
|
BLAKE2b-256 checksum How to use checksums |
694588827cfb41e0c716097afce1f590332bee4a95daf34a54257b09cb20f3fb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency log