Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

EvalSuite

CI PyPI Python License: MIT

Unified, reproducible evaluation for machine learning and research.

EvalSuite brings classification and regression metrics (with clinical, statistical, segmentation and object-detection evaluation on the roadmap) into one consistent, validated, documented framework.

Status: beta (0.1.0b2). Feature-complete for 0.1.0; the stable release follows once this beta is verified.

Installation

pip install evalsuite-python

The package is installed as evalsuite-python and imported as evalsuite:

import evalsuite as es

Why EvalSuite

  • One consistent API. Every metric returns a result object that behaves like a number and exports to JSON, pandas, Markdown and LaTeX.
  • Explicit conventions. Averaging, label order, the positive class and zero-division behaviour are stated and recorded in every result, never silently assumed.
  • Validated. Each metric is tested against scikit-learn where definitions coincide, plus property-based tests and edge cases.
  • Documented. Every metric carries its definition, formula, range, input requirements and references, available programmatically through metric_info().
  • Efficient. evaluate() validates inputs once and computes the confusion matrix once for all metrics.
  • Lightweight. Requires only NumPy, SciPy and pandas.

Quick start

import evalsuite as es

y_true = [0, 1, 1, 0, 1, 0]
y_pred = [0, 1, 0, 0, 1, 1]
y_prob = [0.1, 0.9, 0.4, 0.2, 0.8, 0.6]

result = es.evaluate(y_true, y_pred, y_prob=y_prob)
print(result.summary())

result["f1"]  # MetricResult(f1=0.666667)
f"{result['mcc']:.3f}"  # '0.333'
result.to_latex(caption="Test-set performance")
result.to_dataframe()

es.f1(y_true, y_pred)  # individual metrics
es.roc_auc(y_true, y_prob)
es.metric_info("classification.mcc").formula  # documentation
es.list_metrics("regression")

Comparing models

result = es.compare(
    y_true,
    {"logistic": pred_lr, "forest": pred_rf, "boosting": pred_gb},
    probabilities={"logistic": prob_lr, "forest": prob_rf, "boosting": prob_gb},
    random_state=0,
)
print(result.summary())  # estimates with 95% CIs, paired tests, Holm-adjusted p-values
result.to_latex(label="tab:models")

es.bootstrap_ci("f1", y_true, y_pred, average="macro", random_state=0)  # BCa interval for any metric
es.accuracy_ci(y_true, y_pred)  # Wilson interval
es.delong_test(y_true, prob_a, prob_b)  # two correlated AUCs
es.mcnemar_test(y_true, pred_a, pred_b)

Every model is evaluated on the same bootstrap resamples, so differences are paired. Accuracy is compared with McNemar's test, binary ROC AUC with DeLong's test and other metrics with a paired bootstrap test; p-values are adjusted for multiple comparisons (Holm by default).

Classification report

report = es.classification_report(y_true, y_pred)
print(report)  # per-class precision, recall, F1, specificity, support + averages
report.save("report.html")  # also .csv .md .tex .json .txt

Plots

pip install "evalsuite-python[plot]"   # adds matplotlib; importing evalsuite never loads it
es.plot.roc(y_true, {"logistic": prob_lr, "forest": prob_rf})  # AUC in the legend
es.plot.pr(y_true, prob)  # AP and the prevalence line
es.plot.calibration(y_true, prob)  # reliability diagram, ECE, Brier
es.plot.confusion_matrix(y_true, y_pred, normalize="true")
es.plot.residuals(y_reg, pred_reg)  # or kind="predicted"
es.plot.comparison(es.compare(...))  # forest plot with CIs

Each function returns a matplotlib Axes (pass ax= to draw into your own figure). The numbers shown are computed with EvalSuite's metrics, so plots and tables always agree. Several models get distinct colours and line styles, so figures stay readable in greyscale print.

Exports

Every result (evaluate, classification_report, compare, single metrics) exports to summary(), to_json(), to_csv(), to_dataframe(), to_markdown(), to_latex() and to_html(), and save(path) picks the format from the extension. HTML pages are standalone (inline CSS, no scripts) and escape all text.

Command line

evalsuite evaluate predictions.csv --y-true label --y-pred pred --y-prob prob
evalsuite report predictions.csv --y-true label --y-pred pred -o report.html
evalsuite compare predictions.csv --y-true label --pred lr=pred_lr --pred rf=pred_rf \
    --prob lr=p_lr --prob rf=p_rf --plot comparison.png
evalsuite plot roc predictions.csv --y-true label --y-prob prob -o roc.png
evalsuite metrics --category classification
evalsuite info classification.mcc
evalsuite benchmark --quick

Input files can be CSV, TSV, Parquet or JSON. Output format follows --format or the -o extension (text, json, csv, markdown, latex, html). Errors are reported in one line with exit code 2.

Performance

evaluate() validates inputs once and computes the confusion matrix once for all metrics: about 10× faster than the equivalent separate scikit-learn calls, with identical results. See BENCHMARKS.md.

Metrics in this release

Classification (binary, multiclass, multilabel; micro/macro/weighted/samples/per-class averaging; sample weights): accuracy, balanced accuracy, precision, recall, specificity, NPV, F1, F-beta, Jaccard, MCC, Cohen's kappa (unweighted, linear, quadratic), Hamming loss, confusion matrix, ROC AUC (binary, one-vs-rest, one-vs-one), average precision, ROC and PR curves, log loss, Brier score, top-k accuracy, calibration curve and expected calibration error.

Regression (single and multi-output; sample weights): MAE, MSE, RMSE, R², adjusted R², MAPE, sMAPE, MSLE, RMSLE, median absolute error, explained variance, max error, mean bias error, quantile (pinball) loss, Huber loss, relative absolute error, relative squared error.

Conventions

  • average="auto" resolves to "binary" for binary targets and "macro" otherwise; the resolved value is stored in result.params["average"].
  • Labels are sorted unless you pass labels=[...]; that order defines per-class outputs and the columns of 2-D y_prob.
  • Undefined ratios (zero denominators) return 0 with an UndefinedMetricWarning; pass zero_division=np.nan to propagate NaN, or 0/1 to choose silently.
  • Domain violations raise clear errors instead of being patched over (for example MAPE with zero targets).

Development

python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest --cov=evalsuite
ruff check . && ruff format --check . && mypy

License

MIT. See LICENSE.

Metadata

Release files for evalsuite-python 0.1.0b2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for evalsuite-python 0.1.0b2
File Size Uploaded
evalsuite_python-0.1.0b2.tar.gz 82.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for evalsuite-python 0.1.0b2
File Interpreter ABI Platform
evalsuite_python-0.1.0b2-py3-none-any.whl Python 3 none any Details

Total release size: 160.1 kB

Release files / evalsuite_python-0.1.0b2.tar.gz

Download URL evalsuite_python-0.1.0b2.tar.gz
Size 82.9 kB
Tags Source
SHA-256 checksum
How to use checksums
d51ac8c6c79d9cdbf5b605b7430910e7828057de1d6f2950710d277e3a576d2b
BLAKE2b-256 checksum
How to use checksums
fa29a2e667a6a7f4ca3724a290f7a95a6b601fb9547fbc6b00acc613449fa92f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / evalsuite_python-0.1.0b2-py3-none-any.whl

Download URL evalsuite_python-0.1.0b2-py3-none-any.whl
Size 77.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3a67e72ead78299ae0a2fd098c00c50655bc12332716e353e561c1815653e03f
BLAKE2b-256 checksum
How to use checksums
e5224c25d66e551a275eb8e37b6bedbcdbd87bc1f2bfa59cc24ae0e9e32448e7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page