Compute rankings in Python.

These details have not been verified by PyPI

Project links

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Project description

Ranky

Compute rankings in Python.

logo

Get started

pip install ranky

import ranky as rk

Read the documentation.

Main functions

The main functionalities include scoring metrics (e.g. accuracy, roc auc), rank metrics (e.g. Kendall Tau, Spearman correlation), ranking systems (e.g. Majority judgement, Kemeny-Young method) and some measurements (e.g. Kendall's W coefficient of concordance).

Most functions takes as input 2-dimensional numpy.array or pandas.DataFrame objects. DataFrame are the best to keep track of the names of each data point.

Let's consider the following preference matrix:

matrix

Each row is a candidate and each column is a judge. Here is the results of rk.borda(matrix), computing the mean rank of each candidate:

borda

We can see that candidate2 has the best average rank among the four judges.

Let's display it using rk.show(rk.borda(matrix)):

display

Ranking systems

The rank aggregation methods available include:

Random Dictator: rk.dictator(m)
Score Voting (mean): rk.score(m)
Borda Count (average rank): rk.borda(m)
Majority Judgement (median): rk.majority(m)
Pairwise methods. Copeland's method: rk.pairwise(m), Success rate: rk.pairwise(m, wins=rk.success_rate) and more. You can specify your own "wins" function or select one from the rk.duel module.
Optimal rank aggregation using any rank metric: rk.center(m), rk.center(m, method='kendalltau'). Solver used [1].
(Kemeny-Young method is optimal rank aggregation using Kendall's tau as metric.)
(Optimal rank aggregation using Spearman correlation as metric is equivalent to Borda count.)

Metrics

You can use any_metric(a, b, method) to call a metric from any of the three categories below.

Scoring metrics: rk.metric(y_true, y_pred, method='accuracy'). Methods include: ['accuracy', 'balanced_accuracy', 'precision', 'average_precision', 'brier', 'f1_score', 'mxe', 'recall', 'jaccard', 'roc_auc', 'mse', 'rmse', 'sar']
Rank correlation coefficients: rk.corr(r1, r2, method='spearman'). Methods include: ['kendalltau', 'spearman', 'pearson']
Rank distances: rk.dist(r1, r2, method='levenshtein'). Methods include: ['hamming', 'levenshtein', 'kendall', 'winner', 'euclidean']

To add: general edit distances, kemeny distance, regression metrics...

Visualizations

Use rk.show to visualize preference matrix (2D) or ranking ballots (1D).

>>> rk.show(m)

show example 1

>>> rk.show(m['judge1'])

show example 2

Use rk.mds, to visualize (in 2D or 3D) the points in a given metric space. See rk.scatterplot documentation for display arguments.

>>> rk.mds(m, method='euclidean')

MDS example 1

>>> rk.mds(m, method='spearman', axis=1)

MSE example 2

You can use rk.tsne similarly to rk.mds.
Use rk.critical_difference to plot a critical difference diagram, comparing candidates' performance and grouping them by statistical equivalence. Such diagrams can be seen in [2, 3].

>>> rk.critical_difference(m, comparison_func=rk.bayes_wins)

Critical difference example

Other

Rank, rk.rank, convert a 1D score ballot into a ranking.
Bootstrap, rk.bootstrap, sample a given axis.
Consensus, rk.consensus, check if ranking exactly agree.
Concordance, ,rk.concordance, mean rank distance between all judges of a preference matrix.
Centrality, rk.centrality, mean rank distance between a ranking and a preference matrix.
Kendall's W, rk.kendall_w, coefficient of concordance.
Utility: read_codalab_csv to parse a CSV generated by Codalab representing a leaderboard into a pandas.DataFrame.

References

Please cite ranky in your publications if this is useful for your research. Here is an example BibTeX entry:

@misc{pavao2020ranky,
  title={ranky},
  author={Adrien Pavao},
  year={2020},
  howpublished={\url{https://github.com/didayolo/ranky}},
}

[1] Storn R. and Price K., Differential Evolution - a Simple and Efficient Heuristic for Global Optimization over Continuous Spaces, Journal of Global Optimization, 1997, 11, 341 - 359.

[2] Janez Demsar, Statistical Comparisons of Classifiers over Multiple Data Sets, 7(Jan):1--30, 2006.

[3] H. Ismail Fawaz, G. Forestier, J. Weber, L. Idoumghar, P. Muller, Deep learning for time series classification: a review, Data Mining and Knowledge Discovery, 2018.

License

The text of the Apache License 2.0 can be found online at: http://www.opensource.org/licenses/apache2.0.php

Project details

These details have not been verified by PyPI

Project links

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Release history Release notifications | RSS feed

This version

0.3.4

Jul 19, 2023

0.3.3

Mar 24, 2023

0.3.2

Jan 10, 2023

0.3.1

Jan 10, 2023

0.3.0

Jan 5, 2023

0.2.9

Jul 13, 2022

0.2.8

Jun 30, 2022

0.2.7

Mar 2, 2022

0.2.6

Feb 22, 2022

0.2.5

Feb 21, 2022

0.2.4

Jan 21, 2022

0.2.3

Sep 9, 2021

0.2.2

May 14, 2021

0.2.1

Apr 30, 2021

0.1.7

Apr 30, 2021

0.1.5

Apr 30, 2021

0.1.4

Apr 30, 2021

0.1.3

Apr 22, 2021

0.1.2

Apr 22, 2021

0.1.1

Apr 22, 2021

0.1.0

Apr 17, 2021

0.0.9

Apr 1, 2021

0.0.8

Mar 18, 2021

0.0.7

Feb 19, 2021

0.0.6

Feb 8, 2021

0.0.5

Dec 8, 2020

0.0.4

Dec 8, 2020

0.0.3

Dec 6, 2020

0.0.2

Dec 5, 2020

0.0.1

Dec 5, 2020

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

ranky-0.3.4-py3-none-any.whl (25.6 kB view hashes)

Uploaded Jul 19, 2023 Python 3

Hashes for ranky-0.3.4-py3-none-any.whl

Hashes for ranky-0.3.4-py3-none-any.whl
Algorithm	Hash digest
SHA256	`5a0a7eaf8aa5c0e23f59f7be743615d57b71d7fcc997d98ccd06e07fc7eca1e3`
MD5	`54c7a9c5bfc7c166ddb6df7411487ca4`
BLAKE2b-256	`493d055eb088977a421a2580b233763236f83253c113b2131fb92a76912e1a0c`