🤗 Evaluate is a library that makes evaluating and comparing models and reporting their performance easier and more standardized.
It currently contains:
- implementations of dozens of popular metrics: the existing metrics cover a variety of tasks spanning from NLP to Computer Vision, and include dataset-specific metrics for datasets. With a simple command like
accuracy = load("accuracy"), get any of these metrics ready to use for evaluating a ML model in any framework (Numpy/Pandas/PyTorch/TensorFlow/JAX). - comparisons and measurements: comparisons are used to measure the difference between models and measurements are tools to evaluate datasets.
- an easy way of adding new evaluation modules to the 🤗 Hub: you can create new evaluation modules and push them to a dedicated Space in the 🤗 Hub with
evaluate-cli create [metric name], which allows you to see easily compare different metrics and their outputs for the same sets of references and predictions.
🔎 Find a metric, comparison, measurement on the Hub
🤗 Evaluate also has lots of useful features like:
- Type checking: the input types are checked to make sure that you are using the right input formats for each metric
- Metric cards: each metrics comes with a card that describes the values, limitations and their ranges, as well as providing examples of their usage and usefulness.
- Community metrics: Metrics live on the Hugging Face Hub and you can easily add your own metrics for your project or to collaborate with others.
Installation
With pip
🤗 Evaluate can be installed from PyPi and has to be installed in a virtual environment (venv or conda for instance)
pip install evaluate
Usage
🤗 Evaluate's main methods are:
evaluate.list_evaluation_modules()to list the available metrics, comparisons and measurementsevaluate.load(module_name, **kwargs)to instantiate an evaluation moduleresults = module.compute(*kwargs)to compute the result of an evaluation module
Adding a new evaluation module
First install the necessary dependencies to create a new metric with the following command:
pip install evaluate[template]
Then you can get started with the following command which will create a new folder for your metric and display the necessary steps:
evaluate-cli create "Awesome Metric"
See this step-by-step guide in the documentation for detailed instructions.
Credits
Thanks to @marella for letting us use the evaluate namespace on PyPi previously used by his library.
Metadata
Release files for evaluate 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| evaluate-0.2.1.tar.gz | 56.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| evaluate-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 126.5 kB
Release files / evaluate-0.2.1.tar.gz
| Download URL | evaluate-0.2.1.tar.gz |
|---|---|
| Size | 56.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e7e3a01d870579eecae33aeff339f77bbaf611460a89a31d726a28a6d67334da
|
|
BLAKE2b-256 checksum How to use checksums |
48bb392c27bc0e6cc986e5fd14ddfe6064e4eae8d4aa6e07183e2eace841fb84
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.1 CPython/3.9.13
|
Release files / evaluate-0.2.1-py3-none-any.whl
| Download URL | evaluate-0.2.1-py3-none-any.whl |
|---|---|
| Size | 69.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1df4a61b566276f6f3aaadd47237f44629277955350cc962f220efe1f9f9dbce
|
|
BLAKE2b-256 checksum How to use checksums |
97ef8be1efcaa0d1491466db85ca8251a4055cebfb2014a1f7331687d7fbfa03
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.1 CPython/3.9.13
|