Skip to main content

agent-eval

A utility for evaluating agents on a suite of Inspect-formatted evals, with the following primary benefits:

  1. Task suite specifications as config.
  2. Extracts the token usage of the agent from log files, and computes cost using litellm.
  3. Submits task suite results to a leaderboard, with submission metadata and easy upload to a HuggingFace repo for distribution of scores and logs.

Installation

To install from pypi, use pip install agent-eval.

For leaderboard extras, use pip install agent-eval[leaderboard].

Usage

Run evaluation suite

agenteval eval --config-path CONFIG_PATH --split SPLIT LOG_DIR

Evaluate an agent on the supplied eval suite configuration. Results are written to agenteval.json in the log directory.

See sample-config.yml for a sample configuration file.

For aggregation in a leaderboard, each task specifies a primary_metric as {scorer_name}/{metric_name}. The scoring utils will look for a corresponding stderr metric, by looking for another metric with the same scorer_name and with a metric_name containing the string "stderr".

Weighted Macro Averaging with Tags

Tasks can be grouped using tags for computing summary statistics. The tags support weighted macro averaging, allowing you to assign different weights to tasks within a tag group.

Tags are specified as simple strings on tasks. To adjust weights for specific tag-task combinations, use the macro_average_weight_adjustments field at the split level. Tasks not specified in the adjustments default to a weight of 1.0.

See sample-config.yml for an example of the tag and weight adjustment format.

Score results

agenteval score [OPTIONS] LOG_DIR

Compute scores for the results in agenteval.json and update the file with the computed scores.

Publish scores to leaderboard

agenteval lb publish [OPTIONS] LOG_DIR

Upload the scored results to HuggingFace datasets.

View leaderboard scores

agenteval lb view [OPTIONS]

View results from the leaderboard.

To save plots:

agenteval lb view --save-dir DIR [OPTIONS]

Administer the leaderboard

Prior to publishing scores, two HuggingFace datasets should be set up, one for full submissions and one for results files.

If you want to call load_dataset() on the results dataset (e.g., for populating a leaderboard), you probably want to explicitly tell HuggingFace about the schema and dataset structure (otherwise, HuggingFace may fail to propertly auto-convert to Parquet). This is done by updating the configs attribute in the YAML metadata block at the top of the README.md file at the root of the results dataset (the metadata block is identified by lines with just --- above and below it). This attribute should contain a list of configs, each of which specifies the schema (under the features key) and dataset structure (under the data_files key). See sample-config-hf-readme-metadata.yml for a sample metadata block corresponding to sample-comfig.yml (note that the metadata references the raw schema data, which must be copied).

To facilitate initializing new configs, agenteval lb publish will automatically add this metadata if it is missing.

Development

See Development.md for development instructions.

Metadata

Release files for agent-eval 0.1.55

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-eval 0.1.55
File Size Uploaded
agent_eval-0.1.55.tar.gz 51.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-eval 0.1.55
File Interpreter ABI Platform
agent_eval-0.1.55-py3-none-any.whl Python 3 none any Details

Total release size: 98.1 kB

Release files / agent_eval-0.1.55.tar.gz

Download URL agent_eval-0.1.55.tar.gz
Size 51.9 kB
Tags Source
SHA-256 checksum
How to use checksums
a61d8879da8338bcb30a8312f31d47c1bef89fbc90cf8819ead9d5cd44d3f965
BLAKE2b-256 checksum
How to use checksums
25696482ad46c4f6c4ab039bdcce26629d6a9bc957a108bdd863326f8fc1891a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.17

Release files / agent_eval-0.1.55-py3-none-any.whl

Download URL agent_eval-0.1.55-py3-none-any.whl
Size 46.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cdfa42dae19acd682cc85209fc0879c26642d194c0f259408c0ff6573ad30f89
BLAKE2b-256 checksum
How to use checksums
a4eec2e231a97d38b78d144a4c2f680cf69c548cf232d2b5bee15abc55638730
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.17

Release history Release notifications | RSS feed

This release

0.1.55 This release

2 release files

0.1.53

2 release files

0.1.52

2 release files

0.1.51

2 release files

0.1.50

2 release files

0.1.49

2 release files

0.1.48

2 release files

0.1.46

2 release files

0.1.45

2 release files

0.1.44

2 release files

0.1.43

2 release files

0.1.42

2 release files

0.1.41

2 release files

0.1.40

2 release files

0.1.39

2 release files

0.1.38

2 release files

0.1.37

2 release files

0.1.36

2 release files

0.1.35

2 release files

0.1.34

2 release files

0.1.33

2 release files

0.1.22

2 release files

0.1.21

2 release files

0.1.20

2 release files

0.1.19

2 release files

0.1.18

2 release files

0.1.17

2 release files

0.1.16

2 release files

0.1.15

2 release files

0.1.14

2 release files

0.1.10

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page