flyai-bench
An evaluation framework for LLM agents. It loads evaluation tasks from a
HuggingFace dataset, runs an agent plus verifier inside isolated Docker
containers, and produces standardized scores (reward in the range 0.00 – 1.00).
Prerequisites
- Python >= 3.10
- Docker (with access to the image registry that hosts the benchmark images;
the target platform is
linux/amd64) - An OpenAI-compatible LLM endpoint for the agent, and one for the judge/verifier (they can be the same endpoint)
Installation
# From a clone of this repository
pip install -e .
# Or install the runtime dependencies directly
pip install pyyaml huggingface_hub openai
Installing the package exposes a flyai-bench console command. The examples
below also work by invoking run_eval.py directly.
Quick start
# 1. Copy the default config and fill in your values
cp eval_config.yaml eval_config.local.yaml
# edit eval_config.local.yaml: set the image registry, the LLM/judge
# base URLs, API keys, and the model name
# 2. Run the evaluation
python run_eval.py run --config eval_config.local.yaml
# 3. Inspect results
python run_eval.py report --config eval_config.local.yaml
Supported benchmarks
| Benchmark | Dataset | Tasks | Description |
|---|---|---|---|
| ecommerce_last_exam | FlyaiLab/ecommerce_last_exam | 1000 | Travel + e-commerce agent tool-use evaluation (500 travel, 500 e-commerce) |
Configuration
Key fields in eval_config.yaml (copy to eval_config.local.yaml before editing):
| Field | Meaning |
|---|---|
dataset.repo_id / dataset.config / dataset.split |
HuggingFace dataset, config (travel / e_commerce), and split |
docker.registry |
Image registry that hosts the benchmark images. Set to your own registry, or override at runtime with --registry. Leave empty to use image names as-is. |
docker.platform |
Container platform (default linux/amd64) |
agent.cmd |
Command that runs the agent inside the container |
agent.llm_base_url / agent.llm_api_key / agent.llm_model |
OpenAI-compatible endpoint, key, and model for the agent |
verifier.judge_base_url / judge_api_key / judge_model |
Endpoint, key, and model for the judge/verifier |
runner.concurrency / runner.limit / runner.skip_done |
Parallelism, cap on number of instances (null = all), and whether to skip instances that already have a reward |
API keys are read from your local config and injected into the containers as environment variables. Never commit
eval_config.local.yaml— it is already in.gitignore.
Evaluation flow
HuggingFace Dataset
│
▼
┌─────────────────────────────────┐
│ For each instance: │
│ 1. docker pull <image> │
│ 2. docker run (start env) │
│ 3. agent calls tools -> answer │
│ 4. verifier (test.sh) scores │
│ 5. emit reward.txt (0 – 1) │
└─────────────────────────────────┘
│
▼
scores.jsonl + summary.json
CLI commands
# Run the evaluation
python run_eval.py run [--dataset-config travel|e_commerce] [--limit N] [--dry-run]
# Check progress
python run_eval.py status
# Generate a report
python run_eval.py report
# Package results for leaderboard submission
python run_eval.py submit --model deepseek-v4-flash --provider deepseek
Output
| File | Contents |
|---|---|
scores.jsonl |
One line per instance: {instance_id, domain, reward, duration_sec} |
summary.json |
Aggregate statistics (avg_reward, breakdown by_domain) |
reward is a float in 0.00 – 1.00 produced by the verifier for each instance.
Submitting results to the leaderboard
-
After evaluating, package the results with
submit:python run_eval.py submit \ --dataset-config travel \ --model deepseek-v4-flash \ --provider deepseek \ --agent-type mini-swe-agent
-
Review
experiments/evaluation/travel/<slug>/metadata.yamland complete the model and agent details. -
Open a pull request to the
mainbranch of the flyai-bench repository. -
CI validates the format automatically; once merged, the leaderboard updates.
Submission package layout
evaluation/travel/20260830_deepseek-v4-flash_mini-swe-agent/
├── metadata.yaml # model/agent info + evaluation statistics
├── scores.jsonl # per instance: {instance_id, domain, reward, duration_sec}
└── summary.json # aggregate statistics (avg_reward, by_domain)
Writing a custom agent
Implement a script that, inside the container:
- Reads
/app/tool_defs.jsonfor the tool definitions - Reads
/app/system.md+/app/instruction.mdfor the task - Calls
/app/tools/<name> --arg valto execute a tool - Writes the result to
/app/answer.json
Point the runner at your agent with --agent-cmd:
python run_eval.py run --agent-cmd "python /app/my_agent.py"
agent.py (a minimal LLM agent) and mini_swe_agent.py (a terminal/bash agent)
are provided as reference implementations.
Project structure
flyai-bench/
├── README.md
├── pyproject.toml # packaging (installs the `flyai-bench` CLI)
├── run_eval.py # evaluation CLI (run / status / report / submit)
├── eval_config.yaml # default config (copy to eval_config.local.yaml)
├── benchmark.yaml # benchmark metadata
├── agent.py # reference LLM agent
├── mini_swe_agent.py # reference terminal agent
├── tool_server.py # in-sandbox tool proxy (permission isolation)
├── sandbox_setup.sh # container permission setup
├── validate_submission.py # submission validation + leaderboard rebuild
├── experiments/ # submitted results + leaderboard.json
├── leaderboard_space/ # HuggingFace Space (leaderboard UI)
└── src/flyai_bench/ # installable package mirror of the above
Note: the top-level scripts and the
src/flyai_bench/package currently hold parallel copies of the same code. Prefer editing one and keeping them in sync (or consolidating on the package) to avoid drift.
License
Released under the MIT license. See LICENSE for details.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file flyai_bench-0.1.0.tar.gz.
File metadata
- Download URL: flyai_bench-0.1.0.tar.gz
- Upload date:
- Size: 44.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1b1a34f9888fd77cbd27efa6999d798d1d4d5fa9f64198a656078e4a6ae772ea
|
|
| MD5 |
e6a7fad43539f825924e62932f596afe
|
|
| BLAKE2b-256 |
a9fd0172f255e875c97b841fe3b2581c101ac88ca92f964717024e2066a71ad2
|
File details
Details for the file flyai_bench-0.1.0-py3-none-any.whl.
File metadata
- Download URL: flyai_bench-0.1.0-py3-none-any.whl
- Upload date:
- Size: 25.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1ebcd4e86c69f114f1bf0b51566515d32dbb862344a3d8b22750e5748cb254a5
|
|
| MD5 |
940347f888f54eced511996a04f227f8
|
|
| BLAKE2b-256 |
5a4033175132870172712ec452f9014a6592a9be2044cdb11a32e87c14b69d2c
|