Flow Benchmark Tools
Create and run LLM benchmarks.
Installation
Just the library:
pip install flow-benchmark-tools:1.5.0
Library + Example benchmarks (see below):
pip install "flow-benchmark-tools[examples]:1.5.0"
Usage
-
Create an agent by inheriting BenchmarkAgent and implementing the
run_benchmark_casemethod. -
Create a Benchmark by compiling a list of BenchmarkCases. These can be read from a JSONL file.
-
Associate agent and benchmark in a BenchmarkRun.
-
Use a BenchmarkRunner to run your BenchmarkRun.
Running example RAG benchmarks
Two end-to-end benchmark examples are provided in the examples folder: a LangChain RAG application and an OpenAI Assistant agent.
To run the LangChain RAG benchmark:
python src/examples/langchain_rag_agent.py
To run the OpenAI Assistant benchmark:
python src/examples/openai_assistant_agent.py
The rag benchmark cases are defined in data/rag_benchmark.jsonl.
The two examples follow the typical usage pattern of the library:
- define an agent by implementing the BenchmarkAgent interface and overriding the
run_benchmark_casemethod (you can also override thebeforeandaftermethods, if needed), - create a set of benchmark cases, typically as a JSONL file such as data/rag_benchmark.jsonl,
- use a BenchmarkRunner to run the benchmark.
Running example criteria benchmark
An application of a criteria benchmark is also provided in examples folder: a Criteria application that assesses the quality of pre-computed LLM outputs based on the criteria defined in each benchmark case.
To run the Criteria benchmark:
python src/examples/criteria_evaluation_agent.py
The criteria benchmark cases are defined in data/criteria_benchmark.jsonl.
This example follows a different application of the library:
- define an agent implementing the BenchmarkAgent interface. In this application, each case already has the output we want to evaluate, so we override the
run_benchmark_casemethod to simply repackage eachBenchmarkCaseasBenchmarkCaseResponse. - create a set of quality benchmark cases, typically as a JSONL file such as data/criteria_benchmark.jsonl. In this application, each case's "extra" dictionary includes a "criteria" string.
- use a custom CriteriaBenchmarkRunner which overrides the
_execute_benchmark_casemethod, to run the benchmark using an evaluator that inherits fromCriteriaEvaluator
Metadata
Release files for flow-benchmark-tools 1.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| flow_benchmark_tools-1.5.0.tar.gz | 855.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| flow_benchmark_tools-1.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 882.8 kB
Release files / flow_benchmark_tools-1.5.0.tar.gz
| Download URL | flow_benchmark_tools-1.5.0.tar.gz |
|---|---|
| Size | 855.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
fd845e29e278a98e9099e7d3475ce14cf375a4aac7d7d6f39d7b65c454fdb6b4
|
|
BLAKE2b-256 checksum How to use checksums |
3b6ab73d1267120af320b8b10552fef68bc6e92bcb2a5c98d66c1cf9d525e799
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.1.1 CPython/3.11.8
|
Release files / flow_benchmark_tools-1.5.0-py3-none-any.whl
| Download URL | flow_benchmark_tools-1.5.0-py3-none-any.whl |
|---|---|
| Size | 26.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
710dff2646a17326a346f739e3c56e671dc2e2b15e90e13d0be7c91082d1a149
|
|
BLAKE2b-256 checksum How to use checksums |
bcbb09ae6ddf9abdacf7325f9d80b625d0f7493d27de2dad458c0608647399be
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.1.1 CPython/3.11.8
|