Autobench
When an application changes, you need to know whether the new version is more correct, faster, cheaper, or more reliable. Teams usually answer that question with a custom benchmark script: a small program that runs representative inputs, checks the outputs, collects measurements, and prints a comparison. As the application grows, these scripts are repeatedly rewritten and their results become difficult to reproduce or compare.
We built Autobench to solve that problem. Define the cases, variants, task, scores, and reports once in YAML or Python. Autobench runs the experiment, collects the evidence, records exactly what happened, and lets you replay, report, compare, or export the result without rebuilding that infrastructure for every application.
It is a YAML-first Python framework for AI and non-AI systems:
- deterministic dataset x variant execution
- sync and async application tasks
- semantic observations, checks, measurements, artifacts, and ABP traces
- built-in and custom scoring, cost derivation, policies, and paired baselines
- native Pydantic AI, OpenAI, OpenAI Agents, and HTTPX instrumentation
- explicit and automatic prompt/tool/schema/agent asset lineage
- immutable YAML records, replay, Rich reports, comparisons, and exports
- optional immutable-record export to OTLP-compatible telemetry backends
Install
uv add autobench
For native SDK instrumentation:
uv add 'autobench[instrumentation]'
For outbound OTLP HTTP/protobuf export:
uv add 'autobench[otlp]'
First Run
autobench validate examples/minimal/autobench.yaml
autobench run examples/minimal/autobench.yaml --record /tmp/autobench-minimal
autobench replay /tmp/autobench-minimal
autobench report /tmp/autobench-minimal
A task is a normal sync or async callable:
from autobench import Case, RunContext
def run(ctx: RunContext, case: Case) -> Result:
mode = ctx.factor("mode")
with ctx.span("subject", kind="workflow") as span:
result = application(case.input, mode=mode)
span.set_output(result)
return result
The YAML spec owns reusable benchmark infrastructure: cases, variants, scoring, derivation, policies, instrumentation, and reports.
Examples
| Directory | Demonstrates |
|---|---|
examples/minimal |
Inline cases, variants, exact scoring, report and comparison |
examples/basic |
File dataset, spans, checks, artifacts and failure visibility |
examples/mid |
Semantic usage, pricing, cost, policies and distributions |
examples/advanced |
Repeated measurement and paired-baseline speedup |
examples/pydantic_ai |
Live layered instrumentation and automatic asset discovery |
examples/automatic_assets |
Offline Pydantic AI and custom SDK behavioral lineage |
examples/abp_* |
Manual, concurrent, streaming, Agents and replay protocol flows |
examples/otlp_export |
Offline immutable ABP record to OTLP mapping |
examples/codemode |
Migration of a real external benchmark runner |
Run the offline release matrix:
make examples
Documentation
Full documentation: vcoderun.github.io/autobench
- Installation
- First Benchmark
- Use Cases
- Architecture
- YAML Spec
- Python API
- Autobench Protocol
- OTLP Export
LLM-readable indexes are available at
llms.txt and
llms-full.txt.
Development
uv sync --extra dev --extra instrumentation --extra openai-agents --extra otlp
make prod
make pre-commit
The release gate enforces Python 3.11-3.14, strict lint and typing, strict documentation builds, offline examples, and 100% source line and branch coverage.
Release files for autobench 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| autobench-0.3.0.tar.gz | 699.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| autobench-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.0 MB
Release files / autobench-0.3.0.tar.gz
| Download URL | autobench-0.3.0.tar.gz |
|---|---|
| Size | 699.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
152dc5a997ef8878f4d3981136be75d9153317fb17c2b2ebd8feb89068f52299
|
|
BLAKE2b-256 checksum How to use checksums |
06dac1c7bbd3f159af7ccdb6801bd8a424ea4f88d595f56e9a5fef049c69f974
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.13
|
Release files / autobench-0.3.0-py3-none-any.whl
| Download URL | autobench-0.3.0-py3-none-any.whl |
|---|---|
| Size | 328.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
09c5804c0c450f6d9d9ada7db15abe1b7240bc3afbf222a9a55e09fa29df2e2c
|
|
BLAKE2b-256 checksum How to use checksums |
79869da3b06b53ef906fbca577dec5e662ca5e0032508716b2cf1c14d5b3cae3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.13
|