Skip to main content

Autobench

When an application changes, you need to know whether the new version is more correct, faster, cheaper, or more reliable. Teams usually answer that question with a custom benchmark script: a small program that runs representative inputs, checks the outputs, collects measurements, and prints a comparison. As the application grows, these scripts are repeatedly rewritten and their results become difficult to reproduce or compare.

We built Autobench to solve that problem. Define the cases, variants, task, scores, and reports once in YAML or Python. Autobench runs the experiment, collects the evidence, records exactly what happened, and lets you replay, report, compare, or export the result without rebuilding that infrastructure for every application.

It is a YAML-first Python framework for AI and non-AI systems:

  • deterministic dataset x variant execution
  • sync and async application tasks
  • semantic observations, checks, measurements, artifacts, and ABP traces
  • built-in and custom scoring, cost derivation, policies, and paired baselines
  • native Pydantic AI, OpenAI, OpenAI Agents, and HTTPX instrumentation
  • explicit and automatic prompt/tool/schema/agent asset lineage
  • immutable YAML records, replay, Rich reports, comparisons, and exports
  • optional immutable-record export to OTLP-compatible telemetry backends

Install

uv add autobench

For native SDK instrumentation:

uv add 'autobench[instrumentation]'

For outbound OTLP HTTP/protobuf export:

uv add 'autobench[otlp]'

First Run

autobench validate examples/minimal/autobench.yaml
autobench run examples/minimal/autobench.yaml --record /tmp/autobench-minimal
autobench replay /tmp/autobench-minimal
autobench report /tmp/autobench-minimal

A task is a normal sync or async callable:

from autobench import Case, RunContext


def run(ctx: RunContext, case: Case) -> Result:
    mode = ctx.factor("mode")
    with ctx.span("subject", kind="workflow") as span:
        result = application(case.input, mode=mode)
        span.set_output(result)
        return result

The YAML spec owns reusable benchmark infrastructure: cases, variants, scoring, derivation, policies, instrumentation, and reports.

Examples

Directory Demonstrates
examples/minimal Inline cases, variants, exact scoring, report and comparison
examples/basic File dataset, spans, checks, artifacts and failure visibility
examples/mid Semantic usage, pricing, cost, policies and distributions
examples/advanced Repeated measurement and paired-baseline speedup
examples/pydantic_ai Live layered instrumentation and automatic asset discovery
examples/automatic_assets Offline Pydantic AI and custom SDK behavioral lineage
examples/abp_* Manual, concurrent, streaming, Agents and replay protocol flows
examples/otlp_export Offline immutable ABP record to OTLP mapping
examples/codemode Migration of a real external benchmark runner

Run the offline release matrix:

make examples

Documentation

Full documentation: vcoderun.github.io/autobench

LLM-readable indexes are available at llms.txt and llms-full.txt.

Development

uv sync --extra dev --extra instrumentation --extra openai-agents --extra otlp
make prod
make pre-commit

The release gate enforces Python 3.11-3.14, strict lint and typing, strict documentation builds, offline examples, and 100% source line and branch coverage.

Release files for autobench 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for autobench 0.3.0
File Size Uploaded
autobench-0.3.0.tar.gz 699.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for autobench 0.3.0
File Interpreter ABI Platform
autobench-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.0 MB

Release files / autobench-0.3.0.tar.gz

Download URL autobench-0.3.0.tar.gz
Size 699.4 kB
Tags Source
SHA-256 checksum
How to use checksums
152dc5a997ef8878f4d3981136be75d9153317fb17c2b2ebd8feb89068f52299
BLAKE2b-256 checksum
How to use checksums
06dac1c7bbd3f159af7ccdb6801bd8a424ea4f88d595f56e9a5fef049c69f974
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.13

Release files / autobench-0.3.0-py3-none-any.whl

Download URL autobench-0.3.0-py3-none-any.whl
Size 328.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
09c5804c0c450f6d9d9ada7db15abe1b7240bc3afbf222a9a55e09fa29df2e2c
BLAKE2b-256 checksum
How to use checksums
79869da3b06b53ef906fbca577dec5e662ca5e0032508716b2cf1c14d5b3cae3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.13

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page