Skip to main content

Artificial Analysis

agentperf-local

Benchmark how fast your machine serves an AI agent.
An open-source tool from Artificial Analysis.

License: Apache-2.0 Python 3.12+

agentperf-local measures how fast your machine serves an AI agent. It replays recorded agent conversations against an OpenAI-compatible model server and reports throughput and latency. Each request carries the full conversation so far, as a real agent's would. It measures speed, not answer quality.

It can benchmark a server you already run, or it can download a pinned model, start the server, and benchmark it for you.

Quick start

You need uv. It fetches Python 3.12 if you do not have it. Docker and Rust are optional. macOS, Linux, and Windows are supported. On Windows, managed runs need an NVIDIA GPU and llama.cpp; vLLM and SGLang run only on Linux.

uv tool install agentperf-local
agentperf-local

To try it without installing, run uvx agentperf-local. pipx install agentperf-local also works.

To work from a source checkout instead:

git clone https://github.com/ArtificialAnalysis/aa-agentperf-local.git
cd aa-agentperf-local
uv sync
uv run agentperf-local

The examples below use uv run agentperf-local from a checkout. With an installed tool, drop the uv run.

The TUI walks you through choosing a model, setting up, and running. Arrow keys move, Enter continues, Escape goes back, ? opens help, and q quits. To prefill settings with flags, use agentperf-local tui. See TEXTUAL_TUI.md.

To run without the TUI, pick a catalog profile and a framework you have installed:

uv run agentperf-local managed-run \
  --profile-id gemma4-12b-it-q4-0 \
  --framework llama-cpp \
  --output-dir results/gemma4-12b

managed-run downloads the pinned model into the Hugging Face cache, checks every file's SHA-256, starts the server on localhost, runs the replay, and stops the server. Every recipe it can run is a YAML file in recipes/, by model and then hardware.

The default run

The default replay is agentperf-default-v1: eight recorded agent tasks with 168 model turns. It uses the exact output policy, which makes every turn generate its recorded number of tokens. This is the comparable run.

--replay aa-mini-v1 selects a six-turn synthetic replay. It is a quick install check. Its results are not comparable.

Context requirement

A full run needs a 65,536-token context at batch size 1. Every catalog profile launches at that context. An attached server that serves more, such as 131,072 tokens, also counts as full.

The default replay's largest turn needs about 58,000 tokens, so it never runs below 65,536. A smaller context only serves a replay whose floor allows it, such as aa-mini-v1: pass --context-tokens 32768 to managed-run, or choose a smaller context in the TUI. That run is marked reduced: true and is not comparable with full-context results. An attached server's context is read from the server at the start of the run.

Attached servers

Point run at a server you already run:

uv run agentperf-local run \
  --base-url http://127.0.0.1:8080/v1 \
  --model served-model \
  --output-dir results/my-server

Use your server's base URL and the model name it reports:

Server Typical base URL
llama.cpp http://127.0.0.1:8080/v1
LM Studio http://127.0.0.1:1234/v1
vLLM http://127.0.0.1:8000/v1
SGLang http://127.0.0.1:30000/v1

The exact policy sends ignore_eos, which is not part of the OpenAI API. Before the replay, run checks that the server honours it. If it does not, run stops and names the flag to change.

Ollama cannot honour ignore_eos. The tool detects Ollama before the run and warns you. The run then uses the recorded policy, which lets the model stop on its own and reports end-to-end latency as a normalized estimate. That result is not directly comparable with exact runs.

To send an API key, put it in an environment variable and pass the variable's name with --api-key-env. No flag takes a literal key.

Real tool calling

By default the replay skips tool time between turns. run --tool-mode live runs the recorded shell commands in Docker containers, so tool time is real. This mode is opt-in and needs Docker and prebuilt images. The build scripts are in the source checkout. They are bash scripts, so on Windows run them from Git Bash or WSL:

scripts/install-swebench-validation.sh   # arm64 hosts only, once
scripts/build-default-containers.sh

uv run agentperf-local run \
  --base-url http://127.0.0.1:8080/v1 \
  --model served-model \
  --output-dir results/live-tools \
  --tool-mode live

Security caveats:

  • The recorded commands are untrusted input. Docker reduces the risk but is not a security boundary.
  • Access to the Docker daemon is equivalent to root on most hosts.
  • Containers have no network by default, but a replay manifest can ask for one. Pass --live-network none to force no network.
  • The task workspace is mounted read-write. It lives under the output directory unless you pass --live-workspace-root.
  • The scripts pull or build third-party images and clone the SWE-bench harness from GitHub.

Live-tool runs cannot be submitted.

Rust client (experimental)

Python is the default client. An optional Rust client is available for high-concurrency benchmarking. It needs a Rust toolchain:

uv sync --extra rust
uv run agentperf-local tui --client rust

With an installed tool, use uv tool install 'agentperf-local[rust]' instead. A plain uv sync removes the extension again. Both clients produce the same metrics.

Submitting results

Submitting is optional and nothing is uploaded unless you ask. prepare-submission builds a bundle, submit sends it, and submission-status reads it back. See SUBMITTING.md for what is sent and what stays private.

Commands

Command Purpose
tui Open the guided full-screen app.
run Replay a workload against a server you run.
managed-run Download a catalog model, serve it, and benchmark it.
deployment-options Show which frameworks can serve one catalog profile on this machine (default gemma4-12b-it-q4-0, or --profile-id).
doctor Show local hardware facts without identifiers.
convert Convert an agent recording into a replay manifest.
prepare-submission Build a submission bundle without uploading it.
submit Send a prepared bundle to Artificial Analysis.
submission-status Read a submission's status.

Run uv run agentperf-local <command> --help for every option. python -m agentperf_local is the same program.

Results

A run writes a new directory:

results/my-server/
├── turns.jsonl
├── tasks.json
├── tools.json
├── failures.json
├── summary.json
└── measurement.json

summary.json holds the headline numbers: output tokens per second and median and p95 first-token and turn times. It is written last, so a crashed run has no summary. A managed run also writes the server log and deployment record. FORMATS.md describes every file.

The results are private. They can contain paths, model labels, and errors. Starting a run sends the recorded prompts to the model server, so a remote URL sends them off your machine.

Development

make lint        # ruff check, ruff format --check, ty check
make test        # pytest
make test-rust   # Rust core tests

CI runs these checks and the Python and Rust client equivalence tests. See AGENTS.md for the full list and the code conventions.

Documentation

  • Recipes: every managed-run recipe, and how to add one.
  • Textual TUI: options, keys, and screens.
  • Architecture: measurement rules, evidence boundary, and package layout.
  • Data formats: inputs, outputs, and schemas.
  • Submitting: what is uploaded and how it is used.

agentperf-local is built and maintained by Artificial Analysis. Code is licensed under Apache-2.0. The Artificial Analysis name and logo are not covered by the code licence.

Metadata

Release files for agentperf-local 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agentperf-local 0.2.0
File Size Uploaded
agentperf_local-0.2.0.tar.gz 2.8 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for agentperf-local 0.2.0
File Interpreter ABI Platform
agentperf_local-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 5.4 MB

Release files / agentperf_local-0.2.0.tar.gz

Download URL agentperf_local-0.2.0.tar.gz
Size 2.8 MB
Tags Source
SHA-256 checksum
How to use checksums
36f9f6cf6784a06f87e6123afb2b2a0b3e7629141f7cf264ebad0f34caa311dc
BLAKE2b-256 checksum
How to use checksums
a88f099d643a2d857dc3fc189554fd28fd6326dd51648d26d9413972d064d370
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / agentperf_local-0.2.0-py3-none-any.whl

Download URL agentperf_local-0.2.0-py3-none-any.whl
Size 2.6 MB
Tags Python 3
SHA-256 checksum
How to use checksums
95625ded4368cdcbb720d954f6d02705ef3b087651366d9d3c3578716b45765f
BLAKE2b-256 checksum
How to use checksums
d8a3c9658133213c1fc01328d4b3b30a39517eee5696a45447e60825db02e35f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page