agentperf-local
Benchmark how fast your machine serves an AI agent.
An open-source tool from Artificial Analysis.
agentperf-local measures how fast your machine serves an AI agent. It replays
recorded agent conversations against an OpenAI-compatible model server and
reports throughput and latency. Each request carries the full conversation so
far, as a real agent's would. It measures speed, not answer quality.
It can benchmark a server you already run, or it can download a pinned model, start the server, and benchmark it for you.
Quick start
You need uv. It fetches Python 3.12 if you do
not have it. Docker and Rust are optional. macOS and Linux are supported;
Windows is not supported yet.
uv tool install agentperf-local
agentperf-local
To try it without installing, run uvx agentperf-local. pipx install agentperf-local also works.
To work from a source checkout instead:
git clone https://github.com/ArtificialAnalysis/aa-agentperf-local.git
cd aa-agentperf-local
uv sync
uv run agentperf-local
The examples below use uv run agentperf-local from a checkout. With an
installed tool, drop the uv run.
The TUI walks you through choosing a model, setting up, and running. Arrow keys
move, Enter continues, Escape goes back, ? opens help, and q quits. To
prefill settings with flags, use agentperf-local tui. See
TEXTUAL_TUI.md.
To run without the TUI, pick a catalog profile and a framework you have installed:
uv run agentperf-local managed-run \
--profile-id gemma4-12b-it-q4-0 \
--framework llama-cpp \
--output-dir results/gemma4-12b
managed-run downloads the pinned model into the Hugging Face cache, checks
every file's SHA-256, starts the server on localhost, runs the replay, and
stops the server. MODELS.md lists all 29 profiles.
The default run
The default replay is agentperf-default-v1: eight recorded agent tasks with
168 model turns. It uses the exact output policy, which makes every turn
generate its recorded number of tokens. This is the comparable run.
--replay aa-mini-v1 selects a six-turn synthetic replay. It is a quick
install check. Its results are not comparable.
Context requirement
A full run needs a 65,536-token context at batch size 1. Every catalog profile launches at that context. An attached server that serves more, such as 131,072 tokens, also counts as full.
The default replay's largest turn needs about 58,000 tokens, so it never runs
below 65,536. A smaller context only serves a replay whose floor allows it,
such as aa-mini-v1: pass --context-tokens 32768 to managed-run, or choose
a smaller context in the TUI. That run is marked reduced: true and is not
comparable with full-context results. An attached server's context is read
from the server at the start of the run.
Attached servers
Point run at a server you already run:
uv run agentperf-local run \
--base-url http://127.0.0.1:8080/v1 \
--model served-model \
--output-dir results/my-server
Use your server's base URL and the model name it reports:
| Server | Typical base URL |
|---|---|
| llama.cpp | http://127.0.0.1:8080/v1 |
| LM Studio | http://127.0.0.1:1234/v1 |
| vLLM | http://127.0.0.1:8000/v1 |
| SGLang | http://127.0.0.1:30000/v1 |
The exact policy sends ignore_eos, which is not part of the OpenAI API.
Before the replay, run checks that the server honours it. If it does not,
run stops and names the flag to change.
Ollama cannot honour ignore_eos. The tool detects Ollama before the run
and warns you. The run then uses the recorded policy, which lets the model
stop on its own and reports end-to-end latency as a normalized estimate. That
result is not directly comparable with exact runs.
To send an API key, put it in an environment variable and pass the variable's
name with --api-key-env. No flag takes a literal key.
Real tool calling
By default the replay skips tool time between turns. run --tool-mode live
runs the recorded shell commands in Docker containers, so tool time is real.
This mode is opt-in and needs Docker and prebuilt images. The build scripts
are in the source checkout:
scripts/install-swebench-validation.sh # arm64 hosts only, once
scripts/build-default-containers.sh
uv run agentperf-local run \
--base-url http://127.0.0.1:8080/v1 \
--model served-model \
--output-dir results/live-tools \
--tool-mode live
Security caveats:
- The recorded commands are untrusted input. Docker reduces the risk but is not a security boundary.
- Access to the Docker daemon is equivalent to root on most hosts.
- Containers have no network by default, but a replay manifest can ask for one.
Pass
--live-network noneto force no network. - The task workspace is mounted read-write. It lives under the output directory
unless you pass
--live-workspace-root. - The scripts pull or build third-party images and clone the SWE-bench harness from GitHub.
Live-tool runs cannot be submitted.
Rust client (experimental)
Python is the default client. An optional Rust client is available for high-concurrency benchmarking. It needs a Rust toolchain:
uv sync --extra rust
uv run agentperf-local tui --client rust
With an installed tool, use uv tool install 'agentperf-local[rust]'
instead. A plain uv sync removes the extension again. Both clients produce the same
metrics.
Submitting results
Submitting is optional and nothing is uploaded unless you ask.
prepare-submission builds a bundle, submit sends it, and
submission-status reads it back. See SUBMITTING.md for
what is sent and what stays private.
Commands
| Command | Purpose |
|---|---|
tui |
Open the guided full-screen app. |
run |
Replay a workload against a server you run. |
managed-run |
Download a catalog model, serve it, and benchmark it. |
deployment-options |
Show which frameworks can serve one catalog profile on this machine (default gemma4-12b-it-q4-0, or --profile-id). |
doctor |
Show local hardware facts without identifiers. |
convert |
Convert an agent recording into a replay manifest. |
prepare-submission |
Build a submission bundle without uploading it. |
submit |
Send a prepared bundle to Artificial Analysis. |
submission-status |
Read a submission's status. |
Run uv run agentperf-local <command> --help for every option.
python -m agentperf_local is the same program.
Results
A run writes a new directory:
results/my-server/
├── turns.jsonl
├── tasks.json
├── tools.json
├── failures.json
├── summary.json
└── measurement.json
summary.json holds the headline numbers: output tokens per second and median
and p95 first-token and turn times. It is written last, so a crashed run has
no summary. A managed run also writes the server log and deployment record.
FORMATS.md describes every file.
The results are private. They can contain paths, model labels, and errors. Starting a run sends the recorded prompts to the model server, so a remote URL sends them off your machine.
Development
make lint # ruff check, ruff format --check, ty check
make test # pytest
make test-rust # Rust core tests
CI runs these checks and the Python and Rust client equivalence tests. See AGENTS.md for the full list and the code conventions.
Documentation
- Models: the catalog and per-device caveats.
- Recipes: measured evidence for the device-specific profiles.
- Textual TUI: options, keys, and screens.
- Architecture: measurement rules, evidence boundary, and package layout.
- Data formats: inputs, outputs, and schemas.
- Submitting: what is uploaded and how it is used.
agentperf-local is built and maintained by Artificial Analysis. Code is licensed under Apache-2.0. The Artificial Analysis name and logo are not covered by the code licence.
Metadata
Release files for agentperf-local 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agentperf_local-0.1.0.tar.gz | 2.9 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agentperf_local-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 5.6 MB
Release files / agentperf_local-0.1.0.tar.gz
| Download URL | agentperf_local-0.1.0.tar.gz |
|---|---|
| Size | 2.9 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
988a8c9c8f14a877eab84af81a151ae7b6f32e91ead893983c2bd8e5744f192b
|
|
BLAKE2b-256 checksum How to use checksums |
e08fcbfcf86e1ffcc88d9b6eb19e91bbc2d1da077b9d4731b4989dfc44cf3ebb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency logRelease files / agentperf_local-0.1.0-py3-none-any.whl
| Download URL | agentperf_local-0.1.0-py3-none-any.whl |
|---|---|
| Size | 2.7 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
66e2b225fa32e28d1268c8092e17a75152077303023144e8cfb74b6f6b3da16f
|
|
BLAKE2b-256 checksum How to use checksums |
bd68e57e4f2203ea4acf5e17a2b593335ffdbd03fce6a09bbbb6ce13ebe1182a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency log