Skip to main content

AssistantEval

A verifiable benchmark for personal AI assistants. Results, methodology and every published run: assistanteval.com.

Each task puts an assistant in a small synthetic world: a calendar, a mailbox, a marketplace and so on, served to it over REST or MCP. A simulated user talks to it over several turns, gives private facts only when asked, and approves only what its rules allow. Every service records what the assistant did, and the grader scores the run from that record alone.

The assistant can be any model in a minimal tool loop, an open agent harness you adapt, or your own product. The models that play the user's reader and writer and that judge the run are yours to choose too: any OpenAI-compatible API, hosted or local.

How a run works

  1. Task (tasks/<id>.json): the world's fixtures, the opening message, what only the user knows, when the user approves an action, and the grading rules. See tasks/README.md.
  2. Runner (assistanteval/run_task.py): starts the task's services on a local server with fresh per-run keys, resets the assistant, connects the services, and runs the conversation with the simulated user (assistanteval/simulated_user.py).
  3. Run record (runs/<run_id>/): the conversation, every service request with its input and output, the start and end state, the simulated user's decisions, and which models played each role. See docs/run-record.md.
  4. Grader (assistanteval/grader/): validity checks, then rules answered by code or by a judge ladder (a saved human answer, the reader, the judge, then a human annotation queue), then a pass-rate report. See docs/grader/README.md.

Run it

uv sync
uv run pytest -q          # includes an offline end-to-end run: no API key needed

# Set the judges and the simulated user's models: see "Two setups" below.
# The assistant: here a model in the built-in tool loop
export ASSISTANTEVAL_BASELINE_MODEL=<model> ASSISTANTEVAL_BASELINE_BASE_URL=<url> ASSISTANTEVAL_BASELINE_API_KEY=<key>

uv run assistanteval run tasks/kingsway-reschedule-crew.json --assistant model_baseline
uv run assistanteval grade runs/<run_id>
uv run assistanteval report runs/ --out grades/

To run tasks in Harbor against an agent harness such as Hermes Agent or OpenCode, see docs/harbor.md.

--assistant also takes module:factory, so an adapter for your own harness or product can live in its own package. See docs/developing-episodes.md for the assistant contract, the services and how to add a task or a service.

Model roles

Role Does Required
JUDGE Answers what the reader is unsure about, and checks proposals against the user's approval rules For a live run or live grading
READER Reads each assistant turn for the simulated user, and answers the grader's questions first No; without it the judge reads too
WRITER Rewords the simulated user's fixed replies in the persona's voice No; without it the fixed replies are sent
BASELINE The model under test in the built-in tool loop For --assistant model_baseline

Each role reads ASSISTANTEVAL_<ROLE>_MODEL, _BASE_URL, _API_KEY, and optionally _API (openai, or jev for TypeSafe's Jev as the reader) and _EXTRA_BODY. Records keep each role's model and base URL, never its key. Scores from different judge models are not directly comparable: report the judges with the scores.

Two setups

Official: the judges of the published AssistantEval results. Use it to compare with them.

export ASSISTANTEVAL_READER_API=jev ASSISTANTEVAL_READER_MODEL=jev-1.13.0 ASSISTANTEVAL_READER_API_KEY=<TypeSafe key>
export ASSISTANTEVAL_JUDGE_MODEL=gpt-5.6-sol ASSISTANTEVAL_JUDGE_API_KEY=<OpenAI key>
export ASSISTANTEVAL_WRITER_MODEL=z-ai/glm-5.3-flash ASSISTANTEVAL_WRITER_BASE_URL=https://openrouter.ai/api/v1 \
       ASSISTANTEVAL_WRITER_API_KEY=<OpenRouter key> ASSISTANTEVAL_WRITER_EXTRA_BODY='{"reasoning": {"effort": "low"}}'

OpenRouter: one key, for development. The same judge and writer models; the reader is TypeSafe's Jev Router, which picks its model for each request rather than pinning jev-1.13.0. Its readings can differ from jev-1.13.0's, so compare with published results only under the official setup.

OR=https://openrouter.ai/api/v1 KEY=<OpenRouter key>
export ASSISTANTEVAL_READER_MODEL=typesafe/jev-router ASSISTANTEVAL_READER_BASE_URL=$OR ASSISTANTEVAL_READER_API_KEY=$KEY
export ASSISTANTEVAL_JUDGE_MODEL=openai/gpt-5.6-sol ASSISTANTEVAL_JUDGE_BASE_URL=$OR ASSISTANTEVAL_JUDGE_API_KEY=$KEY
export ASSISTANTEVAL_WRITER_MODEL=z-ai/glm-5.3-flash ASSISTANTEVAL_WRITER_BASE_URL=$OR ASSISTANTEVAL_WRITER_API_KEY=$KEY \
       ASSISTANTEVAL_WRITER_EXTRA_BODY='{"reasoning": {"effort": "low"}}'

To choose other judges, see docs/grader/README.md.

Repository layout

Path What it is
tasks/ The task catalog and its fixtures.
assistanteval/run_task.py, contracts.py, run_record.py, task.py, dates.py The runner, the assistant contract, the run record, the task model, and the renderer of a task's dates at a run's start.
assistanteval/endpoints.py The model roles.
assistanteval/simulated_user.py The simulated user.
assistanteval/grader/ Grading, annotation and the pass-rate report.
assistanteval/services/, world.py, connectors.py The synthetic services, the engine that serves them over REST and MCP, and the local server the runner opens them on.
assistanteval/assistants/ Built-in assistants: the model baseline and a scripted reference assistant.
assistanteval/harbor/, remote.py, assistants/acp.py Running tasks in Harbor against any ACP harness. See docs/harbor.md.
docs/frameworks.md Which open eval and RL frameworks to integrate with, and why.

Safety rails

  • Assistants only ever reach synthetic services, with per-run keys that are revoked, and proven revoked, after every run. In Harbor, where a task's MCP servers carry no key, the trial's private network isolates the run instead.
  • No service delivers mail; the email service also sends only to its fixture's allowed recipients.

License

The code is licensed under the Apache License 2.0. The tasks and their fixtures in tasks/ are licensed under CC BY 4.0: you may share and adapt them, including commercially, with attribution to AssistantEval by ORO AI. See NOTICE.

Metadata

Release files for assistanteval 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for assistanteval 0.1.0
File Size Uploaded
assistanteval-0.1.0.tar.gz 516.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for assistanteval 0.1.0
File Interpreter ABI Platform
assistanteval-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 640.3 kB

Release files / assistanteval-0.1.0.tar.gz

Download URL assistanteval-0.1.0.tar.gz
Size 516.1 kB
Tags Source
SHA-256 checksum
How to use checksums
e20d8b6e5a99c128fe665658205864f8e84d2949a92bb824a110a27219490486
BLAKE2b-256 checksum
How to use checksums
32633b1f98924ef30735cb3247dcc2a690787d45f803cdaaec219b2af30f5df8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release files / assistanteval-0.1.0-py3-none-any.whl

Download URL assistanteval-0.1.0-py3-none-any.whl
Size 124.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6ab0872f70a29e24fcaa1ad98651562137fe621bcfaf181ef20d14f4665d4eb5
BLAKE2b-256 checksum
How to use checksums
45c5e32a5c96fcacc06e1bd0ba0487de9a37889b372f1a7a5e30d6aa7d8febc6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page