Skip to main content

Tolokaforge

A benchmarking harness for evaluating tool-using LLM agents. Multi-turn agent/user loops, sandboxed execution, deterministic grading, and rich telemetry — across any provider via LiteLLM.

Highlights

  • Agent + User Loop – Multi-turn conversations where both agent and user models call tools.
  • Coding-harness mode – Run any of six vendor coding-agent CLIs (claude-code, codex, gemini-cli, kimi-code, opencode, grok-build) inside the trial container instead of the engine's own loop, so a task pack can measure a CLI's scaffolding, not only a bare model. See docs/CODING_HARNESSES.md.
  • Sandboxed Execution – Tool calls proxy into Dockerized services with no external network access.
  • MCP-Compatible Tooling – Tasks declare tools via Model Context Protocol or built-ins.
  • Deterministic Grading – JSONPath assertions, state hashes, transcript rules, optional LLM judges.
  • Rich Metrics – pass@k, cost/token estimates, latency percentiles, failure attribution.
  • Interactive Terminal – Optional Rich Live panel with trial list, live counters, budgets, and on-demand Docker container log inspection; panel-modal keyboard nav (Tab / j / k / l). See docs/CLI.md § Keyboard navigation.
  • Distributed Runner – SQLite for local runs, Postgres for multi-machine execution.
  • Bring-Your-Own Models – Any provider supported by LiteLLM (OpenAI, Anthropic, Google, Azure, Bedrock, Ollama, OpenRouter, and more).

Installation

pip install tolokaforge                # library only (headless / server)
pip install "tolokaforge[dx]"          # + terminal CLI (Rich panels, banners)
pip install "tolokaforge[browser]"     # + Playwright
pip install "tolokaforge[all]"         # everything

pip install tolokaforge transitively pulls the sibling tolokaforge-models wheel — that second wheel ships the model-data tables, the certificate registry, and the per-model policy subclasses on a release cadence independent of the engine. See ADR-0030.

The [dx] extras install the terminal front-end that owns the tolokaforge CLI — Rich panels, banners, and the Click command tree. Without them the library still imports (from tolokaforge.core.orchestrator import Orchestrator), and the tolokaforge console script prints an install hint pointing at pip install 'tolokaforge[dx]'. Front-end pluggability is recorded in ADR-0019.

Dev install:

curl -LsSf https://astral.sh/uv/install.sh | sh
uv sync
uv tool install --editable '.[dx]' --python 3.12   # exposes `tolokaforge` on PATH

The last line installs the tolokaforge command globally (into ~/.local/bin/). --editable keeps it pointing at your working tree so git pull updates it. All examples below assume tolokaforge is on PATH; if you skip the install step, prefix every command with uv run (e.g. uv run tolokaforge run …).

GitHub Codespaces

Open the repo in a Codespace (Code → Codespaces → Create codespace) for a ready-to-run environment. The container image ships a pinned uv, git-lfs, Node, the gh CLI, and Playwright's OS dependencies. On create and attach it runs uv sync, installs Playwright's Chromium, pulls Git LFS objects, and installs the pre-commit hooks. To run evaluations, supply a provider key as a Codespaces secret or in .envtolokaforge run reads it from the environment.

See Python Package Guide for all extras and programmatic API usage.

To run a single trial in-process from Python, tolokaforge.runner.run_trial(...) returns a typed TrialResult — see docs/API.md and the runnable examples/library/run_trial.py.

To drive a trial from any language over a pipe, tolokaforge run-trial runs one trial as a subprocess speaking a JSON-Lines wire — see docs/API.md and the runnable examples/run-trial/drive_run_trial.py.

Prebuilt images are published to Docker Hub as docker.io/tolokasoft1/tolokaforge-{runner,db-service,rag-service,mock-web} under a coordinated semver tag axis (:X.Y.Z, :X.Y, :latest, :X.Y.Z-rc.N), so a host with only Docker installed can docker pull them instead of building from a checkout — see the Standalone Runner Guide.

Quick Start

# 1. Configure provider keys
cp .env.example .env

# 2. Run one of the included examples
tolokaforge run --config examples/native/coding/run_configs/dev.yaml

# 3. Check results
tolokaforge status --run-dir results/coding_example
tolokaforge analyze --trajectory results/coding_example/trials/<task_id>/0/trajectory.yaml

Running tolokaforge with no subcommand drops into an interactive shell with tab-completion of every subcommand and flag — see docs/CLI.md § Interactive shell.

Under --display=rich (the default on a TTY), tolokaforge run opens a live panel with a trial list, focused-trial summary, budgets, and a components monitor. Tab cycles between panels (Trials / Engine Components / per-trial Infrastructure); j / k walk rows within the active panel; l on any focused row reveals its live Docker stdout / stderr. Full cheatsheet in docs/CLI.md § Keyboard navigation. All non-rich modes (--display=plain|log|none) work unchanged and keep the classic tolokaforge status / tolokaforge analyze / tolokaforge browse subcommands available.

That's it. Docker services for browser / mobile / RAG tasks start automatically via auto_start_services (default: true).

What is a run config?

A run config (e.g. examples/native/coding/run_configs/dev.yaml) is a single YAML file that fully specifies an evaluation. The harness reads it and runs the benchmark:

models:                       # which LLM(s) drive the agent + user simulator
  agent:    {provider: openrouter, name: anthropic/claude-sonnet-4-6, ...}
  user:     {provider: openrouter, name: anthropic/claude-sonnet-4-6, ...}

orchestrator:                 # how the run is executed
  workers: 4                  # parallel trials
  repeats: 1                  # trials per task
  max_turns: 20

evaluation:                   # what to evaluate
  tasks_glob: "**/task.yaml"  # which tasks to load (relative to task_packs root or repo)
  task_packs:                 # optional: directories that contain task.yaml files
    - "examples/native/coding/dataset"
  output_dir: "results/coding_example"

To write your own benchmark, copy a working example as a starting point:

cp examples/native/coding/run_configs/dev.yaml my_run.yaml
$EDITOR my_run.yaml         # change model, tasks_glob, output_dir
tolokaforge run --config my_run.yaml

Every example under examples/ ships a run_config.yaml next to its task data. There is no global "default" config — the run config and the tasks it points at always travel together.

For distributed execution and advanced workflows see the Runner Guide.

Runtime modes

Every trial runs inside a Docker stack. Two runtime backends decide how that stack is scoped:

  • shared (default): one stack materialises at run start and every trial hits it. Fast startup, lowest overhead.
  • per_trial: a fresh stack materialises for each trial via Testcontainers. Strict isolation — DB rows, filesystem edits, service state never leak across trials. Use it whenever a task mutates its environment.

Pick per-run via the CLI flag or the config key:

tolokaforge run --config my_run.yaml --runtime per_trial
# my_run.yaml
orchestrator:
  runtime: per_trial            # or "shared" (default)

A task can also declare its requirement via environment_manifest.isolation in task.yaml — the orchestrator refuses to run a per_trial task on the shared backend so cross-trial contamination is a startup-time error, not a silent grading bug.

Read docs/isolated_trials.md for a walkthrough (when to use each mode, how to opt in, cost tradeoffs), or docs/RUNTIME_BACKENDS.md for the full lifecycle deep-dive.

Multi-container tasks

Tasks aren't limited to the engine's built-in stack (runner + db-service ± RAG ± mock-web). A task can declare its own docker-compose.yaml — the task pack ships the compose file next to task.yaml, and the engine materialises that stack instead. Real databases, real APIs, custom services, whatever the compose file names.

# project.yaml
default_environment:
  stack:
    compose_file: "./environment.compose.yaml"
    runner_service: "runner"
  services:
    app-db:                     # per-service isolation: shared | reset | ephemeral
      isolation: "reset"
      reset: { seed: "postgres_baseline" }
  network_policy: "no_internet"

The multi_service example is the smallest working demonstration (runner + db-service + an nginx serving a product catalog); its two siblings scale up to a multi-endpoint join and a full PostgREST + postgres three-tier stack.

Start with the Multi-Container Guide — how the multi-container runtime works, how to run the six example packs, how to author your own stack, and how substrate grading works. For the case-matrix that decides which mode fits your task, see ADR-0018.

Coding-harness mode

A trial can hand its LLM turn loop over to a vendor coding-agent CLI installed inside the task container. Six harnesses ship in the tolokaforge_coding_harnesses registry:

agent_harness Vendor CLI
claude-code @anthropic-ai/claude-code
codex @openai/codex
gemini-cli @google/gemini-cli
kimi-code @moonshot-ai/kimi-code
opencode opencode-ai
grok-build x.ai/cli

The same run config that drives an engine-loop run drives a harness-mode run — one field, evaluation.harness_adapter.params.agent_harness, switches modes. Per-trial artifacts land in the same bundle shape either way, so downstream tooling is unchanged.

evaluation:
  harness_adapter:
    type: "terminal_bench"
    params:
      agent_harness: "claude-code"                        # switches modes
      agent_model:   "openrouter/anthropic/claude-sonnet-5"

Start with docs/CODING_HARNESSES.md for when to reach for each mode; the full recipe reference is docs/RUNNING_TERMINAL_BENCH.md.

Project Structure

tolokaforge/          # Installable Python package
├── cli/              # CLI commands (run, validate, status, analyze)
├── core/             # Orchestration, grading, metrics, queue
├── tools/            # Built-in + MCP tool system
└── env/              # Environment services (JSON DB, mock web, RAG)
tolokaforge_coding_harnesses/   # Coding-harness registry, installer, middleware proxy
tolokaforge_models/             # Model data + per-model policy subclasses
examples/             # Reference task layouts with runnable run_config.yaml
├── native/           # default `native` adapter
│   ├── browser_task/
│   ├── coding/
│   ├── multi_service/                # task-declared compose (nginx catalog)
│   ├── multi_service_advanced/       # multi-endpoint join
│   ├── multi_service_postgres/       # PostgREST + postgres three-tier
│   ├── multi_service_postgres_reset/ # project layer + per-trial postgres reset
│   ├── multi_service_slow_start/     # startup-order stress (slow-start postgres)
│   ├── example-microservices-pack/   # reference project for the Project schema
│   ├── native_shared_domain/
│   └── tool_use/
└── terminal_bench/   # `terminal_bench` adapter (Docker compose)

Documentation

Topic Link
Getting started docs/GETTING_STARTED.md
Architecture overview docs/ARCHITECTURE.md
Task authoring docs/TASKS.md
Grading system docs/GRADING.md
Tool reference docs/TOOLS.md
Browser/mobile tools docs/BROWSER_TOOLS.md
Runner & distributed execution docs/RUNNER.md
Standalone runner (embed / docker pull) docs/STANDALONE_RUNNER.md
Runtime backends (shared vs per-trial) docs/RUNTIME_BACKENDS.md
Projects — top-level abstraction docs/PROJECTS.md
Reset recipes (seed-backed per-trial reset) docs/RESET_RECIPES.md
Isolated trials guide docs/isolated_trials.md
Multi-container task guide docs/MULTI_CONTAINER_GUIDE.md
Coding-harness mode (vendor CLIs) docs/CODING_HARNESSES.md
Running terminal-bench (harness + engine loop) docs/RUNNING_TERMINAL_BENCH.md
Adapter architecture docs/ADAPTER_ARCHITECTURE.md
Analytics & failure attribution docs/ANALYTICS.md
Python package API docs/PYTHON_PACKAGE.md
Task packs docs/TASK_PACKS.md
Configuration reference docs/REFERENCE.md
Security model docs/SECURITY.md
Docker runtime docs/BENCHMARK_BACKEND_DESIGNS.md
Benchmark types docs/BENCHMARK_TYPES.md
Releasing (PyPI + Docker images) docs/RELEASING.md
Testing guide tests/README.md

Examples

Single-container tasks — one runner container plus the engine's built-in services (db-service, mock-web, RAG on demand):

Example Description
examples/native/coding/ Simplest native pattern — file-write grading
examples/native/tool_use/ Structured tool-call grading
examples/native/browser_task/ Browser tool against mock-web fixtures
examples/native/native_shared_domain/ _shared/domain.yaml + FastMCP pattern
examples/terminal_bench/ Docker-compose stacks with terminal_bench adapter

Multi-container tasks — the task ships its own compose file declaring additional services (see MULTI_CONTAINER_GUIDE.md):

Example Description
examples/native/multi_service/ Smallest multi-service task — runner + db-service + nginx product catalog
examples/native/multi_service_advanced/ Multi-endpoint aggregation across two task-specific HTTP APIs
examples/native/multi_service_postgres/ Realistic three-tier stack — PostgREST + postgres, no application code in the task pack
examples/native/multi_service_postgres_reset/ Project layer — per-trial postgres reset via a named sql_dump seed
examples/native/multi_service_slow_start/ Cross-service startup-order stress — slow-start postgres + depends_on: service_healthy
examples/native/example-microservices-pack/ Reference project for the full Project schema (see docs/PROJECTS.md)

Testing

make test              # all tests
make test-unit         # fast, isolated
make test-functional   # mocked externals

See tests/README.md for integration/E2E tests and contribution guidelines.

License

Apache-2.0 — see LICENSE.

Contributing

See CONTRIBUTING.md.

Citation

Use CITATION.cff or CITATION.bib when referencing Tolokaforge in papers or reports.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tolokaforge-0.21.0.tar.gz (5.0 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

tolokaforge-0.21.0-py3-none-any.whl (1.5 MB view details)

Uploaded Python 3

File details

Details for the file tolokaforge-0.21.0.tar.gz.

File metadata

  • Download URL: tolokaforge-0.21.0.tar.gz
  • Upload date:
  • Size: 5.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.6 {"installer":{"name":"uv","version":"0.12.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for tolokaforge-0.21.0.tar.gz
Algorithm Hash digest
SHA256 77544f5d0d3da2e895066594af2fb575cc17fe20fa8d38591a8788b186097c92
MD5 8187f434efb758cfdc8bf583489f6433
BLAKE2b-256 dbc4d08cac7ac5dfc66c726ca367eaf80a2074ef0c14c284cbda2b894519d1c0

See more details on using hashes here.

File details

Details for the file tolokaforge-0.21.0-py3-none-any.whl.

File metadata

  • Download URL: tolokaforge-0.21.0-py3-none-any.whl
  • Upload date:
  • Size: 1.5 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.6 {"installer":{"name":"uv","version":"0.12.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for tolokaforge-0.21.0-py3-none-any.whl
Algorithm Hash digest
SHA256 954e3a55b65d5beee452d30ab3b88dc20414c60b9e943e62d4fd9b082379a5c7
MD5 9ff35177ec5a16b69784962b762318ab
BLAKE2b-256 db8d0fea042994c0562e5e55bbad504b6dba861bc40a77f50f9df1a88e1a2771

See more details on using hashes here.

Release history Release notifications | RSS feed

0.22.5

2 files

0.22.4

2 files

0.22.3

2 files

0.22.2

2 files

0.22.1

2 files

0.22.0

2 files

0.21.4

2 files

0.21.3

2 files

0.21.2

2 files

0.21.1

2 files

This release

0.21.0 This release

2 files

0.20.0

2 files

0.19.1

2 files

0.19.0

2 files

0.18.1

2 files

0.18.0

2 files

0.16.1

2 files

0.15.0

2 files

0.14.2

2 files

0.14.1

2 files

0.14.0

2 files

0.13.1

2 files

0.13.0

2 files

0.12.0

2 files

0.11.2

2 files

0.11.1

2 files

0.11.0

2 files

0.10.0

2 files

0.9.3

2 files

0.9.2

2 files

0.9.1

2 files

0.9.0

2 files

0.8.4

2 files

0.8.3

2 files

0.8.2

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

0.4.1

2 files

0.4.0

2 files

0.3.1

2 files

0.3.0

2 files

0.2.11

2 files

0.2.10

2 files

0.2.9

2 files

0.2.8

2 files

0.2.7

2 files

0.2.6

2 files

0.2.5

2 files

0.2.4

2 files

0.2.3

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.0.2

1 file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page