Skip to main content

harness-run

PyPI Python versions CI License

harness-run is a Python library for running coding agents locally and in remote sandbox environments through a common API.

It provides an abstraction over agent harnesses such as Claude Code and Codex, with support for multiple model providers. Your application can switch harnesses or models without being tightly coupled to one vendor's agent runtime.

The same Engine → Session → Run API works locally for development and remotely on GCP Agent Sandbox for production. With a warm sandbox pool, remote agents can start and complete small tasks in just a few seconds.

Install

pip install harness-run

To also run agents locally:

pip install "harness-run[local]"

Requirements

  • Python 3.12 or newer.
  • Credentials for the model the agent uses: an Anthropic, OpenAI or OpenRouter API key, or Vertex AI on the sandbox. See Getting started.
  • For remote runs, a GCP project prepared with harness-run-gcp-setup --project <your-project>. See GCP setup.

Example

Define a coding agent explicitly using the Codex harness:

from harness_run import AgentSpec

spec = AgentSpec(
    name="code-agent",
    harness="codex",
    model="gpt-5.6-luna",
    checkpoint=True,
)

Run it locally:

from harness_run import local

engine = local.deploy(spec)
session = engine.start_session()

result = await session.run(
    "Create a Python function fibonacci(n), add a few tests, and run them."
)

print(result.text)
print(result.cost_usd)

A session can continue across multiple turns:

result = await session.send(
    "Now make it iterative instead of recursive and run the tests again."
)

Run remotely

The same agent can run remotely on GCP Agent Sandbox. harness-run handles building and deploying the agent environment and running each turn inside an isolated sandbox. Credentials are passed to each run rather than baked into the deployed agent:

import os
from harness_run import sandbox

# Deployment is normally done by CI / ops.
sandbox.deploy(
    spec,
    project="my-project",
    location="us-central1",
    warm_pool=True,
)

# Application code looks up the deployed engine by name.
engine = sandbox.get_engine(
    "code-agent",
    project="my-project",
    location="us-central1",
)

session = engine.start_session()

result = await session.run(
    "Create a Python function fibonacci(n), add a few tests, and run them.",
    secrets={"OPENAI_API_KEY": os.environ["OPENAI_API_KEY"]},
)

session_id = session.session_id

With warm_pool=True, a ready sandbox is kept available for low-latency execution. Small turns typically start in about a second and can complete in a few seconds.

Another process can later re-attach to the session and continue the conversation:

engine = sandbox.get_engine(
    "code-agent",
    project="my-project",
    location="us-central1",
)

session = engine.get_session(session_id)

result = await session.send(
    "Add type hints and explain the time complexity.",
    secrets={"OPENAI_API_KEY": os.environ["OPENAI_API_KEY"]},
)

Want Claude Code instead? The application-level API stays the same. On GCP Agent Sandbox, Claude models are called through Vertex AI by default, so these runs need no API key at all:

spec = AgentSpec(
    name="code-agent",
    harness="claude-code",
    model="claude-sonnet-4-6",
    checkpoint=True,
)

The model can also be overridden per session or per turn, and an engine deployed with both harnesses baked in can pick either one per session, so a deployed engine does not need to represent a single fixed model configuration.

Customize the agent

Skills, MCP servers, a repository to work on and extra packages are all declared on the spec. Secrets are not: the spec names them, and each run supplies the values.

import os
from harness_run import AgentSpec, McpServer, RepoSource, SkillSource, SystemPrompt, local

spec = AgentSpec(
    name="repo-fixer",
    harness="codex",
    model="gpt-5.6-luna",
    system_prompt=SystemPrompt.inherit(append="Keep changes small and add tests."),
    # Skills from a git repository or a local directory, staged for the harness.
    skills=[SkillSource.git("https://github.com/zytedata/codex-skills")],
    # Cloned into the working directory before the agent runs; `auth` makes it push-ready.
    repos=[RepoSource.git("https://github.com/acme/some-repo", ref="main", auth="GH_TOKEN")],
    mcp_servers=[
        McpServer.github(),  # lets the agent open pull requests, using GH_TOKEN
        McpServer.remote(
            "search",
            "https://mcp.example.com",
            header_secrets={"Authorization": "SEARCH_AUTH"},
        ),
    ],
    packages=["httpx", "beautifulsoup4"],  # installed into the agent's environment
    checkpoint=True,
)

session = local.deploy(spec).start_session()

result = await session.run(
    "Fix the failing test in tests/test_parser.py and open a pull request.",
    secrets={
        "OPENAI_API_KEY": os.environ["OPENAI_API_KEY"],
        "GH_TOKEN": os.environ["GH_TOKEN"],
        "SEARCH_AUTH": "Bearer " + os.environ["SEARCH_API_KEY"],
    },
)

The same spec deploys unchanged to the sandbox. Every other field has a default; Configuration covers the rest, from environment variables and tool allow-lists to budgets and reasoning effort.

Features

Beyond the basic example, harness-run supports:

  • Multiple agent harnesses and models — Claude Code with Claude models, Codex with OpenAI models, and either harness with models served through OpenRouter.
  • Local and remote execution through the same API, with remote runs on GCP Agent Sandbox.
  • Low-latency remote runs using warm sandbox pools.
  • Multi-turn sessions with checkpoint/resume and durable remote history.
  • Per-session and per-turn configuration for models, prompts, budgets, tools, output schemas, and more.
  • Streaming and polling as alternatives to awaiting a completed run.
  • Structured output with Pydantic models or JSON Schema.
  • Repository and workspace setup, MCP servers, and agent skills.
  • Per-run secrets without baking credentials into deployed agents.
  • Steering, interruption, and shell access (exec()) while an agent is running.
  • Usage, cost, and resource reporting for production workloads.

Documentation

Start with Getting started.

For contributors, see TESTING.md. For implementation details and platform architecture, see DESIGN.md. Release notes and upgrade instructions are in CHANGELOG.md.

License and provenance

harness-run is developed by Zyte and released under the Apache-2.0 license.

Much of the code was written with AI coding agents, chiefly Claude Code, with the maintainers directing the work and reviewing the result.

Release files for harness-run 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for harness-run 0.4.0
File Size Uploaded
harness_run-0.4.0.tar.gz 457.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for harness-run 0.4.0
File Interpreter ABI Platform
harness_run-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 683.3 kB

Release files / harness_run-0.4.0.tar.gz

Download URL harness_run-0.4.0.tar.gz
Size 457.8 kB
Tags Source
SHA-256 checksum
How to use checksums
0f17125455d8691c8875ca34f0045efe775cc52390b86498619cd06117703bae
BLAKE2b-256 checksum
How to use checksums
c547e1247562529973ec74e54bdf6c29a52f445294c939d1e38441661486aea8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release files / harness_run-0.4.0-py3-none-any.whl

Download URL harness_run-0.4.0-py3-none-any.whl
Size 225.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
33d8dd035567a1d9ac3c7d27987942dc6a0f9dd2098128edf590675fdb53a195
BLAKE2b-256 checksum
How to use checksums
be5a4df2af47433d2f699ce853e0289988b752cfc4d42547906b0ce8d01886a1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page