Skip to main content

Attune Harness

Run agent tasks, check the results independently, and keep the evidence.

Attune Harness checks work from AI agents, commands, or your own code against criteria you supply. Each run returns a receipt with the output, check result, and any errors.

In a JSONL exporter experiment, checks through the real command line caught three defects that serializer-only tests missed. Read the experiment and its limits. For a small working example, see how checks and receipts work.

1.2.0 · A refused Claude turn can be retried. The v1 compatibility contract is in effect; experimental features and platform limits are listed below. Release notes. Qualification status.

Install Attune Harness · User guide

Plan, build, review, fix, and test

attune-harness --help

Getting started
  init         Write a starter participant registry

Task execution
  plan         Define intent and accept its scope
  build        Execute accepted tasks and protected checks
  review       Assess a document against project evidence
  fix          Repair scoped files and check the result
  test         Test a captured change and retain the evidence

Task controls
  status       Inspect a saved task
  resume       Continue a saved task

AI tools and integration
  --help-all   Browse operational tools and compatibility commands

Every verb follows one pattern. You accept a scope before anything runs, the host applies the effects, and checks fixed in advance decide the outcome. fix is the clearest example: it freezes an acceptance probe before the worker starts, applies the worker's proposed replacement itself, and keeps the failed-before and passed-after probe evidence. A saved task can be inspected with status and continued with resume; completed operations replay from saved evidence instead of running again.

Reaching outside the process takes explicit permission. External command participants need --allow-external, and native model participants also need --allow-native. Approving a plan does not authorize paid calls.

Many of these commands are there for the agent and its integrations to call. Learning their syntax is not the price of entry, and attune-harness COMMAND --help covers direct use. Full usage, exit codes and recovery controls are in the CLI guide.

Use Harness with Codex, Claude, or other models

This repository includes an Attune Harness skill for the existing plan, build, review, fix, test, status and resume workflows. Invoke $attune-harness and describe the task; the agent prepares the CLI inputs and reports the saved evidence. It preserves the workflow's scope and execution permissions. The older /attune command belongs to Attune AI.

For an Attune Harness entry in Codex Plugins, use the plugin packaging and installation guide. You can also use standalone skill discovery in this checkout or another project. Installing the Python package alone installs neither integration.

In Claude Code, add this repository as a plugin marketplace, then install the plugin. It carries the Harness skill plus cross-review and smart-test, and it calls the attune-harness command installed above:

/plugin marketplace add Smart-AI-Memory/attune-harness
/plugin install attune-harness@attune-harness

Claude Code does not read the .agents/skills/ folder that Codex discovers, so this plugin is how the skill reaches it. Installing the plugin authorizes no paid calls.

Installation

pipx install 'attune-harness[all]==1.2.0'

or uv tool install 'attune-harness[all]==1.2.0', or pip install 'attune-harness[all]==1.2.0' into an environment of its own. This is the recommended install: everything the review, test, MCP and acceptance journeys need, plus Redis and Voyage retrieval. Python 3.10 or later. The worked example needs no API key or attune-ai installation; Voyage retrieval needs a Voyage API key and makes paid calls. Harness and attune-ai cannot share one environment: they pin different lines of the MCP SDK, and installing Harness over attune-ai replaces attune-ai's; an isolated install avoids that, and mcp-serve says so if it finds the two side by side.

What the install carries

You want Install
Recommended: the base package plus the Redis reader and Voyage retrieval. The experimental extra below stays explicit pip install 'attune-harness[all]==1.2.0'
The contracts and CLI; the evidence-review, test, acceptance and MCP journeys: forms (attune-forms 0.17.0), document claim verification (attune-verify 0.6.0), local Markdown retrieval with source hashes (attune-rag 1.2.0), MCP stdio serving (mcp 2.2.0) and token counting (tiktoken 0.12.0). No model calls pip install 'attune-harness==1.2.0'
Read the Redis memory a hydration keeps warm: the recall digest, related nodes, one record, full-text search over the index (redis 5.3.1); read-only, the text body of a file, lesson or rule pointer is never served; needs a reachable Redis Stack with the hydration's index and function library. Also the Redis backend for working memory (memory scratch), shared across processes and machines; the file backend is in the base pip install 'attune-harness[redis]==1.2.0'
Repository-first retrieval on Voyage embeddings. Needs a Voyage API key, makes paid calls pip install 'attune-harness[voyage]==1.2.0'
Experimental: memory proposals from a Claude model over a pinned, data-only Anthropic API transport (anthropic 1.6.0, httpx2 2.13.0). POSIX only, needs ANTHROPIC_API_KEY, makes paid calls pip install 'attune-harness[memory-native]==1.2.0'

Before 0.4.0 the base had no dependencies and verify, rag, review, mcp and tokens were extras; they were empty from 0.4.0 and are gone since 0.6.0, so an install that still names one gets pip's warning that the extra does not exist and the base install; drop the bracket. Every dependency is pinned exactly and loaded on first use, so a wheel installed without its dependencies still returns an actionable unavailable report for each missing piece instead of a traceback. Keep the quotes around an extra: zsh and bash treat square brackets as glob characters.

Your first five minutes

From a Git project with an uncommitted change and a virtual environment that has pytest, after the install above:

cd your-project
attune-harness init
attune-harness test --project . --scope src/example.py --interpreter .venv/bin/python --task-dir ../test-task --format markdown
attune-harness review --goal "Check README.md against project evidence" --intake-only

init writes a participants.json with two offline demo participants. test previews exactly which tests it will run for your change and saves that plan; accept it with the checkpoint it prints to run them. review shows the intake form for an evidence review: the document, its evidence and who assesses it. Nothing here calls a model. CI runs these commands, as written, on every platform job.

What is qualified and what is not

I would rather you find the limits here than in your own checkout. Green software tests and model quality are different claims, and this project keeps them apart.

Area Qualified Not qualified
Platforms CI builds and installs the wheel on macOS, Ubuntu and Windows with Python 3.10 and 3.12, and exercises timeouts, cancellation, bounded output, crash-released locks and recovery (guide) Other Python versions are outside the matrix. On Windows, a process that holds a run's record.json open for more than about two seconds still fails that run closed
Models CI calls no model provider. Native Claude and Codex adapters have recorded comparisons Native planning and building are experimental. In the September 18, 2026 comparison the original reply contract accepted 1 of 24 replies; after the contract was corrected it accepted 12 of 12. Two repetitions per role do not establish a reliability rate
fix and test Local POSIX Git checkouts, regular files, default pytest discovery File creation, deletion and renames, linked worktrees, custom pytest collectors, committed revision ranges
fix on Windows Nothing yet. New in 0.2.0 and experimental: fix runs on a fixed local NTFS volume instead of refusing, and its native tests pass in CI on windows-2022 and windows-2025 (design note) Everything beyond those tests: deletion and renames, files with their own ACL or nonstandard attributes, files over 64 KiB, crash recovery, concurrent writers, power-loss durability, and any run against a real project. test on Windows is unchanged and unqualified
Isolation Commands and probes run as supervised processes with deadlines and bounded output This is not a security sandbox. Use a dedicated checkout and commands you trust
Receipts Receipts retain the task, output and check evidence locally They are local values, not signed attestations. Constructing a Receipt directly certifies nothing
Plan acceptance Core imports, help and the library run standalone. plan --accept runs from the base install with no Attune AI; CI exercises its gate with Attune AI blocked Acceptance through a live MCP host; CI submits the console approval
Protocols MCP (2025-11-25 and 2026-07-28 profiles) and A2A 1.0 have local independent-client receipts Remote authentication and automatic host installation
Signed Python plugins (1.0.0rc1) The declared cooperating-plugin profile checks signatures and revocation, effective grants, import closure, bounded subprocesses and host-owned journals. The installed-wheel tests passed on macOS, Ubuntu and Windows with Python 3.10 and 3.12 (qualification). This is not an OS sandbox for hostile same-user code. An explicitly accepted artifact receives only its accepted effective grants; paid dispatch needs separate authority. Installation does not authorize execution. Arbitrary plugins, remote authentication and automatic installation are not qualified
Voyage plugin path (1.0.0rc1) One fixed public-corpus query completed eight live stages total across two paths (four matching request/result pairs). The direct SDK and signed-plugin paths produced identical normalized requests/results and five ranked sources. Offline replay of recorded results and a synthetic interrupted stage passed in the six installed-wheel jobs (recorded fixture). This does not establish general ranking or model quality, capture raw HTTP responses, or imply a live provider call on every platform. A dispatched stage with unknown effects is not retried automatically
Memory memory recall, resolve and refresh read explicit raw, personal and curated roots with Harness's own reader, nothing from attune-ai on the path; the raw tier needs only the standard library and the document tiers attune-rag from the base install; CI reads a raw root from the installed wheel on every POSIX platform job and from the --no-deps release gate, and the document tiers where attune-rag is installed; the Windows jobs record the reader's refusal. The legacy formats are frozen as read; native is the only reader. With the redis extra, memory redis reads a hydrated Redis Stack keyspace: digest, related, node and search, as evidence packets; a pointer's text body is never served, a curated node's own record is. memory serve checks active curated membership before printing the digest, includes source and untrusted-evidence framing, and exits 0 whether or not Redis answers. serve --for PROMPT searches prompt terms and suppresses file pointers with a wrong verdict or without an authorized verdict scope; a UserPromptSubmit bridge is documented in the CLI guide. memory scratch is working memory, bounded JSON under short keys with an optional time to live: a file store in the base that declares no sharing, or the Redis store with the extra; the backend is chosen at startup and an unreachable Redis is never replaced by the file store. CI exercises the extra absent and the extra present with no server on all three platforms; the path against a hydrated server passed once on a maintainer's machine on 2026-09-22 and runs only where ATTUNE_TEST_REDIS_URL is set The native reader is POSIX-only: on Windows it reports the refusal instead of reading, until the Phase 4 decision. The immutable native format fixture pins the compatibility contract. The memory-native extra ships in the wheel but is experimental and not activated for live memories. The native transport is POSIX only, accepts two exact model IDs, refuses any other SDK version, and is never exercised in CI
Roadmap ship and reflect are planned routes and do not exist yet

Harness and attune-ai

Harness is the successor I am building to attune-ai. It starts from a constraint attune-ai never had: it runs with no provider SDK and no attune-ai installation. The Attune libraries it does use, forms, claim verification and local retrieval, are part of the install, pinned exactly, and each loads only when the command that needs it runs. Installing an extra does not call a model; invoking a paid provider path requires separate permission and credentials.

Harness can read and save memory through its documented paths. Attune AI still provides the hydrate writer and broader Claude Code and multi-agent workflows; Harness does not replace those today. If you need them, keep attune-ai in its own environment: the two pin different lines of the MCP SDK and cannot share one.

The migration guide lists each journey, its current boundary and what to keep using while a successor is qualified.

Apache License 2.0.

Built by Patrick Roebuck, working with Codex and Claude.

Release files for attune-harness 1.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for attune-harness 1.2.0
File Size Uploaded
attune_harness-1.2.0.tar.gz 355.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for attune-harness 1.2.0
File Interpreter ABI Platform
attune_harness-1.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 758.0 kB

Release files / attune_harness-1.2.0.tar.gz

Download URL attune_harness-1.2.0.tar.gz
Size 355.6 kB
Tags Source
SHA-256 checksum
How to use checksums
2f36d6cd492738a8c737c493e222c9a66d252a9fe89d38f1bdb0a554bea638bc
BLAKE2b-256 checksum
How to use checksums
2272bfc6ed3aef04e999836f0f88cdecde0fc9d7495b65b506cf5b42e7976e7c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / attune_harness-1.2.0-py3-none-any.whl

Download URL attune_harness-1.2.0-py3-none-any.whl
Size 402.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
17c3f0953d136314b199af3b72d7ce6773c9766066ac8c086742aa0d863351f7
BLAKE2b-256 checksum
How to use checksums
6044d981a9ea77ccc8b444c55848430a4adf7e66655bb8c20d85070fb971d191
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.2.0 This release

2 release files

1.1.0

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page