Skip to main content

Find where an AI agent run started going wrong, not just where it stopped

Project description

Runopsy

Find where an AI agent run started going wrong — not just where it stopped.

CI PyPI Python Licence

pip install runopsy && runopsy demo

That runs a worked example end to end: no agent, no API key, no configuration.


The problem

When an agent run fails, the last error is rarely the problem. A config written wrongly at step 9 surfaces as a failing test at step 14, and reading the log bottom-up sends you to fix the test.

Runopsy records agent runs, localizes the step where things actually broke, shows the evidence behind that claim, and can re-run the trace with one thing changed to test it.

It runs locally, spends no tokens for its core analysis, and never claims a cause it has not validated.

What it looks like

Run demo_run — fix the failing integration test in the payments service
demo · 16 events · failure

Observed failure  (what the run visibly got wrong)
  step 14 pytest
  tool 'pytest' failed with exit code 1

Suspected onset  (where it may have started going wrong, unverified)
  step 9 write_config
  tool 'write_config' failed with exit code 1
  47% confidence, unverified
  may have affected step 10 edit_file, step 11 restart_service, and 3 more
  evidence: runopsy evidence demo_run --step 9

  No cause has been confirmed. To test this candidate, replay from it:
    runopsy replay demo_run --from-step 9

Quick start

pip install runopsy          # or: uv tool install runopsy / pipx install runopsy

runopsy demo                 # a worked example, no setup
runopsy record -s "make" -s "pytest"    # wrap any pipeline you already run
runopsy diagnose latest      # where it started going wrong
runopsy evidence latest --step 9        # the command, the output, why it was flagged
runopsy ui                   # timeline and failure map in a browser, loopback only

Driving an agent and diagnosing it in one command:

runopsy adapter hermes       # prints the config block to paste
runopsy run "fix the failing test"

What it is, and is not

It is not a coding assistant and does not replace the one you use. Runopsy has no chat, writes no code and makes no suggestions. It attaches to whatever already runs your work — an agent, a CI pipeline, a Makefile — records what happened, and tells you where it started going wrong.

you already have Runopsy adds
an agent (Hermes today) runopsy run "task" drives it and diagnoses the session
a pipeline or test suite runopsy record -s "make" -s "pytest" wraps it
Inspect AI eval logs runopsy-inspect import reads them
nothing yet runopsy demo, in one command

There is no model to pick and no key to enter for the core product: deterministic diagnosis spends zero tokens and makes zero network calls. A provider key buys exactly one optional thing — --mode hybrid, which asks a model about the few steps already found suspicious.

Does it work? Four measurements, including the ones that go badly

what was measured result
20 labelled traces, onset top-1 (runopsy bench --compare) 94.4%
the same, versus blaming the last failing step 22.2%
faults injected into a real recorded run (--inject --store) 100%
TRAIL — expert-labelled SWE-Bench agent traces (--trail) 0.0%
Who&When — expert-labelled multi-agent traces (--corpus) 0.0%

The last two are published because they say where this does not work, and that is worth more to you than a single flattering number.

Runopsy localizes onsets that were themselves failures — a step that errored before the visible symptom did. On TRAIL, not one of the 30 annotated onsets carries an error status of any kind: they are formatting mistakes, instruction non-compliance, a wrong assumption about a file path. The deterministic layers read exit codes and tool statuses, so they are blind to those by construction. --mode hybrid exists for that case.

Zero false positives on healthy runs, exactly rather than approximately: a spurious finding is what gets a diagnosis tool switched off.

What it will not do

  • claim a cause it has not validated — a finding stays suspected onset until a replay or a named human says otherwise
  • send anything anywhere without --mode hybrid
  • execute a replay outside a disposable sandbox, or perform a blocked external side effect
  • report a finding on a healthy run

Commands

demo · run · record · runs · diagnose · evidence · replay · graph · export · ui · verify · label · bench · prune · doctor · setup · config · adapter

runopsy --help lists them all; runopsy on its own tells you where this machine stands and what to type next.

Packages

pip install runopsy gives you everything. If you are writing code against it, depend on the piece you need:

package what it is
runopsy-core schema, normalizer, 15 detectors, ranking — framework-agnostic, needs only Pydantic
runopsy-collector JSONL journals, DuckDB index, payload vault, retention
runopsy-replay checkpoint restore and counterfactual execution behind a fail-closed gate
runopsy-adapter shell and Hermes runtime adapters
runopsy-bench labelled cases, metrics, fault injection, external benchmarks
runopsy-semantic the optional paid layer, budget-capped
runopsy-server local API and the bundled web view
runopsy-inspect reads Inspect AI eval logs
pip install "runopsy[inspect]"   # adds the Inspect AI reader

Privacy

Traces, state and artifacts stay on your machine. Command text is kept in a local vault with secrets redacted before anything is written; the trace itself stores hashes only. Journals are sealed as they are written, so runopsy verify can tell you a trace is byte-for-byte the one that was recorded.

Links

Built by Vahit Feryad (ORCID 0000-0002-3282-339X).

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

runopsy-0.1.8.tar.gz (5.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

runopsy-0.1.8-py3-none-any.whl (5.2 kB view details)

Uploaded Python 3

File details

Details for the file runopsy-0.1.8.tar.gz.

File metadata

  • Download URL: runopsy-0.1.8.tar.gz
  • Upload date:
  • Size: 5.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.12.0 {"installer":{"name":"uv","version":"0.12.0","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for runopsy-0.1.8.tar.gz
Algorithm Hash digest
SHA256 5f658ae0b0905e92b862fdfbfaafe99e02838382d2c4d2d1a1a4de174d79b283
MD5 58a640d07dfaac2c340bc864d44033c0
BLAKE2b-256 182383b27ded0d5ebfa3d015648bc3cbe32eadff918e5c2f24798e682de69cba

See more details on using hashes here.

File details

Details for the file runopsy-0.1.8-py3-none-any.whl.

File metadata

  • Download URL: runopsy-0.1.8-py3-none-any.whl
  • Upload date:
  • Size: 5.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.12.0 {"installer":{"name":"uv","version":"0.12.0","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for runopsy-0.1.8-py3-none-any.whl
Algorithm Hash digest
SHA256 3a5e3f861b69e0514c4911d3295e50e21ed139a81af10b0e07088c5533c460ef
MD5 176ca915b19900fb8a61919755bae2d6
BLAKE2b-256 d5d3bb7e35fd6b5adc6e3b68a6c5a0ed708a03505852246d01b4201114c33496

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page