Skip to main content

Production readiness platform for AI agent pipelines — detects silent failures, captures full state, enables step-level replay.

Project description


Website PyPI version Python 3.9+ Beta

Catch silent failures in AI agent pipelines before production.

Your LangGraph pipeline runs fine — no exception. But three nodes later, something crashes with a KeyError. The real cause? A node upstream silently dropped a field. ARGUS catches this.


Install

pip install argus-agents

Quick Start

from argus import ArgusWatcher

watcher = ArgusWatcher(graph)       # attach to your StateGraph
app = graph.compile()
result = app.invoke(initial_state)
watcher.finalize()                  # persist the run to .argus/runs/

ARGUS monitors every node, detects failures, and saves the run. No changes to your node functions.

Always call watcher.finalize() after app.invoke(). Required for cyclic graphs, safe for all. Without it the run stays in memory and won't appear in argus list or the dashboard.


What It Catches

Problem Example
Silent failures Node returns {} or drops a required field — no exception, pipeline keeps running broken
Semantic failures Output structure is fine but values are wrong (placeholders, refusals, degraded text)
Loop stalls Agent retries 5 times producing identical output — stuck loop burning tokens
Unnecessary retries Loop produces correct answer on attempt 2, but validator forces 3 more iterations
Crash root cause Traces KeyError at node 5 back to the upstream node that actually dropped the field
Contract violations Output types don't match the next node's expected input schema

Detection Layers

Runs in order, each more expensive — only fires when needed:

  1. Heuristics — 150+ failure signatures (placeholders, empty results, error keys, semantic degradation). Zero cost.
  2. Validators — custom per-node business-logic constraints. Deterministic.
  3. Anomaly detector — statistical checks for output size anomalies, timing outliers. Deterministic.
  4. Correlator — traces failure propagation across nodes. Points at the origin, not the crash site.
  5. LLM semantic judge — evidence-aware final ruling. Receives all signals from layers 1–4 before deciding. Cannot override validator failures or critical anomalies.
  6. LLM investigator — root cause explanations and debugging suggestions. Only on ambiguous failures.
  7. Loop analyzer — LLM analysis for looped nodes: summarizes iterations, detects stalls, flags wasted retries.

Loop-Aware Inspection

Pipelines with loops (LLM -> compiler -> if fail, retry) get special treatment:

  • Earlier iterations that self-corrected are marked retried (not counted as failures)
  • Only the final iteration determines pass/fail
  • LLM analyzes every loop: what went wrong, what changed between attempts, whether retries were necessary
  • Dashboard shows iteration badges, collapse/expand, and natural-language loop summaries

Replay

Fix a bug, re-run from the failing node. Skip upstream nodes entirely:

argus replay <run-id> node_7          # re-run from node_7 onward
argus replay <run-id> node_7 --only   # just that one node
argus diff <rerun-id>                 # compare vs original

External API calls (OpenAI, etc.) are recorded by default — replays are free and deterministic.


Semantic Judge

For subtle quality issues that pattern matching can't catch:

watcher = ArgusWatcher(graph, semantic_judge=True)  # enabled by default

LLM evaluates output quality on every node. Catches wrong tone, unhelpful responses, outdated info. Requires OPENAI_API_KEY.

The judge receives all prior evidence — validator failures, anomaly signals, inspection results — so it rules with full context, not just input/output. Every decision includes an audit trail:

{
  "pass": false,
  "reason": "Validator correctly identified missing resolution_ticket",
  "confidence": 0.85,
  "evidence_considered": ["validator:payment_check", "anomaly:BA-003"],
  "overridden_signals": []
}
  • evidence_considered — which prior signals the LLM weighed
  • overridden_signals — which signals the LLM disagreed with (passed despite the flag)

Custom Validators

watcher = ArgusWatcher(graph, validators={
    "classify": lambda o: (o.get("label") in ["yes", "no"], "unexpected label"),
    "*":        lambda o: ("error" not in o, "error key present"),  # runs on every node
})

CLI

argus list                           # all recorded runs
argus show last                      # most recent run
argus show <id>                      # inspect a specific run
argus inspect <id> --step <node>     # dump raw input/output for a node
argus replay <id> <node>             # re-run from a node
argus diff <id-a> <id-b>             # compare two runs
argus ui                             # web dashboard
argus doctor                         # check setup health
argus login                          # sign in for cloud sync
argus logout                         # clear stored credentials
argus whoami                         # show current login status
argus update                         # check for newer release

Web Dashboard

argus ui    # opens at localhost:7842

Shows all runs, node-level detail, AI analysis, replay diffs, loop iteration badges, and comparison views. No account needed for local use.


Without LangGraph

from argus import ArgusSession

session = ArgusSession()
session.set_edges({"fetch": ["classify"], "classify": ["process"]})

fetch    = session.wrap("fetch",    fetch_fn)
classify = session.wrap("classify", classify_fn)
process  = session.wrap("process",  process_fn)

state = fetch(initial_state)
state = classify(state)
state = process(state)
session.finalize()

Works with any framework — Prefect, Temporal, plain Python.


Requirements

  • Python 3.9+
  • LangGraph 0.2+ (only for ArgusWatcher)
  • OPENAI_API_KEY in env for semantic features (optional — all heuristic detection works without it)

v0.8.6changelog

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

argus_agents-0.8.6.tar.gz (2.9 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

argus_agents-0.8.6-py3-none-any.whl (2.9 MB view details)

Uploaded Python 3

File details

Details for the file argus_agents-0.8.6.tar.gz.

File metadata

  • Download URL: argus_agents-0.8.6.tar.gz
  • Upload date:
  • Size: 2.9 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for argus_agents-0.8.6.tar.gz
Algorithm Hash digest
SHA256 ced3b372fa002448863fc64e8f65d7b716e097204e413c704b0783999cdcc52e
MD5 e075b512a71d803e3d6a7d53c36d2396
BLAKE2b-256 b37c77c1c8acfd287dc7b167b20819aa94f35a472bb86a4b1ee71e9ee15b72b0

See more details on using hashes here.

Provenance

The following attestation bundles were made for argus_agents-0.8.6.tar.gz:

Publisher: publish.yml on VaradDurge/ARGUS

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file argus_agents-0.8.6-py3-none-any.whl.

File metadata

  • Download URL: argus_agents-0.8.6-py3-none-any.whl
  • Upload date:
  • Size: 2.9 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for argus_agents-0.8.6-py3-none-any.whl
Algorithm Hash digest
SHA256 107dd76bcea1964fb7d9a75f166164ab70baa377834f05f3d37f1dbe5b31ba92
MD5 fb4c6c6d097215a71202525a9e84d1a3
BLAKE2b-256 40724c15ca72153d2c7272b8bc3b6333c611dc9d1414c76193b5468eae37f13d

See more details on using hashes here.

Provenance

The following attestation bundles were made for argus_agents-0.8.6-py3-none-any.whl:

Publisher: publish.yml on VaradDurge/ARGUS

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page