Skip to main content

Agent Chaos

Chaos engineering for autonomous AI agents.

AI agents increasingly depend on unreliable model APIs, tools, databases, and HTTP services. Agent Chaos intentionally disrupts those dependencies so developers can measure whether an agent-like workload tolerates a fault, retries successfully, or fails.

Agent Chaos is an early open-source vertical slice. It is framework-independent and does not require an OpenAI or Anthropic API key.

flowchart LR
    A["Agent workload"] --> C["Agent Chaos proxy"]
    C --> D["HTTP dependency"]
    C -. "inject fault" .-> C

Install

The published Python distribution is named agent-chaos-runner; the product and installed command remain Agent Chaos and agentchaos.

uv tool install agent-chaos-runner
agentchaos --version

Agent Chaos supports Python 3.12+ on macOS and Linux.

Quick start

Clone the repository to run the deterministic local demo without API keys:

git clone https://github.com/Coroz2/agent-chaos.git
cd agent-chaos
uv sync --extra dev --locked
uv run agentchaos run examples/scenarios/api_503_recovery.yaml

The scenario starts its deterministic fake dependency automatically. A successful run ends with RECOVERED and writes its artifacts under .agentchaos/runs/<run-id>/.

Other examples:

uv run agentchaos run examples/scenarios/no_fault.yaml
uv run agentchaos run examples/scenarios/api_latency_recovery.yaml
uv run agentchaos run examples/scenarios/api_429_recovery.yaml
uv run agentchaos run examples/scenarios/api_429_failure.yaml
uv run agentchaos run examples/scenarios/api_503_failure.yaml
uv run agentchaos run examples/scenarios/api_503_schedule_recovery.yaml
uv run agentchaos run examples/scenarios/api_503_schedule_incomplete.yaml
uv run agentchaos run examples/scenarios/http_disconnect_recovery.yaml
uv run agentchaos run examples/scenarios/http_disconnect_failure.yaml
uv run agentchaos run examples/scenarios/http_malformed_json_recovery.yaml
uv run agentchaos run examples/scenarios/http_malformed_json_failure.yaml

The 429, 503, disconnect, malformed-JSON, and incomplete-schedule examples deliberately exit with status 1 because the experiment's required recovery or schedule completion is not observed.

Scenario

schema_version: 1
name: api-503-recovery

dependency:
  type: http
  base_url: http://127.0.0.1:19103
  start:
    command: [python, fake_api.py, --port, "19103"]
    cwd: ..
    readiness:
      path: /health

workload:
  name: demo-agent
  command: [python, demo_agent.py]
  cwd: ..
  proxy_url_env: CUSTOMER_API_URL

fault:
  type: http_error
  target:
    method: GET
    path: /customer/*
  trigger:
    occurrence: 2
  status_code: 503

success:
  exit_code: 0

The optional managed dependency is intended for local tests. Omit dependency.start when the upstream already exists. Agent Chaos always exposes the generated proxy URL as AGENTCHAOS_PROXY_URL; proxy_url_env maps it into the variable an existing workload expects.

Agent Chaos also supports deterministic HTTP rate-limit injection:

fault:
  type: http_rate_limit
  target:
    method: GET
    path: /customer/*
  trigger:
    occurrence: 2
  retry_after_seconds: 1

The selected request receives HTTP 429 with an integer Retry-After value and X-Agent-Chaos-Fault: http_rate_limit; the upstream is not contacted for that request.

To test application-level JSON handling despite a successful HTTP transport status, configure the fixed malformed-JSON fault:

fault:
  type: http_malformed_json
  target:
    method: GET
    path: /customer/*
  trigger:
    occurrence: 2

On the selected occurrence, Agent Chaos returns status 200 with Content-Type: application/json, X-Agent-Chaos-Fault: http_malformed_json, and one fixed invalid JSON body. It does not contact the upstream or accept configurable response content. A matching retry that receives valid JSON is classified as recovery; exiting without a successful matching retry is a failed experiment.

To test recovery from an abruptly terminated HTTP connection, configure the disconnect fault:

fault:
  type: http_disconnect
  target:
    method: GET
    path: /customer/*
  trigger:
    occurrence: 2

Agent Chaos does not forward the selected request upstream and terminates the client connection before a complete response arrives. The portable guarantee is a client-visible transport or HTTP protocol failure; a literal TCP reset and a particular client-library exception are not guaranteed. A matching retry that succeeds through the upstream is classified as recovery.

To inject the same configured fault more than once, use an occurrence schedule:

fault:
  type: http_error
  target:
    method: GET
    path: /customer/*
  trigger:
    occurrences: [2, 4]
  status_code: 503

Schedule entries must be positive, unique, and strictly increasing. The scalar occurrence: 2 form remains supported as a one-entry schedule. Target-matching retries count toward the schedule and can themselves be selected for injection. A schedule is complete only when every configured occurrence injects; reaching only part of it fails with FAULT_SCHEDULE_INCOMPLETE even if every observed failure recovered.

Results

  • PASSED: a baseline succeeds, or the workload tolerates injected latency without failure.
  • RECOVERED: every fault-related failure has its own successful matching retry, and the workload succeeds.
  • FAILED: execution fails, the fault never fires, a schedule is incomplete, or any required recovery is not observed.

An injected disconnect always records a failed operation. It produces RECOVERED only when a matching retry succeeds; an expected workload exit without that retry produces FAILED.

Successful and recovered experiments exit 0. Experiment failures exit 1, invalid scenarios and invalid saved reports exit 2, setup or unexpected internal failures exit 3, and interruptions exit 130. Inspecting a structurally valid saved report exits 0 even when its recorded experiment result is FAILED.

Every valid run contains:

.agentchaos/runs/<run-id>/
├── scenario.yaml
├── events.jsonl
├── stdout.log
├── stderr.log
├── dependency.stdout.log
├── dependency.stderr.log
└── report.json

events.jsonl remains a schema-1, sequence-ordered event stream. New runs write schema-2 report.json files with the configured and completed occurrence schedules plus one evidence row per fault-related failed operation. Each row identifies its successful retry and recovery latency, or uses null values when recovery was not observed.

Commands

uv run agentchaos --help
uv run agentchaos --version
uv run agentchaos version
uv run agentchaos validate examples/scenarios/api_503_recovery.yaml
uv run agentchaos run examples/scenarios/api_503_recovery.yaml
uv run agentchaos inspect .agentchaos/runs/<run-id>
uv run agentchaos inspect .agentchaos/runs/<run-id>/report.json

Use --output-dir PATH to place run directories somewhere other than .agentchaos/runs.

inspect accepts either a run directory or its report.json file and prints the same result summary as run. It strictly validates schema-1 and schema-2 saved reports without running a workload, starting a dependency, contacting the network, reading other artifacts, migrating the report, or modifying the run directory.

Documentation

Development

uv sync --extra dev
uv run pytest
uv run ruff check .
uv run ruff format --check .
uv run mypy

See CONTRIBUTING.md for the branch, pull request, verification, and release workflow.

Limitations

Agent Chaos supports one HTTP dependency and zero or one fault on macOS and Linux. It is a reverse proxy, not transparent network interception: the workload must accept the proxy base URL through its configuration. Request bodies and responses are buffered up to 10 MiB. Disconnect injection does not provide packet-level reset controls or partial-response faults. Streaming, SSE, WebSockets, CONNECT tunneling, TLS interception, multiple faults, probabilistic triggers, and model-specific grading are not implemented.

Retry classification uses a deterministic fingerprint of method, path, hashed query, and body. It is useful black-box evidence, not proof of the workload's internal intent.

Roadmap

The next logical steps include trigger windows and reproducible probabilistic policies, multi-fault campaigns, then another dependency adapter such as MCP. These are broad, nonbinding directions; detailed release scope begins only in an approved version specification.

Licensed under Apache-2.0.

Metadata

Release files for agent-chaos-runner 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-chaos-runner 0.4.0
File Size Uploaded
agent_chaos_runner-0.4.0.tar.gz 139.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-chaos-runner 0.4.0
File Interpreter ABI Platform
agent_chaos_runner-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 182.2 kB

Release files / agent_chaos_runner-0.4.0.tar.gz

Download URL agent_chaos_runner-0.4.0.tar.gz
Size 139.7 kB
Tags Source
SHA-256 checksum
How to use checksums
38c44010116c77bd83c0d1bc1d2015b4b50fc2fa24f42abab5857c8214e34ef1
BLAKE2b-256 checksum
How to use checksums
9b863bca4dfb7e851b5cd9479aa91bd605a7ee18b20c5c1c84012a45d9d13d3c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 20, 2026.

Transparency log

Release files / agent_chaos_runner-0.4.0-py3-none-any.whl

Download URL agent_chaos_runner-0.4.0-py3-none-any.whl
Size 42.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2742a9bc215baf4ccbe6c604e8700c123b2f61be0f89574974b235c44003b4ac
BLAKE2b-256 checksum
How to use checksums
c5125907a72aa712d46313f9ee76b3838c91079e56377504a241c2617d38232a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 20, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.0

2 release files

This release

0.4.0 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page