Skip to main content

Argus

An autonomous agent that finds one class of vulnerability — IDOR / broken access control — and proves each finding by generating and executing a real exploit in an isolated sandbox. It returns a verdict with evidence: exploit confirmed, or false positive rejected.

The proof loop is the product. Argus doesn't flag "possible" issues and leave you to sort through them — it tries the attack, and shows you what happened.


The idea: prove it, don't guess

Most scanners produce a long list of maybes. Real issues get buried under false alarms, and a human has to manually confirm each one.

Argus takes the opposite bet on a single, narrow problem and goes deep:

  1. It finds candidate access-control issues.
  2. It generates a concrete proof-of-concept exploit.
  3. It runs that exploit against an authorized target inside a sandbox that can reach only that target.
  4. It returns a verdict with the evidence that justifies it:
    • ✅ Confirmed — here is the exact request/response showing one user holding another user's private data.
    • ❌ Rejected (false positive) — the access control held; here's the 403. Not a bug.

That second outcome — confidently, verifiably saying "this is not a vulnerability" — is the signature behavior, and the hard part almost nothing does well.

The bug it hunts: IDOR / broken access control

IDOR (Insecure Direct Object Reference) is when an app lets you reach someone else's object just by changing an identifier in the request — e.g. opening /api/invoices/124 when only /api/invoices/123 is yours, because the handler forgot to check "does this belong to the caller?" It is one of the most common and most serious bugs in fast-built apps, precisely because that ownership check is so easy to omit.

Benchmark result

Argus ships with a deliberately-vulnerable target that contains real holes and decoys (endpoints that take an object id and look exploitable but are correctly authorized). On that benchmark:

Confirmed Rejected
Real IDORs (6) 6 ✅ 0
Decoys / traps (6) 0 6 ✅

Precision 1.00 · Recall 1.00 · Zero false positives. For each confirmed finding, Argus also drafts a remediation suggestion.

Honest caveat: this is a controlled target with a known answer key — on purpose. That's how you plant decoys and get a real precision number. Crucially, the oracle decides from behavior alone (did the attacker actually receive the victim's data?); the answer key is used only to score the run, never during detection.

How it works

 CLI ─▶ Agent (single LangGraph state machine) ─▶ Runner (Docker, egress-locked) ─▶ Target (bundled vuln app)
            │                                           ▲
            ├── LLM (provider-agnostic) ────────────────┘   proposes candidates + PoCs
            └── Oracle (deterministic) ── verdict + evidence ─▶ Trace (JSON)

The loop, as explicit steps: recon → candidates → PoC synthesis → execute → classify (oracle) → remediate.

Three design choices carry the whole thing:

  • The LLM proposes; a deterministic oracle disposes. The model may hypothesize candidates and write exploits, but it never renders the verdict — a plain, deterministic function does, from evidence. The AI never grades its own homework.
  • One agent, one tight loop — not a swarm. LangGraph models the loop as an explicit state machine of steps, which stays legible and matches the single-family scope.
  • The guardrail lives in the network, not a prompt. Exploits run in a Docker container on an --internal network (no gateway to the internet). By construction, the sandbox can reach the authorized target and nothing else — and a self-check proves the internet is unreachable from inside.

Quickstart

Requires uv. Docker is needed only for the full sandboxed run.

uv sync                                 # create the env (Python 3.11+)

uv run pytest                           # the test suite (no Docker, no API key)
uv run python -m evals.harness          # the scorecard: precision/recall + confusion matrix

# Full end-to-end run (needs Docker running + an LLM key, see below):
cp .env.example .env                     # then put your key in .env
uv run python -m cli.argus scan          # builds the sandbox, runs the loop, writes traces/run-*.json

The LLM is provider-agnostic (any OpenAI-compatible API). The default is DeepSeek — set DEEPSEEK_API_KEY in .env (see .env.example). The scorecard and tests need no key and no Docker.

Project layout

target/   the deliberately-vulnerable Flask app + its ground-truth manifest (the answer key)
agent/    the loop: schemas, LLM client, the 6 nodes, the LangGraph graph, and the oracle
runner/   the Docker sandbox + the network-level egress guardrail
cli/      `argus scan`
evals/    the scorecard (precision/recall + confusion matrix)
traces/   captured run evidence (JSON)

Scope — narrow on purpose

Argus is deliberately one vulnerability family, done deeply. These are explicit non-goals, not missing features:

  • ❌ Other vulnerability families (until the IDOR loop is genuinely solid and measured).
  • ❌ A "graph of agents" swarm — one agent, one orchestrated loop.
  • ❌ Pointing the tool at arbitrary live targets — it attacks only the bundled/authorized target. This is a hard legal + safety guardrail, enforced at the network layer.

Roadmap

  • Vulnerable target + ground-truth manifest
  • The validation loop (recon → … → oracle), CLI, end-to-end
  • Docker sandbox + egress guardrail
  • Eval harness (precision/recall + confusion matrix)
  • Replay dashboard (web UI that replays a captured run)
  • CI (run the scan on push, publish the evidence artifact)

Contributing

Contributions are welcome — please read CONTRIBUTING.md first, especially the scope section (it's what keeps the project focused). Good first areas: harder target endpoints and decoys, more eval cases, the dashboard, and new object-reference patterns for recon to discover.

License

MIT © 2026 Vatsalya Soni.

Metadata

Release files for argus-idor 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for argus-idor 0.1.0
File Size Uploaded
argus_idor-0.1.0.tar.gz 164.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for argus-idor 0.1.0
File Interpreter ABI Platform
argus_idor-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 205.4 kB

Release files / argus_idor-0.1.0.tar.gz

Download URL argus_idor-0.1.0.tar.gz
Size 164.9 kB
Tags Source
SHA-256 checksum
How to use checksums
f37e27124cfe71e4decd009e44af8e064f7d703dda0079485c0968b002c96fb3
BLAKE2b-256 checksum
How to use checksums
e2b0297c7a1a3829bf4eebc5bfc058ae51aef122349f628c65ae4ecf83cdcfe7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.5 {"installer":{"name":"uv","version":"0.10.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / argus_idor-0.1.0-py3-none-any.whl

Download URL argus_idor-0.1.0-py3-none-any.whl
Size 40.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
890d685dbb0b0be7c83664521c394c15af814250bfa438a193556f6a8ab5d9b3
BLAKE2b-256 checksum
How to use checksums
7ebba87d9915087c68d32fb95f089b274eea2ddbed0002eff6f38ac3ece5b057
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.5 {"installer":{"name":"uv","version":"0.10.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page