Skip to main content

Internal Affairs 4 AI 🕵️

AI agent forensics & cost governance. Instrument agents, reconstruct why they chose what they did — and cut cost by swapping expensive models for cheaper ones that produce the same result, per job, per tool, per action.

"Police investigation" for autonomous systems: behavioral forensics + a budget audit.

What it does

  1. Connect to any agent system (LangGraph first, then any OpenTelemetry GenAI source via the OTLP endpoint).
  2. Reconstruct the case file — what the job was, what the agent claimed to do, what it actually did, and what independent artifacts prove it (the claim / behavior / provenance triad).
  3. Emit findings with an explicit epistemic level, so we never present a hypothesis as an observation.
  4. Audit cost and recommend when a cheaper model is a credible candidate — and, just as importantly, say "no clear winner" when the evidence is weak.
  5. Diff two runs (git-for-agents) to see exactly which decisions, tools, and models changed — and classify the root cause of a failed or anomalous run.
  6. Enforce policies — configurable governance rules (max cost, PII presence, expensive-model justification, tool allow/deny lists) that emit violations as findings.
  7. Evaluate rigorously — multi-seed golden-set evaluation (≥3 seeds) with mean ± std, so the "cheaper model?" verdict is statistically honest, not a single-run guess.
  8. Evaluate live — plug a real model client (OpenAI-compatible) into the evaluator and answer "is gpt-4o-mini good enough for this job?" against live models.
  9. Stay current — a syncable model catalog (ia catalog --sync <url>) keeps prices, tiers, context windows, and new models up to date.
  10. Authenticate & authorize — JWT + API-key + OIDC (RS256/JWKS) auth with RBAC roles (admin / investigator / viewer).
  11. Route intelligently — cascade recommender: try the cheap model first, escalate to the expensive one on low confidence, with projected savings.
  12. Judge safely — an injection-screened, structured-input LLM judge (sandboxed, no tools) that treats its own output as evidence, never a verdict.
  13. Ingest asynchronously — OTLP batches are store-and-ack'd; a background worker builds case files off the request path and enforces retention (IA_TRACE_RETENTION).
  14. Reconcile the bill — compare metered cost against Anthropic/OpenAI usage reports to surface unaccounted (shadow) usage.
  15. Alert on findings — HMAC-signed webhooks fire when a case meets a severity threshold.

The core value: per-job, tool-aware cost cutting

For every job, the system reads which tools the agent used and which models it called, then recommends the cheapest model that can produce the same result:

  • Capability-matched — filters candidates by tool-calling support, context window, and deprecation, so you never swap to a model that can't do the job.
  • Token-mix priced — candidates are ranked by projected cost on the job's actual input/output token counts, not list prices: an input-heavy job (RAG, long-document summarization) ranks input price first; a generation-heavy job ranks output price first.
  • Quality-verified — with a golden set + runner, each candidate is run through the multi-seed evaluator; a cheaper model that fails the task is skipped, and the cheapest tied-or-better candidate wins.
  • Free included — zero-cost models (e.g. NVIDIA NIM open models) are ranked first, so "use the free model" is the default suggestion when it holds up.
from internal_affairs.verdict.cost_cutter import analyze_cost_cuts

report = analyze_cost_cuts(case)  # capability-matched
report = analyze_cost_cuts(case, golden_set=golden, runner=runner, seeds=5)  # quality-verified

GET /case/{id}/cost-cuts returns the same report for the dashboard.

Security-first

Redaction happens at ingest, before storage: secrets and PII never reach the store or the (future) LLM judges. A tamper-evident, hash-chained, signed evidence log detects any post-hoc alteration. See SECURITY.md and docs/threat-model.md.

Install the server (Python)

The server is a plain Python package. It installs a console script also named ia (ia serve, ia demo, ia catalog) — the Go client is a different program with the same name, covered in CLI (Go client) below.

uv venv && source .venv/bin/activate
uv pip install -e ".[dev]"          # core + dev/test deps
uv pip install -e ".[langgraph]"    # optional: LangGraph adapter

Run the demo

python -m examples.demo     # full pipeline: record → redact → investigate → verdict
ia catalog                 # list the model catalog
ia catalog --sync <url>    # merge new models/prices from a remote registry

Run the API

cp .env.example .env        # set IA_EVIDENCE_SIGNING_KEY in prod
ia serve                    # or: uvicorn internal_affairs.api.app:app --reload

Web UI: open http://127.0.0.1:8000/ for a dashboard (cases, findings/verdicts, model catalog, audit log, activity feed, run graph, terminal) served by FastAPI — no build step.

CLI (Go client)

A Go client for the same API — ia — in cli/. It is a client, never a server and never a shell: every command maps onto one REST route, and the server decides what a credential may do.

ia is two different programs. The Python package installs a console script called ia (ia serve, ia demo, ia catalog …) — that is the server. This Go binary is also called ia (ia doctor, ia case, ia run, ia smoke-test …) — that is the client. Whichever directory comes first on $PATH wins, so pick one to own the name, or rename one: make -C cli build && mv cli/bin/ia /usr/local/bin/iactl.

Option 1 — install a released binary (no Go toolchain needed). Recommended:

curl -fsSL https://raw.githubusercontent.com/smoolzone/ia4ai/main/cli/scripts/install.sh | sh

The script detects linux/darwin × amd64/arm64, installs to /usr/local/bin (or ~/.local/bin when that is not writable), and refuses to install unless the archive's SHA256 matches the release's SHA256SUMS.

This repository is private, and GitHub answers 404 (not 403) to a client without a token — for the raw script, the release assets and the API alike. Export a token with read access first and pass it on the script fetch:

export GITHUB_TOKEN=...   # any token with read access to this repo
curl -fsSL -H "Authorization: Bearer $GITHUB_TOKEN" \
  https://raw.githubusercontent.com/smoolzone/ia4ai/main/cli/scripts/install.sh | sh

Pin a version, or choose the prefix, through the environment:

VERSION=v0.1.0 PREFIX="$HOME/.local" sh -c \
  'curl -fsSL https://raw.githubusercontent.com/smoolzone/ia4ai/main/cli/scripts/install.sh | sh'

VERSION defaults to the latest published release (v0.1.0 is live); REPO (owner/repo) overrides where the archive is fetched from. This path needs a published release — before v0.1.0 existed the script stopped with could not determine the latest release. On a private repo with no token it now says so directly instead of failing on a bare 404.

Option 2 — build from source (needs Go 1.22+).

make -C cli install                        # -> /usr/local/bin/ia
make -C cli install PREFIX="$HOME/.local"  # no sudo -> ~/.local/bin/ia
make -C cli build                          # just compile -> cli/bin/ia

Build through the Makefile, not a bare go build: the Makefile is what injects the Version/Commit/BuildDate metadata, so a bare build reports dev forever. make -C cli dist cross-compiles every supported platform into cli/dist/.

Then confirm which binary you got — ~/.local/bin is not on $PATH by default:

ia version                          # prints the release tag, or "dev" for a bare build
ia init --base-url http://127.0.0.1:8000
ia login                            # stores a credential, encrypted at rest
ia doctor                           # config, connectivity, role
ia smoke-test                       # api → credential → ingest → read-back → priced
ia case list
ia run show <trace|case>

Credentials are AES-256-GCM sealed under a local master key; --workspace selects a stored credential (the server binds tenant to the credential, so it never widens scope). ia console is dispatched against the same server-side allow-list as the dashboard terminal. The release workflow cross-compiles linux/darwin × amd64/arm64 on v* tags; the module path lives in one Makefile variable if it needs repointing.

Deploy with PM2

No containers required — the app is a plain Python package, so PM2 is used only as a supervisor (autostart, crash-restart, log rotation). A ready-to-edit process file ships as ecosystem.config.js.

npm i -g pm2                        # PM2 is only a process manager, not a runtime
pm2 start ecosystem.config.js       # start the API (ia serve)
pm2 save                            # persist the process list
pm2 startup                         # print/enable the boot hook, then re-run its command
pm2 logs internal-affairs           # tail ./logs/ia.{out,err}.log
pm2 reload internal-affairs         # zero-downtime restart after a deploy

Set real values via the environment (or a .env next to the process file) rather than committing secrets: at minimum IA_EVIDENCE_SIGNING_KEY must be set in production (python -c "import secrets; print(secrets.token_hex(32))").

Scaling. PM2's cluster mode (-i) is a Node feature and does not apply to Python. Keep instances: 1 with the default SQLite backend (IA_DB_PATH) — SQLite is single-writer, and the tamper-evident evidence/access logs rely on a single store. To run multiple workers, switch to Postgres (IA_DB_DSN, the [postgres] extra) and use uvicorn's own --workers N (the commented override in ecosystem.config.js). Either way, terminate TLS at a reverse proxy (nginx/Caddy) in front of the loopback bind.

docker-compose.yml is an optional dev extra (a local Postgres for the Postgres backend only) — it is not part of the app's runtime and can be ignored entirely.

Endpoints:

  • POST /ingest/trace — ingest a raw trace (+ optional job) → normalize, redact, log, store.
  • POST /ingest/otlp — ingest OpenTelemetry Protocol (JSON or Protobuf) trace batches from any OTel source. Also served at the standard OTLP/HTTP path /v1/traces, so a stock SDK pointed at the base URL works unchanged. Store-and-ack: spans are persisted and queued, and the case file is built asynchronously by the worker.
  • GET /case/{case_id} — fetch a reconstructed case file (read is audit-logged).
  • GET /case/{case_id}/cost-cuts — recommend cheaper capable models for a job (the core value).
  • GET /diff/{case_a}/{case_b} — diff two case files (decisions/tools/models/cost).
  • GET /catalog — the model catalog (providers, tiers, prices, capabilities, notes).
  • POST /catalog/models — add/update a model (admin) so new models are live immediately.
  • POST /catalog/sync — sync the catalog from configured feed URLs (admin).
  • GET /audit/evidence-log/verify — prove the evidence log has not been tampered with.
  • GET /audit/access-log — who accessed what (immutable, signed access audit).
  • GET /audit/access-log/verify — prove the access audit log has not been tampered with.
  • GET /trace/{trace_id} — the span DAG a run actually executed (parent/child + timings; read is audit-logged).
  • GET /case/{case_id}/trace — the same, resolved from a case (tenant-checked via the case).
  • POST /graph/register — register a static agent blueprint, e.g. extract_graph_topology(compiled_app).
  • GET /graph/{agent_system} — that blueprint back (the map, not the journey; opt-in, 404 is normal).
  • GET /activity — the pipeline activity feed: the evidence + access logs merged and ordered (admin).
  • GET /activity/stream — the same feed as Server-Sent Events, for the live dashboard terminal (admin).
  • POST /console — the dashboard terminal: one line, dispatched against an allow-list of read-only commands over the same store the API reads (no shell, no subprocess, no writes). Refused with 403 for commands your role does not hold; the line is redacted and appended to the access log either way.
  • GET /cost/summary — aggregate cost: per-model, per-day, per-job rollups + trend drift (viewer).
  • GET /reconcile — compare metered usage vs the provider bill (Anthropic/OpenAI admin usage APIs) (admin).
  • POST /golden-set/bootstrap — build a replay golden set from stored production traces (viewer).
  • GET /alerts/history — recent webhook alert dispatches (viewer).
  • GET /pipeline/stats — ingestion-worker queue/processing/retention stats (viewer).
  • GET /health — liveness.

Test

uv run pytest

The dashboard is inline HTML/JS with no build step, so its Terminal and Run graph tabs are covered by a real headless-browser click-through (tests/test_ui_clickthrough.py) that also fails on any JavaScript error the page logs. It skips unless both playwright and a running server are present:

uv pip install -e ".[ui]" && playwright install chromium
uv run pytest tests/test_ui_clickthrough.py    # needs `ia serve` running

Project layout

internal_affairs/
  schemas/        canonical evidence + case-file models (pydantic)
  security/       redaction + tamper-evident evidence/audit logs
  ingest/         normalization (OTel GenAI semconv) + LangGraph + OTLP adapters
  investigate/    case-file reconstruction + findings + root cause + diffing
  verdict/        cost attribution + "cheaper model?" + policy + evaluation
                  + runner + cascade + judge + cost cutting
  storage/        Store interface: in-memory + SQLite + Postgres
  api/            FastAPI control plane + web UI
examples/         runnable demo + seed data
docs/             architecture + threat model + integration guide

Persistence: set IA_DB_PATH to a .db file for SQLite-backed storage (cases, traces, evidence log, audit log all persist and stay tamper-evident across restarts). Leave it empty for in-memory (dev/tests).

Auth/RBAC: disabled by default. Set IA_AUTH_ENABLED=true and configure IA_AUTH_SECRET (JWT), IA_API_KEYS (key:role:tenant), and/or OIDC (IA_OIDC_AUDIENCE + IA_OIDC_ISSUER). Endpoints require investigator (ingest), viewer (read/diff), or admin (audit logs).

Status

Phase 3 complete (core loop), plus the verdict-honesty layer: non-inferiority three-way verdicts (viable / ruled_out / unproven), governance + latency gating, verdict lineage, capture→replay (ia reverify), and metered verification cost. An operations layer adds an async ingest pipeline (store-and-ack → worker), provider-bill reconciliation, signed webhook alerting, aggregate cost rollups with trend drift, and golden-set bootstrap. Deployment is PM2-based (ecosystem.config.js) — no containers required. See docs/architecture.md for the full roadmap.

Metadata

Release files for internal-affairs-4-ai 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for internal-affairs-4-ai 0.1.0
File Size Uploaded
internal_affairs_4_ai-0.1.0.tar.gz 152.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for internal-affairs-4-ai 0.1.0
File Interpreter ABI Platform
internal_affairs_4_ai-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 288.9 kB

Release files / internal_affairs_4_ai-0.1.0.tar.gz

Download URL internal_affairs_4_ai-0.1.0.tar.gz
Size 152.4 kB
Tags Source
SHA-256 checksum
How to use checksums
c5d39a8a1404e890ac35ed25ea97a2abdc380a156f83fa3a6a002fb4740dcd0f
BLAKE2b-256 checksum
How to use checksums
1e4e2342cf31622471658f3d7082be0e077b79a3aaeb9f57b8e30c1710e33112
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.9

Release files / internal_affairs_4_ai-0.1.0-py3-none-any.whl

Download URL internal_affairs_4_ai-0.1.0-py3-none-any.whl
Size 136.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
800a6a9eb13b9b2f7961b9761e5ee0f5993c04d5168a433b29ffa2bb8ed6139f
BLAKE2b-256 checksum
How to use checksums
9c354deba00799e48a4fe85551f4241f810cdca996595d0db3657600d7c6c73d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.9

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page