Skip to main content

Scrutineer

Agent Behavioral Testing Platform — Tests what agents DO, not just what they SAY.

Python 3.11+ License: MIT Tests Code style: ruff

The Problem

88% of AI agents fail in production. The dominant failure modes are operational:

  • Tool errors (28%)
  • Memory/state issues (22%)
  • Edge cases (18%)

Yet the entire evaluation ecosystem (DeepEval, LangSmith, MS AGT) focuses on output quality or observability. Nobody tests agent behavior in production-like environments before deployment.

The Solution

Scrutineer fills that gap. It's a behavioral testing platform that:

  1. Mocks your agent's environment — tools, APIs, databases with configurable latency, errors, and rate limits
  2. Injects chaos — tool failures, context degradation, cascading errors, spec drift under pressure
  3. Asserts behavior — 20+ assertions across tool calls, state consistency, governance compliance, resilience, and performance
  4. Reports regressions — structural diffing, baseline comparison, HTML + JUnit reports

Quick Start

Distribution name: this project publishes to PyPI as scrutineer-agents. Do not run pip install scrutineer — that name belongs to an unrelated project and installs a different library. The import package and the console script are both scrutineer.

# Install (zero dependencies by default)
pip install scrutineer-agents

# Or with framework adapters
pip install "scrutineer-agents[adapters]"

# Or with the WebUI dashboard
pip install "scrutineer-agents[web]"

# Run the quickstart example
python examples/langchain_quickstart.py

# Run a YAML scenario
scrutineer run --path examples/basic_scenario.yaml

WebUI Dashboard

Scrutineer includes a browser-based dashboard for running scenarios, viewing traces, and comparing baselines — all wrapping the core Python API.

# Install with web dependencies
pip install "scrutineer-agents[web]"

# Start the dashboard
scrutineer serve

# Or with a custom port
scrutineer serve --port 9090

Then open http://localhost:8080 in your browser.

Features:

  • Dashboard — pass/fail stats, recent runs, quick actions
  • Scenarios — browse, inspect, and run test scenarios
  • Runs — live execution with step-by-step trace visualization
  • Baselines — saved results with regression diff comparison
  • Live Console — real-time log streaming via SSE during test runs

See src/scrutineer/web/README.md for the full WebUI guide.

Architecture

src/scrutineer/
├── env.py          # MockTool, MockAPI, MockDatabase, EnvironmentBuilder
├── chaos.py        # ToolFailureInjector, ContextDegradation, CascadingFailures
├── assertions.py   # 20+ behavioral assertions
├── runner.py       # @scrutineer_test decorator, ScenarioRunner
├── reporting.py    # Regression reports, JUnit XML, HTML
├── baseline.py     # JSON baseline storage with git integration
├── otel.py         # OpenTelemetry span model
├── cli.py          # Full CLI: run, list, info, baseline, diff, report
├── adapters/       # LangChain, CrewAI, OpenAI SDK, Generic
└── web/            # FastAPI WebUI dashboard (optional)
    ├── app.py          # FastAPI application factory
    ├── server.py       # Uvicorn entry point
    ├── api/            # REST API routers (scenarios, runs, baselines)
    ├── services/       # Service layer wrapping core modules
    ├── schemas/        # Pydantic request/response models
    └── static/         # Frontend (HTML, CSS, JS)

The Chaos Module (Differentiator)

Scrutineer's chaos injection is what sets it apart:

  • ContextDegradation — Quadratic acceleration curve matching real context window pressure (last 20% is much worse than first 20%)
  • CascadingFailures — Multi-agent error propagation with dependency graphs (database → api_server → ui)
  • SpecDrift — Agent improvisation under pressure with intensity levels and cumulative drift scoring

No other tool tests these production failure modes.

CLI Commands

scrutineer run <scenario>          # Run a test scenario
scrutineer list                    # List available scenarios
scrutineer info <scenario>         # Show scenario details
scrutineer baseline record         # Record current state as baseline
scrutineer baseline show           # Show recorded baseline
scrutineer diff                    # Compare current vs baseline
scrutineer report                  # Generate regression report
scrutineer trace <run-id>          # Show execution trace
scrutineer serve                   # Start WebUI dashboard
scrutineer serve --port 9090       # Custom port

Framework Adapters

Scrutineer ships adapters for specific frameworks, plus a generic hook adapter for anything else:

# LangChain — rebinds your_agent.tools so the agent's own call path hits the mocks
from scrutineer.adapters.langchain import wrap_agent
wrapped = wrap_agent(your_agent, tool_map={...}, trace=trace)

# Agents whose tools are bound internally (a create_react_agent Runnable, say)
# cannot be rebound. wrap_agent raises AgentInterceptionError rather than
# quietly letting the real tools run — build the agent against the mocks instead:
wrapped = wrap_agent(agent=None, tool_map={...}, trace=trace, intercept=False)
agent = create_react_agent(model, wrapped.tools.values())

# CrewAI
from scrutineer.adapters.crewai import wrap_crew_agent
wrapped = wrap_crew_agent(your_crew, tool_map={...}, trace=trace)

# OpenAI SDK
from scrutineer.adapters.openai import wrap_agent
wrapped = wrap_agent(your_agent, tool_map={...}, trace=trace)

# Generic (any framework)
from scrutineer.adapters.generic import HookAdapter
adapter = HookAdapter(mock=your_mock, before=hook_fn)

Chaos Example

from scrutineer.chaos import (
    ToolFailureInjector,
    ContextDegradation,
    CascadingFailures,
    ChaosBudget,
)

# Fail 30% of tool calls with timeout errors
injector = ToolFailureInjector(
    failure_type="timeout",
    probability=0.3,
)

# Degrade context with quadratic acceleration
degradation = ContextDegradation(strategy="TRUNCATION")

# Cascade failures from database to API to UI using a custom dependency graph
cascade = CascadingFailures(
    cascade_probability=0.7,
    max_cascade_depth=3,
    dependency_graph={
        "database": "api_server",
        "api_server": "user_interface",
    },
)

# Cap total failures per run
budget = ChaosBudget(max_failures=10)

Documentation

Testing

# Run all tests
pytest tests/ -v

# Run with coverage
pytest tests/ --cov=scrutineer --cov-report=html

# Run integration tests only
pytest tests/scrutineer/test_integration_langchain.py -v

# Lint
ruff check src/ tests/

License

MIT

Metadata

Release files for scrutineer-agents 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scrutineer-agents 0.3.0
File Size Uploaded
scrutineer_agents-0.3.0.tar.gz 651.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for scrutineer-agents 0.3.0
File Interpreter ABI Platform
scrutineer_agents-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 863.1 kB

Release files / scrutineer_agents-0.3.0.tar.gz

Download URL scrutineer_agents-0.3.0.tar.gz
Size 651.2 kB
Tags Source
SHA-256 checksum
How to use checksums
649abbb6247d398a07da3512b5c9e0ba85c0c2144db514626a5906d888110b6b
BLAKE2b-256 checksum
How to use checksums
cbf7997f604acba1739206e3769c293d72e61fc5c31cea25238fa511a83e5855
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release files / scrutineer_agents-0.3.0-py3-none-any.whl

Download URL scrutineer_agents-0.3.0-py3-none-any.whl
Size 211.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
567a8390d4093ad6dc4d5b000208fd3d4771e01d0e04eca37447caf33c4b91dc
BLAKE2b-256 checksum
How to use checksums
c47086311f2effa7d6f5080b770b2bf010a132148cd2e5b286356637beba294c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page