Scrutineer
Agent Behavioral Testing Platform — Tests what agents DO, not just what they SAY.
The Problem
88% of AI agents fail in production. The dominant failure modes are operational:
- Tool errors (28%)
- Memory/state issues (22%)
- Edge cases (18%)
Yet the entire evaluation ecosystem (DeepEval, LangSmith, MS AGT) focuses on output quality or observability. Nobody tests agent behavior in production-like environments before deployment.
The Solution
Scrutineer fills that gap. It's a behavioral testing platform that:
- Mocks your agent's environment — tools, APIs, databases with configurable latency, errors, and rate limits
- Injects chaos — tool failures, context degradation, cascading errors, spec drift under pressure
- Asserts behavior — 20+ assertions across tool calls, state consistency, governance compliance, resilience, and performance
- Reports regressions — structural diffing, baseline comparison, HTML + JUnit reports
Quick Start
Distribution name: this project publishes to PyPI as
scrutineer-agents. Do not runpip install scrutineer— that name belongs to an unrelated project and installs a different library. The import package and the console script are bothscrutineer.
# Install (zero dependencies by default)
pip install scrutineer-agents
# Or with framework adapters
pip install "scrutineer-agents[adapters]"
# Or with the WebUI dashboard
pip install "scrutineer-agents[web]"
# Run the quickstart example
python examples/langchain_quickstart.py
# Run a YAML scenario
scrutineer run --path examples/basic_scenario.yaml
WebUI Dashboard
Scrutineer includes a browser-based dashboard for running scenarios, viewing traces, and comparing baselines — all wrapping the core Python API.
# Install with web dependencies
pip install "scrutineer-agents[web]"
# Start the dashboard
scrutineer serve
# Or with a custom port
scrutineer serve --port 9090
Then open http://localhost:8080 in your browser.
Features:
- Dashboard — pass/fail stats, recent runs, quick actions
- Scenarios — browse, inspect, and run test scenarios
- Runs — live execution with step-by-step trace visualization
- Baselines — saved results with regression diff comparison
- Live Console — real-time log streaming via SSE during test runs
See src/scrutineer/web/README.md for the full WebUI guide.
Architecture
src/scrutineer/
├── env.py # MockTool, MockAPI, MockDatabase, EnvironmentBuilder
├── chaos.py # ToolFailureInjector, ContextDegradation, CascadingFailures
├── assertions.py # 20+ behavioral assertions
├── runner.py # @scrutineer_test decorator, ScenarioRunner
├── reporting.py # Regression reports, JUnit XML, HTML
├── baseline.py # JSON baseline storage with git integration
├── otel.py # OpenTelemetry span model
├── cli.py # Full CLI: run, list, info, baseline, diff, report
├── adapters/ # LangChain, CrewAI, OpenAI SDK, Generic
└── web/ # FastAPI WebUI dashboard (optional)
├── app.py # FastAPI application factory
├── server.py # Uvicorn entry point
├── api/ # REST API routers (scenarios, runs, baselines)
├── services/ # Service layer wrapping core modules
├── schemas/ # Pydantic request/response models
└── static/ # Frontend (HTML, CSS, JS)
The Chaos Module (Differentiator)
Scrutineer's chaos injection is what sets it apart:
- ContextDegradation — Quadratic acceleration curve matching real context window pressure (last 20% is much worse than first 20%)
- CascadingFailures — Multi-agent error propagation with dependency graphs (database → api_server → ui)
- SpecDrift — Agent improvisation under pressure with intensity levels and cumulative drift scoring
No other tool tests these production failure modes.
CLI Commands
scrutineer run <scenario> # Run a test scenario
scrutineer list # List available scenarios
scrutineer info <scenario> # Show scenario details
scrutineer baseline record # Record current state as baseline
scrutineer baseline show # Show recorded baseline
scrutineer diff # Compare current vs baseline
scrutineer report # Generate regression report
scrutineer trace <run-id> # Show execution trace
scrutineer serve # Start WebUI dashboard
scrutineer serve --port 9090 # Custom port
Framework Adapters
Scrutineer ships adapters for specific frameworks, plus a generic hook adapter for anything else:
# LangChain — rebinds your_agent.tools so the agent's own call path hits the mocks
from scrutineer.adapters.langchain import wrap_agent
wrapped = wrap_agent(your_agent, tool_map={...}, trace=trace)
# Agents whose tools are bound internally (a create_react_agent Runnable, say)
# cannot be rebound. wrap_agent raises AgentInterceptionError rather than
# quietly letting the real tools run — build the agent against the mocks instead:
wrapped = wrap_agent(agent=None, tool_map={...}, trace=trace, intercept=False)
agent = create_react_agent(model, wrapped.tools.values())
# CrewAI
from scrutineer.adapters.crewai import wrap_crew_agent
wrapped = wrap_crew_agent(your_crew, tool_map={...}, trace=trace)
# OpenAI SDK
from scrutineer.adapters.openai import wrap_agent
wrapped = wrap_agent(your_agent, tool_map={...}, trace=trace)
# Generic (any framework)
from scrutineer.adapters.generic import HookAdapter
adapter = HookAdapter(mock=your_mock, before=hook_fn)
Chaos Example
from scrutineer.chaos import (
ToolFailureInjector,
ContextDegradation,
CascadingFailures,
ChaosBudget,
)
# Fail 30% of tool calls with timeout errors
injector = ToolFailureInjector(
failure_type="timeout",
probability=0.3,
)
# Degrade context with quadratic acceleration
degradation = ContextDegradation(strategy="TRUNCATION")
# Cascade failures from database to API to UI using a custom dependency graph
cascade = CascadingFailures(
cascade_probability=0.7,
max_cascade_depth=3,
dependency_graph={
"database": "api_server",
"api_server": "user_interface",
},
)
# Cap total failures per run
budget = ChaosBudget(max_failures=10)
Documentation
- Quickstart — 5-minute guide from install to first test
- Chaos Guide — Deep dive on failure injection patterns
- Adapters Guide — How to write custom adapters
- Integration Testing — Proof of value with real LangChain tools
- API Reference — Module documentation
- WebUI Design — Architecture and implementation plan
- WebUI Guide — Getting started with the dashboard
Testing
# Run all tests
pytest tests/ -v
# Run with coverage
pytest tests/ --cov=scrutineer --cov-report=html
# Run integration tests only
pytest tests/scrutineer/test_integration_langchain.py -v
# Lint
ruff check src/ tests/
License
MIT
Metadata
Release files for scrutineer-agents 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| scrutineer_agents-0.3.0.tar.gz | 651.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| scrutineer_agents-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 863.1 kB
Release files / scrutineer_agents-0.3.0.tar.gz
| Download URL | scrutineer_agents-0.3.0.tar.gz |
|---|---|
| Size | 651.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
649abbb6247d398a07da3512b5c9e0ba85c0c2144db514626a5906d888110b6b
|
|
BLAKE2b-256 checksum How to use checksums |
cbf7997f604acba1739206e3769c293d72e61fc5c31cea25238fa511a83e5855
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.15
|
Release files / scrutineer_agents-0.3.0-py3-none-any.whl
| Download URL | scrutineer_agents-0.3.0-py3-none-any.whl |
|---|---|
| Size | 211.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
567a8390d4093ad6dc4d5b000208fd3d4771e01d0e04eca37447caf33c4b91dc
|
|
BLAKE2b-256 checksum How to use checksums |
c47086311f2effa7d6f5080b770b2bf010a132148cd2e5b286356637beba294c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.15
|