Skout Scan
Find what your AI agent evals aren't testing.
Skout Scan statically scans your agent code and eval suite to surface high-confidence behaviors that may be insufficiently tested.
Local-first. No code execution. No source upload. No API key required.
Try Skout Scan
1. Install
Skout Scan requires Python 3.12 or newer.
pip install skout-scan
2. Scan your agent repository
cd your-agent-repo
skout scan .
Skout Scan discovers repository artifacts, extracts eval scenarios and deterministic agent behaviors, matches behaviors against existing eval evidence, and surfaces selected high-confidence potential gaps.
3. Review the findings
skout review
Run skout review from the same repository you scanned; review state is stored
locally in that repository. From another directory, use
skout review --repository path/to/your-agent-repo.
The full local workflow is:
cd your-agent-repo
skout scan .
skout review
skout metrics --repository .
The interactive review shows the source, evidence, coverage assessment, and a suggested eval for each high-confidence potential gap. Skout Scan lets you classify each finding as:
add_eval— you intend to add or modify an evalvalid_later— the gap is valid but is not a current priorityalready_covered— adequate coverage already existsnot_relevant— the behavior does not need an evalsuppressed— hide the finding from normal output
4. See your validation metrics
skout metrics --repository .
skout metrics --repository . --json
The JSON form provides a machine-readable export that you can share with the Skout Scan team during V0 validation.
5. Give feedback
We're validating Skout Scan with engineers building real AI agents. If you try it, we'd really value 2 minutes of feedback:
Generate the JSON to paste into the feedback form with:
skout metrics --repository . --json
Skout Scan does not automatically upload metrics, source code, or repository contents.
What Skout Scan finds
A finding is an explainable potential gap, for example:
Potential eval gap
Tool:
create_jira_ticket
Why flagged:
Skout Scan found the tool behavior but no eval that both exercises the behavior
and verifies its expected outcome.
Suggested test:
Invoke create_jira_ticket with valid inputs and verify the expected mutation.
Each detected behavior receives one of these assessments:
covered— an eval meaningfully exercises the behavior and verifies its expected outcome or invariantpartially_covered— an eval exercises the functionality, but does not fully verify a relevant condition, branch, failure path, or expected outcomepotentially_uncovered— no discovered eval appears to materially exercise and verify the behavior
Skout Scan selects high-confidence actionable findings from these assessments. It does not turn every raw assessment into a finding or calculate an overall behavioral coverage percentage.
Supported patterns
V0 extracts eval evidence from:
- pytest
test_*functions and test methods inTest*classes - direct
assertstatements andpytest.raises - statically resolvable calls and literal arguments
- JSONL eval scenarios with an
input, optionalexpected, and optional metadata such as a name or tool list
V0 extracts deterministic behaviors from:
- functions decorated with
@tool, qualified@*.tool, or@function_tool - local functions registered through literal tool collections,
tools=[...], orbind_tools([...]) - explicit raised exceptions and explicit failure returns
- deterministic conditional branches with visible return, raise, escalation, or handoff outcomes
- supported LangGraph
add_edgetransitions and literaladd_conditional_edgesroutes on statically assignedStateGraphinstances
Calls through .invoke, .ainvoke, and .coroutine are recognized as tool
invocations only when the receiver is already known statically as a tool. The
wrapper call alone does not prove that an expected outcome or failure path was
verified.
V0 also supports common statically analyzable CrewAI patterns:
- direct
Agent(...),Task(...), andCrew(...)construction @CrewBaseclasses with@agent,@task,@crew,@before_kickoff, and@after_kickoffmethods- statically linked CrewAI YAML and plain JSON agent/task configuration
- CrewAI
@toolfunctions,BaseToolsubclasses, and literal agent/task tool attachments - literal sequential task order, task context dependencies, and hierarchical crew metadata without inferred runtime delegation paths
- explicit Flow
@start,@listen, and@routerrelationships and literal router targets
CrewAI role, goal, backstory, task description, and expected-output text are retained as structured evidence. Skout does not interpret that prose as a behavioral guarantee.
How it works
Repository
↓
Discover artifacts
↓
Extract eval scenarios
↓
Extract deterministic agent behaviors
↓
Match behaviors to eval evidence
↓
Surface high-confidence potential gaps
↓
Engineer reviews findings
↓
Persist local lifecycle and feedback
Analysis and matching are static and deterministic. Skout Scan never imports or
executes Python from the target repository. It stores finding history, review
feedback, and scan history in .agentguard/agentguard.db within the scanned
repository.
Behavior and finding IDs are designed to survive formatting, whitespace, and line movement where possible. On later complete scans, Skout Scan can observe when a new or modified eval covers a previous finding and record that finding as resolved. Review choices and lifecycle history persist locally across scans.
Limitations
- Static analysis is intentionally conservative. Indirect calls, fixture indirection, aliases, and wrappers outside the supported patterns may be missed.
- Dynamically constructed tool, agent, or workflow registration may not be discovered.
- Dynamic CrewAI task/agent lists, runtime orchestration, computed config paths, and non-literal Flow routes may not be recoverable. CrewAI JSONC configuration is detected but is not parsed in V0.
- Dynamic pytest parametrization has limited support.
- Prompt files are discovered, but natural-language prompt obligations are not extracted as behaviors in V0.
- File moves, symbol renames, and major restructuring may change stable IDs.
- V0 does not use semantic, embedding, or LLM-based matching.
- V0 does not analyze production traces.
- Skout Scan does not calculate an overall numeric behavioral coverage percentage.
- Findings are potential testing gaps. They do not certify that an agent is safe, unsafe, production-ready, or inadequately tested.
Configuration
Configuration is optional. When present, agentguard.toml must be in the root
of the repository being scanned.
include = ["**/*.py", "**/*.jsonl"]
exclude = ["**/.venv/**", "**/.agentguard/**", "**/build/**", "**/dist/**"]
include selects candidate artifacts and exclude removes matching paths;
exclusions take precedence. Patterns match POSIX-style paths relative to the
repository root and support recursive ** segments.
Without a configuration file, Skout Scan includes Python, text, Markdown, and JSONL files and excludes common Git, virtual-environment, dependency, cache, build, distribution, and local-state directories. Symbolic links are not followed.
Development
Create a Python 3.12 environment and install the project with development tools:
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
The public distribution and command are skout-scan and skout. The current
internal Python package remains agentguard under src/agentguard/.
Run the contributor checks:
pytest
ruff check .
ruff format --check .
mypy src/agentguard
Validate release artifacts with:
python -m build
python -m twine check dist/*
Release maintainers should follow the release guide.
Metadata
Release files for skout-scan 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| skout_scan-0.1.2.tar.gz | 53.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| skout_scan-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 115.9 kB
Release files / skout_scan-0.1.2.tar.gz
| Download URL | skout_scan-0.1.2.tar.gz |
|---|---|
| Size | 53.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e96b442a96702b43ddff0236d4c538b21fa5ee1c5e624256e29296b4030f934f
|
|
BLAKE2b-256 checksum How to use checksums |
cf3fc78b7d77b4e766c69e02dc63c82e3d8dc1c7ea2a7cbcf4b157907fd32e1e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / skout_scan-0.1.2-py3-none-any.whl
| Download URL | skout_scan-0.1.2-py3-none-any.whl |
|---|---|
| Size | 62.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
cda657bc76ef174d051890d83494f70a96c4140dcc1a062d9dc772c75b0980b0
|
|
BLAKE2b-256 checksum How to use checksums |
15245d91d8584e5df07fb7dff05d38e1548ab374b93cb7edc9437c5e5e0b860f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log