Skout Scan
Find what your AI agent evals aren't testing.
Skout Scan statically scans your agent code and eval suite to surface high-confidence behaviors that may be insufficiently tested.
Local-first. No code execution. No source upload. No API key required.
Try Skout Scan
1. Install
Skout Scan requires Python 3.12 or newer.
pip install skout-scan
2. Check repository compatibility
Run the read-only diagnostic before your first scan:
cd your-agent-repo
skout doctor
skout doctor explains what Skout can discover, which framework and eval
patterns it recognizes, whether the repository is ready for a meaningful
coverage assessment, and any scope changes worth making first. It does not
create .agentguard/ state. From another directory, use
skout doctor --repository path/to/your-agent-repo.
Use skout doctor --json for machine-readable diagnostics.
3. Scan your agent repository
cd your-agent-repo
skout scan .
Skout Scan discovers repository artifacts, extracts eval scenarios and deterministic agent behaviors, matches behaviors against existing eval evidence, and surfaces selected high-confidence potential gaps.
4. Review the findings
skout review
Run skout review from the same repository you scanned; review state is stored
locally in that repository. From another directory, use
skout review --repository path/to/your-agent-repo.
The full local workflow is:
cd your-agent-repo
skout scan .
skout review
skout metrics --repository .
The interactive review shows the source, evidence, coverage assessment, and a suggested eval for each high-confidence potential gap. Skout Scan lets you classify each finding as:
add_eval— you intend to add or modify an evalvalid_later— the gap is valid but is not a current priorityalready_covered— adequate coverage already existsnot_relevant— the behavior does not need an evalsuppressed— hide the finding from normal output
5. See your validation metrics
skout metrics --repository .
skout metrics --repository . --json
The JSON form provides a machine-readable export that you can share with the Skout Scan team during V0 validation.
6. Give feedback
We're validating Skout Scan with engineers building real AI agents. If you try it, we'd really value 2 minutes of feedback:
Generate the JSON to paste into the feedback form with:
skout metrics --repository . --json
Skout Scan does not automatically upload metrics, source code, or repository contents.
What Skout Scan finds
A finding is an explainable potential gap, for example:
Potential eval gap
Tool:
create_jira_ticket
Why flagged:
Skout Scan found the tool behavior but no eval that both exercises the behavior
and verifies its expected outcome.
Suggested test:
Invoke create_jira_ticket with valid inputs and verify the expected mutation.
Each detected behavior receives one of these assessments:
covered— an eval meaningfully exercises the behavior and verifies its expected outcome or invariantpartially_covered— an eval exercises the functionality, but does not fully verify a relevant condition, branch, failure path, or expected outcomepotentially_uncovered— no discovered eval appears to materially exercise and verify the behavior
Skout Scan selects high-confidence actionable findings from these assessments. It does not turn every raw assessment into a finding or calculate an overall behavioral coverage percentage.
Supported patterns
V0 extracts eval evidence from:
- pytest
test_*functions and test methods inTest*classes - direct
assertstatements andpytest.raises - statically resolvable calls and literal arguments
- JSONL eval scenarios with an
input, optionalexpected, and optional metadata such as a name or tool list
JSONL files are first classified with a bounded deterministic schema probe.
Files under eval-oriented paths or whose sampled object records consistently
contain input are parsed as eval datasets. Application JSONL used for memory,
caches, logs, or other data is ignored for eval coverage when its sampled schema
is clearly not an eval schema.
V0 extracts deterministic behaviors from:
- functions decorated with
@tool, qualified@*.tool, or@function_tool - local functions registered through literal tool collections,
tools=[...], orbind_tools([...]) - explicit raised exceptions and explicit failure returns
- deterministic conditional branches with visible return, raise, escalation, or handoff outcomes
- supported LangGraph
add_edgetransitions and literaladd_conditional_edgesroutes on statically assignedStateGraphinstances
Calls through .invoke, .ainvoke, and .coroutine are recognized as tool
invocations only when the receiver is already known statically as a tool. The
wrapper call alone does not prove that an expected outcome or failure path was
verified.
V0 also supports common statically analyzable CrewAI patterns:
- direct
Agent(...),Task(...), andCrew(...)construction @CrewBaseclasses with@agent,@task,@crew,@before_kickoff, and@after_kickoffmethods- statically linked CrewAI YAML and plain JSON agent/task configuration
- CrewAI
@toolfunctions,BaseToolsubclasses, and literal agent/task tool attachments - literal sequential task order, task context dependencies, and hierarchical crew metadata without inferred runtime delegation paths
- explicit Flow
@start,@listen, and@routerrelationships and literal router targets
CrewAI role, goal, backstory, task description, and expected-output text are retained as structured evidence. Skout does not interpret that prose as a behavioral guarantee.
Skout reports provenance counts for statically linked CrewAI YAML and JSON configuration. Literal Markdown or text instruction files referenced by those configs may be retained as linked evidence, but their natural-language contents are not converted into behavioral obligations.
How it works
Repository
↓
Discover artifacts
↓
Extract eval scenarios
↓
Extract deterministic agent behaviors
↓
Match behaviors to eval evidence
↓
Surface high-confidence potential gaps
↓
Engineer reviews findings
↓
Persist local lifecycle and feedback
Analysis and matching are static and deterministic. Skout Scan never imports or
executes Python from the target repository. It stores finding history, review
feedback, and scan history in .agentguard/agentguard.db within the scanned
repository.
Behavior and finding IDs are designed to survive formatting, whitespace, and line movement where possible. On later scans with sufficient evidence for a finding, Skout Scan can observe when a new or modified eval covers it and record that finding as resolved. Review choices and lifecycle history persist locally across scans.
Limitations
- Static analysis is intentionally conservative. Indirect calls, fixture indirection, aliases, and wrappers outside the supported patterns may be missed.
- Dynamically constructed tool, agent, or workflow registration may not be discovered.
- Dynamic CrewAI task/agent lists, runtime orchestration, computed config paths, and non-literal Flow routes may not be recoverable. CrewAI JSONC configuration is detected but is not parsed in V0.
- Dynamic pytest parametrization has limited support.
- Prompt files are discovered, but natural-language prompt obligations are not extracted as behaviors in V0.
- Ambiguous JSONL produces one classification warning and is skipped. Malformed records in a recognized eval dataset remain incomplete eval evidence.
- Unsupported or dynamic evidence can make affected assessments unavailable. Unrelated application data and skipped cache directories do not invalidate otherwise assessable behaviors.
- File moves, symbol renames, and major restructuring may change stable IDs.
- V0 does not use semantic, embedding, or LLM-based matching.
- V0 does not analyze production traces.
- Skout Scan does not calculate an overall numeric behavioral coverage percentage.
- Findings are potential testing gaps. They do not certify that an agent is safe, unsafe, production-ready, or inadequately tested.
Repository scope and configuration
Configuration is optional. When present, agentguard.toml must be in the root
of the repository being scanned.
include = ["**/*.py", "**/*.jsonl"]
exclude = ["output/**", "generated/**", "vendor/**"]
include selects candidate artifacts and exclude removes matching paths;
exclusions take precedence. Patterns match POSIX-style paths relative to the
repository root and support recursive ** segments.
Use repeatable --exclude options for temporary scope changes without editing
the repository configuration:
skout doctor --exclude "output/**" --exclude "vendor/**"
skout scan . --exclude "output/**" --exclude "vendor/**"
Exclusions are additive. Skout applies its built-in exclusions first, then
those in agentguard.toml, then every CLI --exclude value. Include patterns
still select candidate artifact types, and any matching exclusion wins. Skout
checks directory exclusions before descending into them, so excluding a large
generated or vendor tree also avoids traversal work.
Without a configuration file, Skout Scan includes Python, text, Markdown, and
JSONL files and excludes common Git, virtual-environment, dependency, cache,
tool-runtime, build, distribution, and local-state directories, including
.uv-cache and .tools. knowledge/, config/, skills/, and output/ are
not excluded by default. Symbolic links are not followed.
Avoid excluding broad source or test trees merely to reduce findings. Scope out paths only when their contents are generated, vendored, cached, or irrelevant to the agent behavior and eval suite you intend to assess.
Troubleshooting
- Doctor reports no supported behaviors: Skout found files but did not recognize deterministic tool or workflow patterns. Check the detected framework and extraction warnings; the repository may use dynamic construction or an unsupported framework.
- Behaviors are found but no evals are found: confirm tests use supported pytest shapes or the documented JSONL eval schema and are inside the active include/exclude scope.
- Behaviors and evals are found but candidate pairs are zero: tests may invoke the behavior indirectly, use unsupported wrappers, or omit statically visible symbol evidence.
- Generated, output, cache, or vendor files dominate the scan: run
skout doctor --exclude "path/**"to verify a temporary scope, then add the useful pattern toagentguard.toml. - The result is partially assessable or inconclusive: review the reported warnings before interpreting findings. These statuses mean Skout lacks enough deterministic evidence for part or all of the repository.
- A scan reports 0 findings: this does not establish complete coverage.
Check
skout doctorreadiness and the scan's behavior, eval, and candidate pair counts before interpreting the result.
Development
Create a Python 3.12 environment and install the project with development tools:
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
The public distribution and command are skout-scan and skout. The current
internal Python package remains agentguard under src/agentguard/.
Run the contributor checks:
pytest
ruff check .
ruff format --check .
mypy src/agentguard
Validate release artifacts with:
python -m build
python -m twine check dist/*
Release maintainers should follow the release guide.
Metadata
Release files for skout-scan 0.1.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| skout_scan-0.1.5.tar.gz | 63.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| skout_scan-0.1.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 135.1 kB
Release files / skout_scan-0.1.5.tar.gz
| Download URL | skout_scan-0.1.5.tar.gz |
|---|---|
| Size | 63.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
fa9bc6df1ee4cab0b3cf7dc6a442a3f5b79aede0dbd9cda261e26798a15a5020
|
|
BLAKE2b-256 checksum How to use checksums |
f102458f7130c0eee792473c34a8a89306121adf425032ad706593d0c39d3cb0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency logRelease files / skout_scan-0.1.5-py3-none-any.whl
| Download URL | skout_scan-0.1.5-py3-none-any.whl |
|---|---|
| Size | 72.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ca81a69a61a138196d0c36232285b4fe0e8f0ff4970229c3133cfd0eec37b86e
|
|
BLAKE2b-256 checksum How to use checksums |
7ef9fab9ecd5ec3f10bd7eab918280c0d678b6060a4e46d18b587212e7c21c4c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency log