Skip to main content

Skout Scan

Find what your AI agent evals aren't testing.

Skout Scan statically scans your agent code and eval suite to surface high-confidence behaviors that may be insufficiently tested.

Local-first. No code execution. No source upload. No API key required.

Try Skout Scan

1. Install

Skout Scan requires Python 3.12 or newer.

pip install skout-scan

2. Check repository compatibility

Run the read-only diagnostic before your first scan:

cd your-agent-repo
skout doctor

skout doctor reports detected frameworks and evals, unsupported or dynamic patterns, effective scope and exclusions, and whether the supported repository surface is ready or partially assessable. It does not measure agent quality or create .agentguard/ state. From another directory, use skout doctor --repository path/to/your-agent-repo.

Use skout doctor --json for machine-readable diagnostics.

3. Scan your agent repository

cd your-agent-repo
skout scan .

Skout Scan discovers repository artifacts, extracts eval scenarios and deterministic agent behaviors, matches behaviors against existing eval evidence, and surfaces selected high-confidence potential gaps.

If generated output or another directory should not be scanned, exclude it for that run:

skout scan . --exclude "output/**"

4. Review the findings

skout review

Run skout review from the same repository you scanned; review state is stored locally in that repository. From another directory, use skout review --repository path/to/your-agent-repo.

The full local workflow is:

cd your-agent-repo
skout doctor
skout scan .
skout review
skout metrics --repository .

The interactive review shows the source, evidence, coverage assessment, and a suggested eval for each high-confidence potential gap. Skout Scan lets you classify each finding as:

  • add_eval — you intend to add or modify an eval
  • valid_later — the gap is valid but is not a current priority
  • already_covered — adequate coverage already exists
  • not_relevant — the behavior does not need an eval
  • suppressed — hide the finding from normal output

5. See your validation metrics

skout metrics --repository .
skout metrics --repository . --json

The JSON form provides a machine-readable export that you can share with the Skout Scan team during V0 validation.

6. Give feedback

We're validating Skout Scan with engineers building real AI agents. If you try it, we'd really value 2 minutes of feedback:

Share feedback

Generate the JSON to paste into the feedback form with:

skout metrics --repository . --json

Skout Scan does not automatically upload metrics, source code, or repository contents.

What Skout Scan finds

A finding is an explainable potential gap, for example:

Potential eval gap

Tool:
create_jira_ticket

Why flagged:
Skout Scan found the tool behavior but no eval that both exercises the behavior
and verifies its expected outcome.

Suggested test:
Invoke create_jira_ticket with valid inputs and verify the expected mutation.

Each detected behavior receives one of these assessments:

  • covered — an eval meaningfully exercises the behavior and verifies its expected outcome or invariant
  • partially_covered — an eval exercises the functionality, but does not fully verify a relevant condition, branch, failure path, or expected outcome
  • potentially_uncovered — no discovered eval appears to materially exercise and verify the behavior

Skout Scan selects high-confidence actionable findings from these assessments. It does not turn every raw assessment into a finding or calculate an overall behavioral coverage percentage.

Supported patterns

V0 extracts eval evidence from:

  • pytest test_* functions and test methods in Test* classes
  • direct assert statements and pytest.raises
  • statically resolvable calls and literal arguments
  • JSONL eval scenarios with an input, optional expected, and optional metadata such as a name or tool list

JSONL files are first classified with a bounded deterministic schema probe. Files under eval-oriented paths or whose sampled object records consistently contain input are parsed as eval datasets. Application JSONL used for memory, caches, logs, or other data is ignored for eval coverage when its sampled schema is clearly not an eval schema.

V0 extracts deterministic behaviors from:

  • functions decorated with @tool, qualified @*.tool, or @function_tool
  • local functions registered through literal tool collections, tools=[...], or bind_tools([...])
  • explicit raised exceptions and explicit failure returns
  • deterministic conditional branches with visible return, raise, escalation, or handoff outcomes
  • supported LangGraph add_edge transitions and literal add_conditional_edges routes on statically assigned StateGraph instances

Calls through .invoke, .ainvoke, and .coroutine are recognized as tool invocations only when the receiver is already known statically as a tool. The wrapper call alone does not prove that an expected outcome or failure path was verified.

V0 also supports common statically analyzable CrewAI patterns:

  • direct Agent(...), Task(...), and Crew(...) construction
  • @CrewBase classes with @agent, @task, @crew, @before_kickoff, and @after_kickoff methods
  • statically linked CrewAI YAML and plain JSON agent/task configuration
  • CrewAI @tool functions, BaseTool subclasses, and literal agent/task tool attachments
  • literal sequential task order, task context dependencies, and hierarchical crew metadata without inferred runtime delegation paths
  • explicit Flow @start, @listen, and @router relationships and literal router targets

CrewAI role, goal, backstory, task description, and expected-output text are retained as structured evidence. Skout does not interpret that prose as a behavioral guarantee.

Skout reports provenance counts for statically linked CrewAI YAML and JSON configuration. Literal Markdown or text instruction files referenced by those configs may be retained as linked evidence, but their natural-language contents are not converted into behavioral obligations.

Skout supports common statically analyzable Pydantic AI patterns:

  • Agent(...), @agent.tool, @agent.tool_plain, and static tools=[...]
  • static Tool(...) registrations
  • RunContext dependency declarations and visible ctx.deps access evidence
  • literal or referenced instructions and structured output_type contracts
  • output validators with explicit conditional ModelRetry paths
  • known-agent run() and run_sync() calls in pytest scenarios
  • Python Pydantic Evals Case and Dataset definitions, including inputs, expected_output, metadata, and statically visible evaluator evidence

Instructions and output schemas are retained as evidence; their prose and field descriptions are not interpreted as behavioral obligations. Pydantic Evals cases become coverage candidates only when Skout can statically associate their dataset task with a known agent.

How it works

Repository
↓
Discover artifacts
↓
Extract eval scenarios
↓
Extract deterministic agent behaviors
↓
Match behaviors to eval evidence
↓
Surface high-confidence potential gaps
↓
Engineer reviews findings
↓
Persist local lifecycle and feedback

Analysis and matching are static and deterministic. Skout Scan never imports or executes Python from the target repository. It stores finding history, review feedback, and scan history in .agentguard/agentguard.db within the scanned repository.

Behavior and finding IDs are designed to survive formatting, whitespace, and line movement where possible. On later scans with sufficient evidence for a finding, Skout Scan can observe when a new or modified eval covers it and record that finding as resolved. Review choices and lifecycle history persist locally across scans.

Limitations

  • Static analysis is intentionally conservative. Indirect calls, fixture indirection, aliases, and wrappers outside the supported patterns may be missed.
  • Dynamically constructed tool, agent, or workflow registration may not be discovered.
  • Dynamic CrewAI task/agent lists, runtime orchestration, computed config paths, and non-literal Flow routes may not be recoverable. CrewAI JSONC configuration is detected but is not parsed in V0.
  • Dynamic Pydantic AI toolsets, MCP/runtime tool discovery, runtime-generated instructions, dynamic output contracts, serialized Pydantic Evals datasets, and general Pydantic Graph workflows are not analyzed. Custom evaluator semantics are not treated as verification unless deterministic evidence is visible. Complex aliases, factories, and cross-language call graphs are also outside the supported static patterns.
  • Skout can analyze supported Python agent/eval code inside a larger mixed-language repository. Coverage results apply only to the supported Python surface; surrounding languages are not analyzed for behavioral coverage.
  • Dynamic pytest parametrization has limited support.
  • Prompt files are discovered, but natural-language prompt obligations are not extracted as behaviors in V0.
  • Ambiguous JSONL produces one classification warning and is skipped. Malformed records in a recognized eval dataset remain incomplete eval evidence.
  • Unsupported or dynamic evidence can make affected assessments unavailable. Unrelated application data and skipped cache directories do not invalidate otherwise assessable behaviors.
  • File moves, symbol renames, and major restructuring may change stable IDs.
  • V0 does not use semantic, embedding, or LLM-based matching.
  • V0 does not analyze production traces.
  • Skout Scan does not calculate an overall numeric behavioral coverage percentage.
  • Findings are potential testing gaps. They do not certify that an agent is safe, unsafe, production-ready, or inadequately tested.

Repository scope and configuration

Configuration is optional. When present, agentguard.toml must be in the root of the repository being scanned. Use it for exclusions that should apply to every scan:

include = ["**/*.py", "**/*.jsonl"]
exclude = [
  "output/**",
  "generated/**",
]

include selects candidate artifacts and exclude removes matching paths; exclusions take precedence. Patterns match POSIX-style paths relative to the repository root and support recursive ** segments.

Temporary CLI exclusions

Use --exclude for a one-time scope change without editing repository configuration:

skout scan . --exclude "output/**"

skout scan . \
  --exclude "output/**" \
  --exclude "generated/**"

skout doctor --repository . --exclude "output/**"

--exclude is repeatable. Patterns are repository-relative and support recursive ** segments. CLI exclusions are additive to built-in defaults and the exclusions in agentguard.toml; any matching exclusion wins. Matching directories are pruned before traversal, so excluded contents are not parsed or counted and do not produce extraction warnings.

Use --exclude for temporary experiments and agentguard.toml for persistent repository scope. Do not exclude directories containing agent code, evals, requirements, guardrails, prompts, or other evidence that Skout should analyze.

Without a configuration file, Skout Scan includes Python, text, Markdown, and JSONL files and excludes common Git, virtual-environment, dependency, cache, tool-runtime, build, distribution, and local-state directories, including .uv-cache and .tools. knowledge/, config/, skills/, and output/ are not excluded by default. Symbolic links are not followed.

Avoid excluding broad source or test trees merely to reduce findings. Scope out paths only when their contents are generated, vendored, cached, or irrelevant to the agent behavior and eval suite you intend to assess.

Troubleshooting

  • A scan finds 0 behaviors: run skout doctor and review detected frameworks, extraction warnings, scope, and readiness. The repository may use dynamic construction or an unsupported framework.
  • Behaviors are found but no evals are found: confirm tests use supported pytest shapes or the documented JSONL eval schema and are inside the active include/exclude scope.
  • Behaviors and evals are found but candidate pairs are zero: tests may invoke the behavior indirectly, use unsupported wrappers, or omit statically visible symbol evidence.
  • Pydantic AI is detected but extraction is limited: agent, tool, toolset, instruction, output, or evaluator construction may be dynamic or outside the supported patterns. Review the doctor warnings.
  • Pydantic AI behaviors are found but no evals are found: add supported pytest evidence or Python Pydantic Evals cases within the configured scope.
  • Pydantic Evals cases are found but candidate pairs are zero: make the dataset task's call to a known Pydantic AI agent statically visible.
  • A mixed-language repository is detected: the readiness result applies to its supported Python agent/eval surface; surrounding source languages are not part of the behavioral assessment.
  • Generated, output, cache, or vendor files dominate the scan: run skout doctor --exclude "path/**" to verify a temporary scope, then add the useful pattern to agentguard.toml.
  • The result is partially assessable or inconclusive: review the reported warnings before interpreting findings. These statuses mean Skout lacks enough deterministic evidence for part or all of the repository.
  • A scan reports 0 findings: this does not establish complete coverage. Check skout doctor readiness and the scan's behavior, eval, and candidate pair counts before interpreting the result.

Development

Create a Python 3.12 environment and install the project with development tools:

python3.12 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

The public distribution and command are skout-scan and skout. The current internal Python package remains agentguard under src/agentguard/.

Run the contributor checks:

pytest
ruff check .
ruff format --check .
mypy src/agentguard

Validate release artifacts with:

python -m build
python -m twine check dist/*

Release maintainers should follow the release guide.

Metadata

Release files for skout-scan 0.1.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for skout-scan 0.1.6
File Size Uploaded
skout_scan-0.1.6.tar.gz 70.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for skout-scan 0.1.6
File Interpreter ABI Platform
skout_scan-0.1.6-py3-none-any.whl Python 3 none any Details

Total release size: 151.7 kB

Release files / skout_scan-0.1.6.tar.gz

Download URL skout_scan-0.1.6.tar.gz
Size 70.9 kB
Tags Source
SHA-256 checksum
How to use checksums
b55aa28613990e6dbd3d983d961aa55617ad3e8e6d8f38770f127ae88b58d155
BLAKE2b-256 checksum
How to use checksums
59cb13016deaa11f1dde256263dcab2738b1cc7e8c1c175d9af2622150e6a39f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / skout_scan-0.1.6-py3-none-any.whl

Download URL skout_scan-0.1.6-py3-none-any.whl
Size 80.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6cfe6e4d61b569ae25e25ea3237a133fb6c33a6a0424a7c5e4e077db3680fa58
BLAKE2b-256 checksum
How to use checksums
129f75580428ee7f21a2d775f24f1629e0b4d0b3f1432262f50f420a1db31813
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.6 This release

2 release files

0.1.5

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page