Skip to main content

HorusTrace

CI License Python

Five-layer policy-as-code security analysis for AI agents.

HorusTrace statically discovers agent configuration and evaluates five connected security layers:

  1. Agent configuration — tools, approvals, guardrails, MCP, code execution and framework-specific controls.
  2. Capability analysis — effective authority, capability budgets, prohibited actions and dangerous combinations.
  3. Identity & permissions — cloud/IAM roles, wildcard permissions, OAuth scopes and credential source.
  4. Data & network reachability — sensitive resources, resource scope, outbound destinations and allowlist violations.
  5. Attack-path analysis — potential risk combinations such as untrusted content → delegated agent → shell, or confidential data → agent → external write.

Status: v0.4.1 on main. Findings are deterministic within supported constructs. HorusTrace does not prove runtime exploitability or complete live cloud authority.

Security model

What is configured?
        ↓
What can the agent do directly or through delegation?
        ↓
What authority does its identity have?
        ↓
What data/resources/destinations are reachable?
        ↓
Which end-to-end attack paths exist?

The scanner is static-first and local-first. It does not import target Python modules and does not launch MCP servers or agents.

Current framework/input coverage

  • Google Agent Development Kit (ADK) Python 2.x — first-class adapter.
  • Google ADK Agent Config YAML (root_agent.yaml and related agent configs).
  • OpenAI Agents SDK Python constructs, including v0.4 handoff normalization.
  • LangGraph StateGraph / MessageGraph normalization, including hardened real-world graph discovery.
  • Common MCP JSON configuration (mcp.json, .mcp.json).
  • Framework-neutral horustrace.manifest.yaml for business/security intent.
  • Terraform (.tf) for an initial GCP/Azure/AWS IAM view.
  • ADK .env credential-source checks without exposing secret values in findings.

Google ADK coverage

The v0.3 release (published as HorusTrace) performs repository-aware analysis of security-relevant ADK composition rather than only matching individual Agent(...) declarations.

Agents and orchestration

  • Agent / LlmAgent.
  • SequentialAgent, ParallelAgent, LoopAgent.
  • sub_agents and transitive delegated authority.
  • AgentTool delegation and include_plugins isolation.
  • RemoteA2aAgent consumption and to_a2a(...) exposure.
  • transfer restrictions and workflow metadata.
  • cross-file Python/config delegation aliases.

Tools and execution

  • plain Python functions used as ADK tools.
  • FunctionTool, LongRunningFunctionTool, AuthenticatedFunctionTool.
  • require_confirmation.
  • ExecuteBashTool and BashToolPolicy.
  • EnvironmentToolset / LocalEnvironment.
  • UnsafeLocalCodeExecutor, BuiltInCodeExecutor, AgentEngineSandboxCodeExecutor, GkeCodeExecutor.
  • ComputerUseToolset.
  • BigQuery/Bigtable/Data Agent toolsets.
  • Google API/Gmail/Calendar/Docs/Sheets/Slides/YouTube toolsets.
  • search/retrieval tools and inferred untrusted external content.
  • OpenAPI/API Hub/Application Integration/REST-style tool surfaces.
  • memory/artifact/MCP-resource tools represented as data capabilities.

MCP

  • McpToolset / MCPToolset.
  • stdio, SSE and Streamable HTTP connection parameters.
  • URL/transport analysis.
  • auth scheme/credential/header-provider detection.
  • recognized authorization headers.
  • tool filters.
  • confirmation requirements.

Safety controls

  • before_agent_callback, after_agent_callback.
  • before_model_callback, after_model_callback.
  • before_tool_callback, after_tool_callback.
  • model/tool error callbacks.
  • security/safety/policy-style plugins attached through App/Runner.
  • tool-level confirmation.

Google identity context

  • google.auth.default(scopes=...).
  • service-account-file use.
  • ADK credentials config objects.
  • Google API toolset additional_scopes.
  • client-secret literals (reported without secret material).
  • .env API-key/service-account-file indicators.
  • Terraform GCP IAM bindings feeding Layer 3.

See docs/google-adk.md for the exact supported surface and limitations.

Quick start

python -m venv .venv
source .venv/bin/activate
pip install horustrace
horustrace scan .

List the built-in rule catalogue without scanning a project:

horustrace rules
horustrace rules --format json --output rules.json

ADK demo

horustrace scan examples/google-adk-vulnerable --fail-on none
horustrace scan examples/google-adk-secure --fail-on none

The vulnerable ADK fixture intentionally exercises all five layers. The secure fixture should return zero findings under the current rule catalogue.

Generate SARIF:

horustrace scan . --format sarif --output horustrace.sarif --fail-on none

Fail CI on high/critical findings:

horustrace scan . --fail-on high

Malformed, unreadable, or structurally invalid policy manifests stop the scan with exit code 1 and an error on stderr. This also applies with --fail-on none; that option only disables failure for security findings. No report is written for a failed scan.

Coverage and strict CI

horustrace scan . --strict --format json --output report.json

All report formats include file counts and coverage diagnostics. JSON exposes a coverage object; SARIF exposes coverage in run properties and diagnostics as tool execution notifications. Console output separates coverage from findings by layer. --strict returns exit code 1 for detected incomplete analysis, regardless of --fail-on. Unlike a fatal manifest error, incomplete analysis still writes the report so CI can retain the diagnostics. Security threshold failures return exit code 2; incomplete analysis takes precedence in strict mode.

Diagnostics use stable ARG-COV-* identifiers and currently cover read/parse failures, unresolved Python tool/MCP references, unresolved delegation, dynamic agent configuration sequences or expanded keyword arguments, unresolved external helper semantics, dynamic MCP endpoints/tool filters, unknown MCP authentication state, and scans with no supported security targets. Files in default ignored directories, and subtrees containing an .horustrace-ignore marker, are excluded from file counts. Other unsupported file types are counted as skipped. A scanned file was read and parsed; that count does not mean its entire application behavior was understood.

No detected coverage gaps is not proof of complete analysis. Runtime-generated behavior, arbitrary function semantics, cloud authorization, and control effectiveness remain outside these diagnostics. A clean report means no supported rules triggered.

Fingerprints, baselines and suppressions

Every finding has an arg-v1: fingerprint based on its rule, agent, repository-relative path, and evidence. Line numbers and checkout roots are excluded, so fingerprints survive routine source movement and different CI workspaces. A material evidence or scope change produces a new fingerprint.

Create an initial baseline of current findings:

horustrace baseline . \
  --output .horustrace.suppressions.yaml \
  --reason "Initial adoption backlog SEC-42" \
  --expires 2026-12-31

Baseline creation scans without applying existing suppressions. It requires a reason and a non-past expiry and refuses to replace an existing file unless --force is passed. It also refuses to write when coverage diagnostics show incomplete analysis. Review the generated entries before committing them; a baseline records temporary risk acceptance rather than making the findings safe.

Suppression files use this schema:

version: 1
suppressions:
  - id: accepted-shell-migration
    reason: Temporary migration path owned by SEC-42
    expires: 2026-12-31
    fingerprint: arg-v1:0123456789abcdef01234567
    rule_id: AGT020

A fingerprint is the narrowest scope. A rule suppression without a fingerprint must also specify agent or a repository-relative path glob. IDs, reasons, and expiry dates are mandatory; unknown fields and duplicate IDs or YAML keys fail closed. IDs use letters, digits, dots, underscores, and dashes; reasons are single-line; path scopes cannot be absolute or escape the scan root. Expired entries never hide findings. Stale, matched, and expired entries remain in console, JSON, and SARIF audit output. --strict also fails when an exception has expired. Use --suppressions path/to/file.yaml to select a non-default file.

Reviewed benchmark

horustrace benchmark benchmarks/cases.yaml

The reviewed corpus declares the exact RULE@agent findings expected for each case. Unexpected findings are measured as false positives and missing findings as false negatives. The v0.3 corpus contains 26 reviewed scenarios across secure, execution, delegation, MCP, identity, data/network, attack-path and dynamic/unresolved analysis. One case intentionally expects incomplete analysis and an exact coverage diagnostic; all other cases fail on incomplete coverage. The command exits nonzero on any drift and supports --format json for CI artifacts. See benchmarks/README.md.

Evidence and control semantics

Findings retain their rule IDs and severity thresholds and now include:

  • assessment: static_configuration, policy_violation, heuristic_risk, or potential_risk.
  • provenance: facts with a subject, origin, and source location. observed means a supported static configuration was discovered; declared means a manifest assertion; inferred means a heuristic or derived relationship.
  • limitations: uncertainty about runtime authority, controls, and exploitability.

Configuration findings include local evidence; aggregate capability/data/path findings include agent context, which is not a formal data-flow trace. A manifest policy violation may involve declared or inferred capabilities; it does not establish runtime authority. Function capabilities are heuristic, while recognized built-in capabilities follow the scanner's supported static semantics.

control_observations in JSON and SARIF distinguish approval configuration, callback hooks, plugin-name inference, sandbox configuration, and possible network destinations. For supported ADK and MCP configuration, they also record static evidence of sandbox timeout/network/filesystem limits, Bash allowlist plus blocklist policies, MCP tool allowlists, and whether every discovered egress destination fits a declared allowlist. Their runtime effectiveness is always not_verified. Callback or plugin presence can satisfy a missing-hook rule, but does not prove that arbitrary callback/plugin code authorizes actions safely. Approval callbacks alone do not establish an approval requirement. Mixed hosted-MCP approval policies are treated as unknown rather than blanket approval.

Attack paths are labeled potential risks with basis: capability_cooccurrence and exploitability: not_verified. Their severity reflects potential impact, not proven exploitability. A literal URL inside a function establishes a possible destination, not an egress allowlist; such functions now trigger the missing-restriction rule. No runtime enforcement is inferred from those literals.

Manifest schema

Manifests support schema version 1. Omitting version retains the legacy v1 behavior. Unknown fields, unsupported versions, duplicate YAML keys, invalid nested types, cyclic YAML aliases, and negative capability limits are rejected. Approval, guardrail, authentication, and restriction fields require actual booleans; capability/scope fields accept strings or lists of strings. Policy errors identify the field and source line/column without printing its value. Existing documented field aliases remain supported. The schema definitions live in manifest_schema.py.

ADK-specific rule highlights

  • ADK001 — privileged ADK agent lacks a detected tool-control callback/plugin/confirmation boundary.
  • ADK002 — unsafe local code executor.
  • ADK003 — LocalEnvironment exposes local shell/file I/O.
  • ADK004 — bash execution without a detected restrictive BashToolPolicy.
  • ADK012 — sandboxed code execution lacks an explicit timeout, network, or filesystem limit.
  • ADK005 — computer-use capability lacks an explicit action boundary.
  • ADK006 — BigQuery write capability is not statically blocked.
  • ADK007 — broad generated/API toolset without a tool filter.
  • ADK008 — delegated AgentTool disables inherited plugins.
  • ADK009 — remote A2A agent card uses plaintext HTTP.
  • ADK010 — remote A2A agent has no detected authentication.
  • ADK011 — privileged agent is exposed over A2A without a detected safety control.

These run in addition to the framework-neutral AGT/CAP/IDN/DATA/NET/PATH rules.

Framework-neutral security manifest

The manifest declares business intent and runtime context that static source parsing cannot prove:

version: 1
agents:
  - name: invoice-agent

    inputs:
      - name: supplier-portal
        kind: web
        trust: untrusted

    data:
      - name: invoices
        classification: confidential
        selector: /finance/invoices/**

    identities:
      - name: invoice-agent-sa
        provider: gcp
        roles: [roles/storage.objectViewer]
        resource_scope: projects/acme/buckets/invoices
        credential_source: workload_identity

    network:
      - target: https://erp.example.com/**
        restricted: true

    policy:
      required: [data.read, external.write, network.external]
      denied_capabilities: [process.execute, destructive.write]
      allowed_resources: [/finance/invoices/**]
      allowed_destinations: [https://erp.example.com/**]
      require_approval_for: [external.write]
      max_privileged_capabilities: 1

This enables least-privilege comparison between required and effective capabilities and lets Layers 4–5 reason about data and network paths.

GitHub Action

- uses: psandhir/horustrace@v0.4.0
  with:
    path: .
    fail-on: high
    strict: "true"
    suppressions: .horustrace.suppressions.yaml

Design principles

  1. Do not execute the target. Static analysis must be safe on untrusted repositories.
  2. Separate observation from policy. Adapters discover facts; policy adds business/security intent.
  3. Analyse effective authority. Direct and delegated tool combinations matter more than isolated calls.
  4. Connect identity, data and egress. Agent risk is an end-to-end property.
  5. Explain the path. Findings include nodes forming the attack chain.
  6. Prefer deterministic CI findings. Semantic/LLM analysis can be additive later.

Scope boundary

The Google ADK adapter is intended to be comprehensive for security-relevant static constructs in current Python ADK 2.x and native Agent Config YAML. It is not a claim that arbitrary third-party tool implementations, dynamically generated Python, runtime cloud authorization, or separate Java/Go/JavaScript/Kotlin ADK SDK syntax is fully analysed. Those require dedicated adapters or runtime/cloud-control-plane enrichment.

HorusTrace is not a runtime firewall, formal taint verifier, malware scanner or proof that a prompt injection is exploitable. It does not execute the application or call cloud control planes during a normal scan.

Project roadmap

See ROADMAP.md for planned live GCP authority resolution, change-aware analysis, reachability enrichment, and framework expansion.

Support

See SUPPORT.md.

Security

See SECURITY.md.

Contributing

See CONTRIBUTING.md.

License

Apache-2.0. See LICENSE.

Repository scanner configuration

Use .horustrace.yaml to tune scanner policy separately from the security intent manifest and temporary suppressions:

version: 1
scanner:
  strict: true
rules:
  ADK007: {severity: high}
  AGT022: {enabled: false}

Use horustrace scan . --config path/to/config.yaml to select a file explicitly. Disabled rules are reported separately from suppressions; severity overrides affect reporting and failure thresholds but not rule metadata or finding fingerprints.

v0.3 release notes

v0.3 (released as HorusTrace) moved the project from primarily file-level ADK parsing toward repository-level security reachability analysis. It adds cross-file tool and helper resolution, conservative factory and collection resolution, static MCP constant resolution, improved ADK execution/control semantics, stronger identity and OAuth linkage, and lower-noise network and capability inference. Coverage gaps remain explicit rather than being treated as safe. See docs/releases/v0.3.0.md for the release summary.

Release files for horustrace 0.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for horustrace 0.4.1
File Size Uploaded
horustrace-0.4.1.tar.gz 134.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for horustrace 0.4.1
File Interpreter ABI Platform
horustrace-0.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 249.7 kB

Release files / horustrace-0.4.1.tar.gz

Download URL horustrace-0.4.1.tar.gz
Size 134.1 kB
Tags Source
SHA-256 checksum
How to use checksums
27454c77fb6e3df5f12ee6ce017eb9834331a1ee5cf66294b24cf3736cc9fa01
BLAKE2b-256 checksum
How to use checksums
887aff365fc089f4abde1e8787d29e8debac2766a196b4fc4870507bb55a6303
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.

Transparency log

Release files / horustrace-0.4.1-py3-none-any.whl

Download URL horustrace-0.4.1-py3-none-any.whl
Size 115.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
340aac5554eda6f2d5ed164936f18a6f56d74a58cc7d88eab0304854e6fd22e4
BLAKE2b-256 checksum
How to use checksums
f559c9ded703972dca22aef75c37801a68cc31eaef5efd62cd5608f7cf552af0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.0

2 release files

0.5.0

2 release files

This release

0.4.1 This release

2 release files

0.4.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page