Skip to main content

Governed Security Hunting (GSH) Framework: open-source threat hunting for agentic AI

Governed Security Hunting (GSH) Framework

CI License Version Paper Website MITRE ATLAS NIST CSF Stars Forks DOI

Author: Sunil Gentyala, IEEE Senior Member | Lead Cybersecurity and AI Security Consultant, HCLTech Contact: sunil.gentyala@ieee.org Website: sunilgentyala.github.io/gsh-framework License: Apache 2.0


Most enterprise security stacks were not built for the threat surface that agentic AI creates. Endpoint agents cannot see what an LLM gateway is doing. SIEMs have no baselines for multi-agent tool call chains. The GSH Framework closes that gap.

GSH is an open-source research artifact for autonomous agentic AI threat hunting. It provides structured detection playbooks, behavioral baselining logic, and a policy-driven enforcement engine (Sovereign Sentinel) designed for the cognitive cyber domain: the operational layer where large language models, autonomous agents, and multi-agent pipelines interact with enterprise infrastructure.

All detection signals are mapped to MITRE ATLAS and NIST CSF 2.0, giving practitioners framework-aligned coverage they can operationalize immediately.


Try It in One Command

cd demo && docker compose up --build --abort-on-container-exit

Watch Hunt-005 catch an MCP rug pull, a poisoned tool description, and an implementation swap that keeps the tool schema identical. Details and limits: demo/README.md.

GSH Hunt-005 demo: rug pull, tool poisoning and implementation swap are all caught

Replay of the real output of python demo/run_demo.py, rendered from captured terminal output (not a live screen recording). Very long lines are truncated and the local temp path is masked.


Contributors Wanted

GSH is actively looking for security engineers, Python developers, detection engineers and AI-security researchers.

Good first contributions (each has a scoped issue with files, acceptance criteria and a difficulty estimate):

Look for good first issue and help wanted. First contribution? Documentation and tests are welcome too, and small fixes can go straight to a PR. See CONTRIBUTING.md.

Especially wanted: people who will try to break the Hunt-005 detection assumptions (false positives, false negatives, bypasses of the Implementation Identity Gate).


Current Status

The hunt playbooks, detection logic, thresholds, and policy schema are complete and documented.

Hunt-001 through Hunt-004 (scripts/gsh-sentinel-deploy.py, scripts/gsh-probe-eval.py) implement the full baselining, drift-scoring, and ZTLV enforcement logic end-to-end, but ship with a synthetic telemetry generator (clearly marked SIMULATION MODE in the script output and # Replace this block in source) so you can see the detection logic run without a live environment first. Wiring --target to a real LLM gateway event stream is the integration step you complete before using this for actual enforcement.

Hunt-005 (adapters/mcp_proxy.py, scripts/gsh-mcp-proxy.py, scripts/gsh-baseline.py) is different: it is a real MCP JSON-RPC stdio proxy that intercepts actual tool definitions and tool calls between a real MCP host and a real MCP server - approval-time schema hashing, drift detection, semantic poisoning scans, and per-call enforcement (permit/alert/block) all run against live traffic, not synthetic data. A captured baseline is never auto-trusted: it starts as UNVERIFIED and only becomes a trusted comparison point through the gsh-baseline.py capture -> review -> approve -> verify workflow; --mode aggressive refuses to even launch the wrapped server without an approved baseline. An approved baseline is also bound to an Implementation Identity (resolved executable/script hashes plus an adjacent dependency-lock file) so a server that keeps an identical, approved tool schema while its underlying implementation is swapped out under the same launch command is refused before it is even started, in any enforcement mode - not just when the baseline itself drifts. See tests/test_mcp_proxy.py and tests/test_gsh_baseline.py for subprocess-driven end-to-end tests of both CLIs. Known gaps: canary/response-asymmetry comparison and tool-return-value scanning are not implemented yet (see playbooks/hunt-005-mcp-tool-poisoning.md section 5.2 for details), and only the stdio transport is supported (not streamable HTTP/SSE MCP servers).

SIEM output (adapters/splunk_hec.py, adapters/elastic_bulk.py, adapters/windows_eventlog.py) is also real: set siem_output: splunk, siem_output: elastic, or siem_output: windows_eventlog in your policy YAML (see configs/sentinel-policy-default.yaml) and both gsh-sentinel-deploy.py and gsh-mcp-proxy.py will send findings there (Splunk HEC / Elasticsearch _bulk over real HTTP, or a registered source in the local Windows Application Event Log). A failed or unconfigured send always falls back to local file output - a finding is never silently dropped. The Windows Event Log adapter is Windows-only and requires pywin32; on any other platform (or without pywin32) it logs a warning and falls back like any other unconfigured destination. See tests/test_siem_adapters.py and tests/test_windows_eventlog.py (the latter includes a test that writes a real event and reads it back, not just a mocked one).

LangChain telemetry (adapters/langchain_callback.py) is a fourth real integration: GSHCallbackHandler attaches to any LangChain Runnable/agent via config={"callbacks": [handler]} and evaluates real tool-call rate, token velocity, unauthorized-tool invocations, and suspicious call parameters against Hunt-001/Hunt-004 thresholds - no synthetic data. Important limitation: LangChain callback handlers are notification hooks, not gates - by default LangChain swallows exceptions raised inside a callback rather than stopping the tool call, so this adapter can only alert, never block. Every finding it emits is explicitly marked enforcement_mode: "alert_only" and action_taken: "ALERTED", regardless of policy mode. It also has no visibility into DNS queries (Hunt-002). See tests/test_langchain_callback.py, tested against langchain-core 1.4.x.

See open issues for remaining work (SARIF reporting and the Hunt-006 playbook). The Docker demo shipped as demo/.


Version History

Full release notes (including known limitations at each release) are on the Releases page. Summary:

Version Highlights
v1.9.0 Sentinel CLI error handling: clearer messages, --log-level DEBUG tracebacks and a remediation hint (#12, #16). Canonical Apache-2.0 license text so GitHub detects it, README banner, Zenodo DOI in CITATION.cff. Dependabot updates, CodeQL scanning and main branch protection enabled
v1.8.0 One-command Docker demo for Hunt-005 (demo/): drives the real baseline and proxy CLIs through a rug pull, a poisoned tool description and an implementation swap that keeps the schema identical, with a CI smoke test. Contribution path relaxed (small fixes go straight to a PR), Contributors Wanted block, scoped good-first-issue guides. No framework code changes
v1.7.0 Implementation Identity Gate for Hunt-005 (adapters/mcp_proxy.py): an approved baseline is now bound to the resolved executable/script hashes and adjacent dependency-lock file behind the launch command, not just the tool schema - a server that keeps an identical schema while its implementation is swapped out under the same command reference is blocked before it is ever launched, not just quarantined after the fact. Closes a gap identified in independent third-party review (schema-only trust). Breaking: pre-1.7.0 approved baselines have no identity data and must be re-captured and re-approved
v1.6.0 Security audit pass: newly-added MCP tools (added after an approved baseline) are no longer auto-authorized for invocation until reviewed; concurrent-write protection for MCP proxy stdout and LangChain adapter alert IDs; DNS-tunneling allowlist (Hunt-002) now raises its entropy/label-length bar for trusted apex domains instead of fully exempting them
v1.5.0 Real Windows Application Event Log output adapter (adapters/windows_eventlog.py); optional and Windows-only, safe no-op elsewhere
v1.4.0 Real LangChain callback adapter (adapters/langchain_callback.py) for Hunt-001/Hunt-004 telemetry - alert-only by design, since LangChain callbacks cannot block a tool call
v1.3.0 Real Splunk HEC and Elastic bulk SIEM output adapters, wired into both the Sentinel and the MCP proxy via a shared dispatcher; a failed/unconfigured SIEM send now always falls back to local file output
v1.2.0 Real MCP JSON-RPC stdio proxy for Hunt-005 (adapters/mcp_proxy.py) - schema-hash drift detection, semantic poisoning scan, and real per-call enforcement against live MCP traffic, not simulated
v1.1.0 Hunt-004 (rogue agent) completed; Hunt-005 (MCP supply chain / tool poisoning) added as a playbook; project website launched
v1.0.0-beta Initial public release: Hunt-001 through Hunt-003 playbooks, Sentinel reference scripts (synthetic telemetry), default policy schema

Framework Components

Component Description
Sovereign Sentinel Policy-driven behavioral enforcement agent deployed alongside LLM gateways
Hunt Playbooks Structured threat detection playbooks for high-severity agentic AI threats
DDI-AI Fusion DNS/DHCP/IPAM telemetry layer with AI-agent-aware baselining
Zero-Trust Logic Validation (ZTLV) Gate Per-invocation tool call authorization engine
Behavioral Baseline Engine Continuous model output drift detection and probe evaluation pipeline

Hunt Playbooks

Playbook Threat Class Severity Status
Hunt-001 Agentic Loop / Resource Exhaustion High Active
Hunt-002 DDI Covert Channel / C2 via DNS Critical Active
Hunt-003 ML Model Poisoning / Behavioral Drift Critical Active
Hunt-004 Rogue Agent / Unauthorized Tool Use Critical Active
Hunt-005 MCP Supply Chain / Tool Poisoning Critical Active

Quick Start

1. Clone the Repository

git clone https://github.com/sunilgentyala/gsh-framework.git
cd gsh-framework
pip install -r requirements.txt

Alternative: install as a package. The repo is also pip-installable from source (not yet published to PyPI, so install from the checkout, not pip install gsh-framework):

pip install .              # adapters + all CLIs, PyYAML only (no SIEM/LangChain/Windows extras)
pip install ".[splunk]"    # + Splunk/Elastic HTTP output (adapters/splunk_hec.py, elastic_bulk.py)
pip install ".[langchain]" # + LangChain callback adapter
pip install ".[windows]"   # + Windows Event Log adapter (Windows only)
pip install ".[llm]"       # + OpenAI-compatible client for gsh-probe-eval.py
pip install ".[dev]"       # + pytest, ruff, mypy, black

This installs gsh-sentinel-deploy, gsh-mcp-proxy, gsh-baseline, gsh-probe-eval, and gsh-ddi-log-parser as commands (equivalent to python scripts/<name>.py), and makes adapters importable without manually adjusting sys.path. Extras can be combined, e.g. pip install ".[splunk,langchain]".

2. Review the Sentinel Policy

cat configs/sentinel-policy-default.yaml

Edit it to set your organization name, SIEM output destination, and egress allowlist before deploying.

3. Deploy a Sovereign Sentinel

Start in passive mode to build a 7-day behavioral baseline, then move to standard enforcement:

python scripts/gsh-sentinel-deploy.py \
  --target "llm-gateway-01" \
  --mode passive \
  --policy configs/sentinel-policy-default.yaml \
  --baseline-window 7d

As shipped, this generates synthetic telemetry (SIMULATION MODE, logged at startup) so you can watch the baselining and scoring logic run immediately. Replace the telemetry-generation block noted in the script (real LLM gateway/API metrics or LangChain callbacks) to run it against live traffic.

4. Run the MCP Proxy (Hunt-005 - real enforcement, not simulated)

Unlike step 3, this runs against real MCP traffic. A captured baseline is never auto-trusted - capture it, review it, then approve it:

python scripts/gsh-baseline.py capture \
  --server-id "corp-tools-mcp-01" \
  --server-cmd "npx -y @modelcontextprotocol/server-filesystem /srv/data"

python scripts/gsh-baseline.py review --baseline baselines/mcp/corp-tools-mcp-01.json

python scripts/gsh-baseline.py approve \
  --baseline baselines/mcp/corp-tools-mcp-01.json --reviewer "your-name-or-email"

Then configure your MCP host to launch the proxy instead of the real server directly:

python scripts/gsh-mcp-proxy.py \
  --server-cmd "npx -y @modelcontextprotocol/server-filesystem /srv/data" \
  --server-id "corp-tools-mcp-01" \
  --mode standard \
  --baseline baselines/mcp/corp-tools-mcp-01.json

The proxy will alert on (or, in --mode aggressive, block) definition drift, poisoned tool descriptions, invisible Unicode content, and unauthorized tool calls. In --mode aggressive, the proxy refuses to launch the wrapped server at all unless the baseline above has been approved - see playbooks/hunt-005-mcp-tool-poisoning.md section 5.1 for why, and section 5.2 for the full detection logic.

5. Wire a LangChain Agent to a Sentinel (real telemetry, alert-only)

pip install langchain-core
from adapters.langchain_callback import GSHCallbackHandler

handler = GSHCallbackHandler(
    target="my-langchain-agent",
    allowlist=["web_search", "calculator"],   # unlisted tools trigger an immediate alert
)

# Attach to any LLM, tool, or chain via the standard LangChain callbacks config:
llm.invoke(prompt, config={"callbacks": [handler]})
my_tool.invoke(args, config={"callbacks": [handler]})

handler.flush()  # evaluate any partial window at the end of a run

This is alert-only, not enforcement - see adapters/langchain_callback.py's module docstring for why LangChain callback handlers cannot reliably block a tool call.

6. Run a Hunt Playbook

Each playbook is a self-contained Markdown document with detection logic, data sources, MITRE ATLAS mapping, triage decision tree, and response actions. Start with Hunt-001 for loop detection:

cat playbooks/hunt-001-agentic-loop-detection.md

Research

A companion research paper covering the full technical rationale, design decisions, and threat model is in preparation and not yet submitted. Per publication policy, the manuscript is not included in this repository. For research inquiries, contact sunil.gentyala@ieee.org.


Repository Structure

gsh-framework/
├── README.md
├── SECURITY.md
├── LICENSE
├── CITATION.cff
├── CONTRIBUTING.md
├── requirements.txt
├── pyproject.toml                  # pip-installable package + console-script CLIs, extras
├── .github/
│   └── workflows/
│       └── ci.yml                  # pytest + ruff + mypy across Python 3.10-3.13
├── adapters/
│   ├── mcp_proxy.py                 # Real MCP JSON-RPC proxy (Hunt-005) + baseline approval governance
│   ├── langchain_callback.py        # Real LangChain telemetry, alert-only (Hunt-001/004)
│   ├── splunk_hec.py                # Real Splunk HTTP Event Collector output
│   ├── elastic_bulk.py              # Real Elasticsearch/OpenSearch _bulk output
│   ├── windows_eventlog.py          # Real Windows Application Event Log output
│   └── siem_dispatch.py             # Shared dispatcher used by all three SIEM adapters
├── configs/
│   └── sentinel-policy-default.yaml
├── docs/
│   └── index.html                  # Project website (GitHub Pages)
├── playbooks/
│   ├── hunt-001-agentic-loop-detection.md
│   ├── hunt-002-ddi-tunneling-anomaly.md
│   ├── hunt-003-model-poisoning-baseline.md
│   ├── hunt-004-rogue-agent-detection.md
│   └── hunt-005-mcp-tool-poisoning.md
├── probes/
│   └── standardized-probe-set-v1.json
├── scripts/
│   ├── _cli_shims.py               # console-script entry points (pyproject.toml) for the scripts below
│   ├── ddi-log-parser-ai.py
│   ├── gsh-baseline.py             # capture/review/approve/verify CLI for MCP baseline governance
│   ├── gsh-mcp-proxy.py            # CLI for adapters/mcp_proxy.py
│   ├── gsh-probe-eval.py
│   └── gsh-sentinel-deploy.py
├── tests/
│   ├── test_ddi_log_parser.py
│   ├── test_gsh_baseline.py
│   ├── test_mcp_proxy.py
│   ├── test_siem_adapters.py
│   ├── test_windows_eventlog.py
│   ├── test_langchain_callback.py
│   └── fixtures/
│       ├── mock_mcp_server.py      # Minimal MCP stdio server for testing
│       └── mock_http_sink.py       # Minimal HTTP server for testing SIEM adapters
├── baselines/
└── reports/

Threat Coverage

Threat MITRE ATLAS MITRE ATT&CK NIST CSF 2.0
Agentic Loop / Resource Exhaustion AML.T0048, AML.T0040 DE.AE-02, DE.CM-01, RS.MI-01
DDI Covert Channel Exfiltration AML.T0048, AML.T0051 T1071.004, T1048, T1568 DE.CM-01, DE.AE-04, PR.DS-01
ML Model Poisoning / Behavioral Drift AML.T0020, AML.T0043, AML.T0044 ID.RA-01, DE.AE-02, DE.CM-06
Rogue Agent / Unauthorized Tool Use AML.T0051, AML.T0053, AML.T0054 PR.PS-04, DE.CM-01, RS.AN-03
MCP Supply Chain / Tool Poisoning AML.T0010, AML.T0051, AML.T0053 T1195 ID.SC-04, PR.PS-04, DE.CM-06

Contributing

Security practitioners, AI safety researchers, and detection engineers are welcome. Read CONTRIBUTING.md before opening a Pull Request.

High-priority contributions include: additional hunt playbooks, refined detection thresholds, and integration adapters for LangChain, AutoGen, CrewAI, and MCP host platforms.


Citation

If you use the GSH Framework in your research, please cite:

@misc{gentyala2026gsh,
  author       = {Gentyala, Sunil},
  title        = {The Governed Security Hunting (GSH): An Autonomous Agentic Framework
                  for Defending the Cognitive Cyber Domain},
  year         = {2026},
  howpublished = {Open Source Research Artifact, GitHub},
  publisher    = {Zenodo},
  doi          = {10.5281/zenodo.21384588},
  url          = {https://github.com/sunilgentyala/gsh-framework}
}

The DOI above is the Zenodo concept DOI, which always resolves to the latest archived release. Machine-readable metadata is in CITATION.cff.


Security Vulnerabilities

To report a vulnerability in the GSH Framework itself, use GitHub's private vulnerability reporting or email sunil.gentyala@ieee.org with the subject [GSH Security Vulnerability] - [brief description]. Do not open a public GitHub Issue. See SECURITY.md for the full policy, supported versions, and response timeline.


  • ContextGuard: Zero-trust middleware for Model Context Protocol (MCP) server security. Precision 100%, Recall 96.7%, F1 98.3% at 1.005ms latency.
  • ARGUS: LLM application security scanner.
  • IEEE Senior Member Profile: ORCID 0009-0005-2642-3479

Metadata

Release files for gsh-framework 1.9.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gsh-framework 1.9.0
File Size Uploaded
gsh_framework-1.9.0.tar.gz 93.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gsh-framework 1.9.0
File Interpreter ABI Platform
gsh_framework-1.9.0-py3-none-any.whl Python 3 none any Details

Total release size: 172.5 kB

Release files / gsh_framework-1.9.0.tar.gz

Download URL gsh_framework-1.9.0.tar.gz
Size 93.1 kB
Tags Source
SHA-256 checksum
How to use checksums
a50ed2378b2b39da0525f4a2be0e71e0d5bc1746dddd1ab79bef0dbf44dce7fb
BLAKE2b-256 checksum
How to use checksums
698ec3b9a3ec76be7e2e9ef62afd98c704a546678daf70ee337bc7098131e360
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / gsh_framework-1.9.0-py3-none-any.whl

Download URL gsh_framework-1.9.0-py3-none-any.whl
Size 79.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f32688f2bc72915170804bd637f9fb4dfbf40aa0f368679461536b9c4e2eeb98
BLAKE2b-256 checksum
How to use checksums
44667f7ae1fad5583a847b08fca59bb2d0893b5882acf01e4c76836b85703edf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.9.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page