Skip to main content

pkgguard

Vet the packages and repos your AI assistant recommended — before you install them.

CI PyPI Python License: MIT No API key Ecosystems


Your AI coding assistant just gave you a pip install line or a list of "recommended libraries." Some of those names don't exist. Some are one keystroke away from a popular package. Some are real but AGPL-licensed, abandoned, riddled with CVEs, or propped up by bought stars. Installing them is how modern supply-chain attacks start.

pkgguard checks every name in one command and tells you which are safe — across PyPI, npm, crates.io, Go, RubyGems, Packagist, NuGet and pub.dev — with no API key and without executing a single line of the packages it inspects.

$ pkgguard requests reqeusts beautifulsoup-4 django==3.0.0 super-fake-pkg-zzz

      ✅ OK      requests           Exists on PyPI · 1.6B downloads/mo · Apache-2.0
      ⚠️ WARN    django             31 known CVEs in 3.0.0 — upgrade
      ❌ DANGER  reqeusts           Resembles 'requests' but does not exist — slopsquat bait
      ❌ DANGER  beautifulsoup-4    Hallucination — real package is 'beautifulsoup4'
      ❌ DANGER  super-fake-pkg-zzz Not found on PyPI or npm — likely invented

1 ok  1 warn  3 danger  0 unknown          # exit code 2 → fails your CI

Existing tools (Snyk, Socket, OSV) scan dependencies you've already chosen. pkgguard answers the newer, earlier question: "the thing the LLM just told me to install — is it even real, and should I trust it?"


🩸 Why this matters: the slopsquatting epidemic

LLMs invent package names. A peer-reviewed USENIX Security 2025 study generated 2.23M code samples and found:

  • 19.7% of AI-generated samples referenced a package that does not exist (up to ~33% for some models).
  • When the same prompt was repeated, 43% of hallucinated names appeared every single time — they are predictable.

That predictability is the attack. An adversary asks an LLM what it hallucinates, pre-registers those exact names on PyPI/npm with malware inside, and waits for the next developer to copy-paste the assistant's answer into a terminal. Security researchers named it slopsquatting, and it is already happening in the wild.

As AI assistants become the default way developers discover dependencies, the moment an AI hands you a package list is now a front-line security boundary. pkgguard guards exactly that moment.


⚡ What it checks

Every package/repo runs through an ordered pipeline of 11 checks. The worst finding sets the verdict, so a single command gives you one clear answer per item.

Check What it catches
Existence Names that don't exist on any registry — the core hallucination / slopsquat signal
Typosquat + homoglyph 1–2 edits from a popular package (reqeustsrequests) and digit/letter look-alikes (dj4ng0django)
Known vulnerabilities Open CVEs/advisories for the resolved version via OSV.dev (GHSA / PyPA / RustSec / RubySec…)
Source malware scan (opt-in --scan) Statically inspects the package archive for install-time code execution, obfuscated payloads, child_process/os.system, credential access — without ever running it
Agent skills Malicious SKILL.md instructions — credential exfiltration, stealth directives, agent-config writes, encoded payloads (--skills)
License traps AGPL / SSPL / BUSL / CC-BY-NC / fair-code / "no license" — and custom / NOASSERTION non-commercial licenses (reads the raw LICENSE text; common for AI model repos). Detects restrictions in 25 languages, not just English — see below
Maintenance Archived, disabled, deprecated, or long-abandoned projects
Popularity Download counts as a legitimacy signal
Fake stars (opt-in --deep) Star-count inflation — with reputable-owner / real-adoption suppression so normal viral growth isn't flagged
Purpose / policy (opt-in --policy) Tools whose purpose is ToS-abuse/attack (account farms, bypass, DDoS) — a heuristic lens, backs off for defensive/detection tooling
Repo health GitHub stars / last commit / license / archived state for the upstream repo
Malware metadata npm install scripts, freshly-registered look-alikes, packages with no auditable source

Supported ecosystems: PyPI · npm · crates.io · Go modules · RubyGems · Packagist · NuGet · pub.dev — plus GitHub repos.


🛡️ Agent-skill security (--skills)

A skill is a folder with a SKILL.md full of instructions your agent follows — with your permissions, on your machine, next to your credentials. Installing one off a leaderboard is closer to running a stranger's shell script than to adding a dependency.

That surface is now the problem: an independent audit of 2,857 published skills found roughly 12% malicious, and Trail of Bits bypassed the existing malicious-skill detectors using prompt injection and bytecode-hidden payloads.

pkgguard --skills owner/repo

Four rules, each requiring its signals to co-occur inside a short window:

Rule Fires on
credential-exfil reads a secret and sends it off-machine, in one instruction
stealth-instruction "do not tell the user", "skip the confirmation prompt"
agent-config-write writes into .claude/, CLAUDE.md, mcp.json, .cursor/
encoded-payload a base64 blob plus a decode step, inside an instruction file

Why proximity matters. The first version paired signals anywhere in the same document and produced five findings against anthropics/skills and obra/superpowersall false. An API-reference skill names ANTHROPIC_API_KEY in one paragraph and a URL in another; "silently" shows up in ordinary prose. Synthetic tests passed; only real repos exposed it. Rewritten to require proximity and imperative phrasing: 0 findings across four real skill repos, all six malicious shapes still caught.

This is a heuristic, not an audit. It never emits ❌ on its own — it quotes the line and leaves the judgment to you. Read the SKILL.md before installing; that reading is the entire security model.

🌍 License checks work in 25 languages

Open source is global; LICENSE files are not always in English. GitHub's classifier returns NOASSERTION for a hand-written license in Korean, Spanish or Russian — and an English-only scanner then reads that as "unclassifiable" and passes it as ✅. That is the worst possible outcome for a tool whose whole job is answering "can I use this commercially?".

A real example: a Korean LICENSE that explicitly bans resale and paid-service bundling graded ✅ OK before v0.2.0.

⚠️  WARN  lbiz-partners/hometax-doum
    License trap: custom / restrictive (non-commercial signals)
    GitHub couldn't classify the license (SPDX=NOASSERTION), but the LICENSE
    (written in Korean) contains restrictive terms — likely non-commercial /
    research-only, unsafe for a commercial product without a separate license.
    Evidence: "...유지한 비상업적 공유 (링크 공유 권장) ## 금지 - 본 스킬 모음 또는 그 수정본의 재판매,..."

Supported: Arabic, Chinese, Czech, Dutch, English, French, German, Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Korean, Persian, Polish, Portuguese, Romanian, Russian, Spanish, Swedish, Thai, Turkish, Ukrainian, Vietnamese.

How it works. Word order is the real problem, not vocabulary. Korean lists the prohibited acts first and negates once at the end; German splits the verb; Arabic and Hebrew read right-to-left. So instead of matching phrases, pkgguard looks for two concepts co-occurring in a short window — a commercial-use term and a prohibition term, in either order. Adding a language means adding two short word lists, not encoding its grammar.

It also catches gated commercial use ("commercial use requires a separate license"), which is legally a gate rather than a ban but answers "can I ship this for free today?" the same way.

On false positives. Calling MIT non-commercial would be worse than missing a restriction — it trains you to ignore ⚠️. Every language ships a paired test asserting its permissive phrasing is not flagged ("상업적 이용을 허용합니다", "Die kommerzielle Nutzung ist gestattet", …), and the suite is validated against the real LICENSE files of React, Kubernetes, Rust, Django, PyTorch and others: 0 false positives.

Findings are written in English. The detection is multilingual; the report is not yet localised. If you'd like localised output, open an issue — it is a deliberate scope call, not an oversight.

📦 Install

pip install pkgguard-cli            # core — zero dependencies, no API key
pip install "pkgguard-cli[rich]"    # + prettier coloured tables
pip install "pkgguard-cli[mcp]"     # + MCP server for AI assistants

The PyPI distribution is pkgguard-cli; the command it installs is pkgguard. Requires Python 3.9+. The core engine is stdlib-only.

🚀 Quick start

# a few names
pkgguard requests numpy pandas

# a manifest — auto-detected: requirements.txt, package.json, Cargo.toml, Gemfile, go.mod
pkgguard requirements.txt
pkgguard package.json Cargo.toml          # several at once

# 🌟 the headline trick: paste whatever ChatGPT / Claude told you
pkgguard --stdin < chat.txt
pbpaste | pkgguard --stdin                # macOS

# go deeper
pkgguard requirements.txt --scan          # download + statically scan source
pkgguard some/repo --deep                 # add fake-star analysis
pkgguard some/repo --policy               # flag abuse/attack-purpose tools

pkgguard mines free text for pip install … / npm i … commands, GitHub links, inline `code spans` and bullet lists — and is careful not to flag plain English prose as packages. Names whose ecosystem isn't stated are checked against both PyPI and npm.

Machine-readable output & CI

pkgguard -f requirements.txt --json               # JSON to stdout
pkgguard -f requirements.txt --markdown -o report.md
pkgguard -f requirements.txt --fail-on warn       # non-zero exit gates your pipeline

Exit codes: 0 clean · 1 a warning (--fail-on warn) · 2 a danger.


🔌 Integrations

pre-commit hook
# .pre-commit-config.yaml
- repo: https://github.com/Highcrypto7/pkgguard
  rev: v0.1.0
  hooks:
    - id: pkgguard          # auto-runs on requirements*.txt and package.json
GitHub Action
# .github/workflows/pkgguard.yml
- uses: Highcrypto7/pkgguard@v1
  with:
    files: "requirements.txt package.json"
    fail-on: danger

🤖 Use it through your AI assistant (MCP)

The highest-leverage way to run pkgguard is to not run it yourself. Register the MCP server and the assistant vets a package before it ever recommends it — slopsquatting stops at the source instead of at your terminal.

pip install "pkgguard-cli[mcp]"

Then add it to your assistant's MCP config:

Claude Codeclaude mcp add pkgguard -- pkgguard-mcp

Claude Desktop (claude_desktop_config.json) / Cursor (.cursor/mcp.json):

{
  "mcpServers": {
    "pkgguard": {
      "command": "pkgguard-mcp"
    }
  }
}

Two tools are exposed:

Tool Use
vet_packages(text) Pass names, a requirements.txt / package.json body, or free text — including the assistant's own draft answer. Returns a verdict per item with reasons.
is_safe_to_install(name) Fast yes/no gate for one package or owner/repo.

A useful standing instruction for your agent:

Before recommending or installing any package, call pkgguard. Do not suggest anything that comes back danger, and surface the reason for anything warn.

It is language-neutral on input. Paste an assistant's answer in any language — the parser reads structure (install commands, owner/repo slugs, code spans, bullet lists), not English prose, so nothing is lost in Korean, Japanese, Spanish, Russian or Arabic, and prose words don't leak in as fake package names:

> 웹 스크래핑에는 requests를 쓰세요:
>   pip install requests
> 3D 좌석 시야는 github.com/thebuggeddev/football-stadium 참고하세요.

⚠️ WARN  thebuggeddev/football-stadium  License trap: custom / restrictive
   … the LICENSE contains restrictive terms — unsafe for a commercial product
✅ OK    requests                        Exists on PyPI (latest 2.34.2)

🥊 How it compares

Other tools are excellent at scanning dependencies you've already chosen. pkgguard is the fast first gate at the moment an AI (or a teammate) hands you a list.

pkgguard sloppy-joe depscope GuardDog Snyk / Socket
Hallucination / existence ~
Typosquat + homoglyph ~ ~
Known CVEs (OSV)
Static source malware scan
License traps (AGPL / NC / …) ~
Maintenance / dead repo ~ ~
Fake-star inflation
Paste a chat answer (free text)
MCP self-check for assistants
Ecosystems 8 2 19 5 many
Open source · No key · Offline-degraded partial

pkgguard's edge: the widest set of checks in a single zero-key OSS gate, framed around AI output — including license and fake-star checks the others skip, and a "paste the chat answer" workflow nobody else has. Honest gap: GuardDog/Snyk/Socket do deeper source-level malware analysis; run them alongside pkgguard for defence in depth.


📊 Proof

  • Benchmark: 100% accuracy on a labeled set of 30 PyPI/npm packages (15 real, 15 hallucinated/typosquat). Reproduce: python benchmark/run_benchmark.py. See BENCHMARK.md.
  • Zero false positives when vetting the 50 most popular real PyPI/npm packages.
  • 67 automated tests, deterministic and offline.

🧭 Verdicts

  • OK — exists and nothing concerning found.
  • ⚠️ WARN — usable, but read the caveat (license, CVE, staleness, look-alike…).
  • DANGER — doesn't exist, or a strong risk signal. Don't install without verifying.
  • UNKNOWN — couldn't determine (offline / rate-limited). Honest about what it didn't check.

🔐 Design principles

  • No API key, ever. Public registry/GitHub metadata over HTTPS. Set GITHUB_TOKEN only to raise rate limits.
  • No code execution. The source scan parses with ast and pattern-matching; it never imports or runs package code, and extracts archives in-memory with strict size/path guards.
  • Honest by default. "Couldn't check" is ❔, never a silent ✅.
  • Fast & offline-friendly. On-disk response cache; a previous run answers even with no network.

⚠️ Limitations

  • Heuristics, not proof. A ✅ means "no red flags found," not a security guarantee.
  • The typosquat reference list is a curated set of popular packages, not all of every registry.
  • Fake-star and source-scan checks are opt-in and intentionally conservative — they complement, not replace, dedicated tools (StarScout, GuardDog).
  • Unauthenticated GitHub is limited to ~60 requests/hour; set GITHUB_TOKEN for large runs.

🗺️ Roadmap

  • More ecosystems (Maven, Hex, CPAN)
  • Large-scale benchmark against the trendmicro/slopsquatting dataset
  • VS Code extension
  • Deeper static source analysis

🤝 Contributing

Issues and PRs welcome — a new ecosystem is just a registry adapter, and a new check is a single module (see src/pkgguard/checks/). Run pytest before submitting.

📄 License

MIT — see LICENSE. Built to make the AI coding era a little safer.

Release files for pkgguard-cli 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pkgguard-cli 0.3.0
File Size Uploaded
pkgguard_cli-0.3.0.tar.gz 93.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pkgguard-cli 0.3.0
File Interpreter ABI Platform
pkgguard_cli-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 184.5 kB

Release files / pkgguard_cli-0.3.0.tar.gz

Download URL pkgguard_cli-0.3.0.tar.gz
Size 93.4 kB
Tags Source
SHA-256 checksum
How to use checksums
9a533bc43ba6f9707867c7dbc6c387a3e3693047434664812412faf5165309af
BLAKE2b-256 checksum
How to use checksums
d95e44cf48f5ab21559f2aa36b316978e95bddc2a391dfaa36b3a7f5150b6530
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 29, 2026.

Transparency log

Release files / pkgguard_cli-0.3.0-py3-none-any.whl

Download URL pkgguard_cli-0.3.0-py3-none-any.whl
Size 91.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
021ab5820e3b45d72d872a20e77c637902d0f9592ef7ccfa75471954b4120a8b
BLAKE2b-256 checksum
How to use checksums
57911a4a609ddf5ea141e5e8d06d4e989cef3e1b658040e35ab7df6d79958e09
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 29, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page