Skip to main content

package-doctor

CI PyPI Python License: MIT

Finds the dependencies that sit at a trust boundary and have no one left to fix them. Then stops your coding agent from adding another.

In late August 2026, Anthropic's coordinated disclosure programme reported 2,300 vulnerabilities across 392 open source projects. 421 had been patched upstream.

That ratio is the problem this tool measures. Discovery is becoming automated. Remediation still needs a human to reproduce the issue, avoid regressions, cut a release and support everyone downstream. Finding a thousand bugs does not produce a thousand safe patches.

So the question worth asking about a dependency is not "is it healthy?" It is:

If a vulnerability lands in this package tomorrow, am I exposed, and is anyone home to fix it?

Two ways to ask it. package-doctor scan answers for everything in your lockfile. package-doctor check answers for one package at the moment it is about to be installed — and as a Claude Code hook it asks before an agent's pip install runs, blocking an invented name, a package registered last week, or an abandoned library at a trust boundary, with the reason put in front of the model so it picks something else.

The two-axis model

Every other scanner in this space scores packages on one axis — how maintained they look — and then has to explain why it flagged six. package-doctor uses two, and only escalates when both fire.

Axis 1 — Exposure. Does this package routinely handle data an attacker can influence? Authentication, session and token handling, file uploads, archive extraction, deserialization, HTML parsing, query building, cryptography, URL parsing, and model loading. This comes from a curated map in exposure.toml, not from a heuristic.

It holds 731 packages across 15 categories, curated against the 3,000 most-downloaded packages on PyPI — 86 of the top 100 have been reviewed one way or the other, and Django, FastAPI/ML, document-processing and web-scraping stacks all scan with zero unclassified packages. Anything outside it falls back to classifier inference, which is marked inferred and flagged with ? in output so you can distrust it.

Axis 2 — Remediation capacity. If a fix were needed, would one ship?

Both must fire. six going quiet is not a finding, because six is not at a trust boundary. legacy-auth going quiet is the whole point.

Which advisory first

"Affected by 35 advisories" is not a decision. A list that long gets skimmed and then ignored, which is how real exposure survives a green-looking pipeline. So findings are ranked by how likely the flaw is to actually be used:

  • CISA KEV — the Known Exploited Vulnerabilities catalogue. Membership is not a prediction: it means the flaw has been used against real targets. Nothing else in the tool outranks it.
  • FIRST EPSS — a daily-refreshed probability of exploitation in the next 30 days, giving an ordering where an advisory count gives none.
EXPOSED + NO ONE HOME   act on these
pillow  10.0.0  file/media parsing  CVE-2023-4863 on CISA's known-exploited list,
                                    and your pinned version is affected
Exploitability of your version
  Known exploited (CISA)    CVE-2023-4863
  CVE-2023-4863             >99% chance of exploitation in 30 days
  CVE-2023-50447            1.7% chance of exploitation in 30 days
  CVE-2024-28219            1.0% chance of exploitation in 30 days
    (+12 more scored)
  No EPSS score             1 of 18
    Unscored means unknown, not low risk.

Across sixty real projects that turns 2,383 advisories affecting pinned versions into 32 worth reading first — the ones on CISA's list or above a 10% exploit probability.

Only advisories affecting your pinned version are scored — a CVE fixed five releases ago is not your problem, and including it would drown the signal. pillow 10.4.0 is silent on CVE-2023-4863 for exactly that reason; 10.0.0 is not.

What it doesn't claim

EPSS scores the CVE, not your usage: a bug in a code path you never touch scores the same as one you hit on every request. It ranks, it doesn't decide — which is why a high score alone never creates a verdict, and why the raw number is always shown next to the wording. A CVE with no EPSS entry is unscored, which is unknown rather than low risk. And KEV fires rarely on Python libraries — high precision, low recall — so its silence means nothing on its own.

Reachability

package-doctor also AST-parses your own source and reports where you import each dependency:

EXPOSED + NO ONE HOME   act on these
pillow  10.0.0  file/media parsing  CVE-2023-4863 on CISA's known-exploited list,
                                    and your pinned version is affected
                                    imported by your code at api/upload.py:2
paramiko 3.5.1  remote access       pinned version 3.5.1 is affected by 1 advisory:
                                    GHSA-r374-rxx8-8654
                                    imported by your code at api/sftp.py:12

Findings your code demonstrably imports sort to the top, because those are the ones you can act on today. Import names are mapped to distributions, so from PIL import Image is reported against pillow and import jwt against pyjwt. Imports that only appear in test code are labelled as such.

What "not imported" does not mean

It does not mean unreachable, and it never lowers a verdict. Your dependencies call each other: a package absent from your source can still run on every request. An import site is positive evidence that something is in play; its absence is not evidence of anything, and the tool says so rather than implying safety:

Reachability
  Imported by your code     no direct import found
    Not a safety finding: your dependencies call each other, so
    this can still run without appearing in your source.

A test asserts this invariant directly — the same package with an import site, without one, and never checked all produce the same verdict.

Use --src PATH to point at source directories explicitly, or --no-reachability to skip the scan.

What it measures, and what it refuses to

Release age is a famously bad signal on its own — it cannot tell an abandoned library from a finished one, or from a tool whose data updates server-side. So it is never sufficient here. Signals are split in two:

Authoritative (any one is enough — these are facts, not inferences):

  • the repository is archived
  • the maintainer set Development Status :: 7 - Inactive
  • an advisory whose affected range still includes the latest release — nobody has shipped a fix

Weak (two must agree before the tool says anything):

  • no release in 2 years
  • no commits in 18 months
  • a history of fixing advisories only after public disclosure

On "time to fix"

The obvious metric — days from advisory to patch — is wrong, and the data says so. Under coordinated disclosure a healthy project ships the patched release at or before the advisory goes public, so the median delta for well-run projects is zero or negative. Measured across real packages, Django fixed 153 of 161 advisories at or before disclosure, with 7 that cannot be dated; requests, 7 of 8.

GitHub's advisory backfill makes a naive reading worse still: an advisory written in 2022 for a fix shipped in 2012 produces a ten-year negative.

So package-doctor measures what the data actually supports:

  • never fixed — advisories whose affected range still includes the latest release. That is the only reading under which "nobody shipped a patch" is a fact: OSV closes ranges with fixed or with last_affected, and GitHub's advisory database encodes most older fixes as the latter. An earlier version of this tool only recognised fixed, and escalated Django 6 for a CSRF bug closed at 1.2.7. The strongest signal here, once read right.
  • closed, no fix named — the range ends before the latest release but no fix version is recorded, so there is nothing to put on a timeline. Shown in explain, and never counted against a package.
  • fixed late — how often a fix landed only after disclosure, and the median size of that window. PyYAML: 1 of 4, median 259 days.
  • fixed timely — the healthy case, shown so a good project reads as good.

One advisory is one advisory: OSV routinely carries a GHSA record and a PYSEC record for the same CVE, and they are merged before anything is counted.

An archived repository is not always an abandoned package

"Repository is archived" is a fact about the URL PyPI declares, and projects move: Google archived python-bigquery when it folded the package into a monorepo, and google-cloud-bigquery ships monthly. So an archived repository only settles the matter when nothing has been released in the stale window either. With a recent release it is stated, but as one weak signal, worded the code may have moved.

Missing data is never a bad score

If a package has no advisory history, that is unknown, not good. If it declares no repository, that is unknown, not bad. Findings with no signal go in their own section and are never counted against a package. Conflating null with zero is the most common flaw in package health tooling, and it is precisely what makes those tools flag quiet infrastructure libraries as failures.

There is no aggregate health score anywhere in this tool. A single number is the thing users cannot act on and maintainers cannot argue with.

Install

pip install package-doctor

A guardrail for coding agents

A scan finds problems after they are in the lockfile. A coding agent adds a dependency in the time it takes to complete an import statement and never reads the PyPI page, so the useful moment is before pip install runs.

package-doctor check reqeusts pillow==10.0.0 requests six
BLOCK     reqeusts
          not on PyPI: an install would fail, or fetch whatever someone has registered under this name since
          one edit from requests; did you mean that?
BLOCK     pillow 10.0.0
          exposed, and CVE-2023-4863 on CISA's known-exploited list, and your pinned version is affected
WARN      requests 2.34.2?
          at a trust boundary (http/network): 7 of 8 past advisories fixed at or before disclosure
OK        six 1.17.0?
          reviewed as not at a trust boundary

One decision per package: block, warn, ok, or unchecked. Exit 1 on a block (--fail-on warn to be stricter), --json for tools.

Agents fail in a way people rarely do: they invent names. So three facts can block on their own, and they are facts rather than inferences: the package is not on PyPI; it was first published within 30 days (--new-days); or it is one edit from a far more common name and not itself common, which is the shape of a typo or of a squat. These are about provenance, not exposure, and they never touch a verdict. Otherwise the scanner's verdict decides: act blocks, watch warns, the rest pass. And when the upstream services cannot answer, the package is unchecked and allowed — a guardrail that fails closed on somebody else's outage is the first thing a team removes.

Claude Code hook

Add to .claude/settings.json in the project, or ~/.claude/settings.json for every project:

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash",
        "hooks": [
          { "type": "command", "command": "package-doctor hook claude-code", "timeout": 60 }
        ]
      }
    ]
  }
}

Every shell command the agent proposes passes through the hook. Commands that install nothing are ignored in a millisecond. For pip install, uv add, poetry add, pipenv install, pdm add and pipx install, each package is checked, and:

  • a block stops the call and puts the reasons in front of the model — which is what makes an agent choose a different package instead of retrying the same one;
  • a warning is attached as context for the model, and the call goes through your normal permission flow — the hook never grants a permission you did not;
  • anything else is silent.

If the lookup itself fails, the install is allowed and the model is told it went unchecked. --warn-blocks makes warnings block too. Responses are cached, so the second check of a package costs nothing.

Edits to dependency files

An agent that writes a name into pyproject.toml and then runs uv sync never types the name into a shell command, so the install hook cannot see it. The same command also runs as a PostToolUse hook on edits:

{
  "hooks": {
    "PreToolUse": [
      { "matcher": "Bash",
        "hooks": [{ "type": "command", "command": "package-doctor hook claude-code", "timeout": 60 }] }
    ],
    "PostToolUse": [
      { "matcher": "Edit|Write|MultiEdit",
        "hooks": [{ "type": "command", "command": "package-doctor hook claude-code", "timeout": 60 }] }
    ]
  }
}

After an edit to pyproject.toml, Pipfile or a requirements*.txt, the hook reads the file and checks only the names the edit introduced: anything not already in a lockfile or in the version of the file git last committed. An edit to any other file, or to a lockfile, costs nothing. The edit has already happened, so this cannot block; the finding goes to the model as context, with the instruction to remove the package from the file before anything installs it — which is enough for an agent to fix its own mistake before a resolver runs.

Use

package-doctor scan
Dependency Risk Report   73 packages · 12 direct
from uv.lock, pyproject.toml

EXPOSED + NO ONE HOME    act on these
  legacy-auth   2.1.0   auth/session       repository is archived
                                           2 advisories with no published fix
  old-parser    1.7.4   html/xml parsing   marked Development Status :: 7 - Inactive

EXPOSED, MAINTAINED      watch
  requests      2.33.1  http/network, url parsing   7 of 8 past advisories
                                                    fixed at or before disclosure

STALE, NOT EXPOSED       low priority
  six           1.17.0  no boundary        reviewed as a finished utility

Then get the working behind any row:

package-doctor explain legacy-auth

Run inside the project, it reads the pinned version from the lockfile so the advisories are matched against what you actually install. Anywhere else, pass it: package-doctor explain pillow --pin 10.0.0.

No lockfile?

Then nothing says which version you run, and the most probable answer is the one a fresh pip install would pick today: the newest release that satisfies the declared range, skipping yanked and pre-release versions. package-doctor assumes that version, and says so everywhere the assumption shows:

8 of 8 without a pinned version: advisories matched against the newest release
instead, marked ? (use a lockfile to pin them)

EXPOSED, MAINTAINED   watch
httpx  0.28.1?  http/network  1 of 1 past advisories fixed at or before disclosure

The ? is the same mark an inferred exposure carries: something to distrust. In explain the version reads 0.28.1 (assumed), an advisory match reads newest release 0.28.1 (assumed: nothing pins this package) is affected, and the JSON carries "version_assumed": true. --no-assume-latest turns the fallback off and skips advisory matching for unpinned packages instead. A lockfile is still the answer; this is a first run, not a substitute.

Responses are cached for a day. package-doctor cache path shows where, and package-doctor cache clear empties it.

In CI

package-doctor scan --fail-on act

Exits 1 when anything lands in act on these, 0 otherwise. Use --fail-on watch to be stricter, --fail-on never to report only.

package-doctor scan --json -o report.json
package-doctor scan --sarif package-doctor.sarif     # for code scanning
package-doctor scan --markdown "$GITHUB_STEP_SUMMARY" # for the job summary

GitHub Actions

- uses: binuka200/package-doctor@v0.3.0

One line scans the checkout, fails the job on anything in act on these, and writes the report to the job's step summary so nobody opens a log. To have findings appear as code scanning alerts on the pull request as well:

permissions:
  security-events: write
steps:
  - uses: actions/checkout@v5
  - uses: binuka200/package-doctor@v0.3.0
    with:
      upload-sarif: true
      fail-on: act          # or watch, or never
      args: --direct-only   # anything `scan` takes

The action is package-doctor scan with --sarif and --markdown, plus a response cache carried between runs. Inputs, outputs and defaults are in action.yml.

SARIF

--sarif PATH writes a SARIF 2.1.0 report that GitHub code scanning, GitLab and most security dashboards ingest. Each act, watch or low finding is one result, at levels error, warning and note. Its location is the first place your own source imports the package, with the other import sites as related locations; a package your code never imports points at the line of the dependency file that declared it. Unknown and ok produce nothing, because an alert that can never be resolved is how a tool gets muted. Accepted risks (below) are carried as SARIF suppressions, so a dashboard shows them as dismissed with the reason rather than not at all.

pre-commit

- repo: https://github.com/binuka200/package-doctor
  rev: v0.3.0
  hooks:
    - id: package-doctor

Runs only when a dependency file changes, and the day-long response cache means the second run of the day is free.

Accepting a risk

The first true positive with a ticket already filed is where a scanner gets removed from CI. So a finding can be accepted, on the record, for a while:

# package-doctor.toml, next to the lockfile
[[accept]]
package = "legacy-auth"
reason = "Replacement lands in Q4, see PROJ-123"
until = 2026-12-31
version = "2.1.0"   # optional: an upgrade brings the finding back for a look

Three rules make this safe to offer. Every entry has a reason, because six months on an entry without one cannot be told from a mistake. Every entry expires, because a permanent suppression is how a known exposure survives every review; when the date passes the build fails again and the row says the acceptance expired, not that something new appeared. And an accepted finding stays in the report, in a section of its own, out of the exit code and never out of sight:

ACCEPTED RISK   on the record, not failing the build
  legacy-auth  2.1.0  act on these  Replacement lands in Q4, see PROJ-123
                                    until 2026-12-31 (in 3 months)

The verdict is not changed by acceptance. It is a fact about the package; the acceptance is a fact about what you will do. A malformed entry is a usage error rather than a warning, since an entry that silently failed to apply would fail a build for no visible reason, and one that silently applied too widely would hide risk. Entries that match nothing are reported so the file stays a list of live decisions. The same table can live under [tool.package-doctor] in pyproject.toml; --config PATH points anywhere else.

Options

Flag Meaning
--direct-only skip transitive dependencies
--src PATH source directory or file to check for imports (repeatable)
--no-reachability skip the import scan of your own source
--show-ok also list packages with no concerns
--sarif PATH also write a SARIF 2.1.0 report
--markdown PATH also write the report as Markdown
--config PATH accepted-risk file (default: package-doctor.toml)
--no-assume-latest with no pin, skip advisory matching rather than assume the newest release
--offline-repo skip repository lookups (faster, fewer signals)
--stale-release-days N tune the weak release-age signal (default 730)
--stale-push-days N tune the weak commit-age signal (default 545)
--max-packages N refuse to look up more than this many (default 2000)
--concurrency N parallel requests to the free upstream APIs (default 8)
--cache-ttl SECONDS how long responses are reused (default 86400)
--no-cache bypass the local response cache

Reads uv.lock, poetry.lock, Pipfile.lock, pyproject.toml (PEP 621, PEP 735 and Poetry), Pipfile, requirements*.txt, requirements/*.txt and one level below that (requirements/<env>/*.txt), following -r includes within the project. scan PATH takes a directory or one of those files.

Dependencies that come from git, a URL or a local path — a git+ssh:// line, an -e editable, a name @ url reference, a git source in a lockfile — are never looked up, because no registry can speak to them. They are listed under Not analysed rather than dropped: a first-party package at a trust boundary is the last thing a report should be quiet about.

Data sources

All free, all unauthenticated, no token setup:

  • PyPI JSON API — release timeline, classifiers, repository URL
  • OSV — advisories, affected ranges, fixed versions
  • ecosyste.ms — repository metadata, at 5,000 req/hour rather than GitHub's unauthenticated 60
  • CISA KEV — vulnerabilities known to be exploited in the wild
  • FIRST EPSS — exploit probability scores

Responses are cached in ~/.cache/package-doctor/ for 24 hours. A 429 or a 5xx is retried with backoff, honouring Retry-After; if requests still fail, the report says which host and how many, so a rate-limited run is visibly degraded rather than quietly incomplete.

The exposure map

exposure.toml is the part of this tool that cannot be scraped, and it is deliberately a plain data file so it can be argued with. A package belongs in it when it:

  1. parses, decodes, renders, verifies or transports data that commonly originates outside the trust boundary, or
  2. makes an authentication or authorisation decision, or
  3. constructs queries or commands from caller-supplied values.

"Popular" is not a criterion. Neither is "sounds security-adjacent" — xxhash and mmh3 are deliberately not in the crypto category, because they are not cryptographic, and tiktoken tokenises text rather than issuing auth tokens.

Deserialization is not the same as parsing

Two kinds of package turn bytes into data, and a flaw in them costs very different things. A pickle loader, an unsafe YAML loader, a config system that instantiates the class a document names: loading is code execution, by design. A JSON, MessagePack, Avro or date parser: the worst a hostile input does is exhaust memory or slip past validation. The map keeps them apart — deserialization for the first, data parsing for the second — because a stale cloudpickle and a stale ujson are not the same finding.

Every category names what a flaw at that boundary tends to cost, from a fixed vocabulary: code execution, memory corruption, file write, account takeover, data access, script injection, request forgery, prompt injection, denial of service. explain shows the word, the JSON carries it as exposure.consequence, and within a report section it breaks ties between findings with the same evidence, worst first. It is a word, not a number: nothing sums it, weights it, or lets it change a verdict.

Reviewed-and-safe is not the same as unreviewed

The map carries a third list, [reviewed] not_exposed, recording packages a human checked and decided are not at a trust boundary — each with a note saying why the obvious guess is wrong. Without it, classifier inference fires on exactly those names. It also means "we looked and it's fine" is distinguishable in the data from "nobody has looked yet", which is the distinction the rest of this tool is built on.

Depth is not the same as breadth

Beyond the top few hundred, most packages genuinely belong in neither list: a plotting library or a CLI helper needs no entry, and adding one would be noise. So raw "percent of PyPI covered" is the wrong measure. The one that matters is whether a real lockfile scans without gaps — which is why the map is curated against download-ranked data and checked against whole stacks rather than grown for its own sake.

One boundary most scanners miss

[category.ml_model] covers torch, transformers, huggingface-hub, joblib and friends. Loading pickle-based weights is arbitrary code execution, and tools aimed at web stacks tend not to model it at all. It is also where advisory data is least conclusive: torch and transformers each carry around thirty advisories, and seven or eight per package are closed in OSV without any fix version named — the unsafe-deserialisation reports their maintainers regard as by-design. The tool shows those as closed, no fix named and counts them against nothing, so that an ML stack is judged on what it actually imports, not on a pile of disputed pickle advisories.

PRs welcome — include the reasoning, not just the name.

A second reader

The map is one person's judgement, which is its stated weakness. research/annotate.py is the instrument for a second one: it draws a blind sample across the categories, the cleared list and unreviewed packages, walks a reader through each with the same evidence the curator had, and reports Cohen's kappa between the two on the distinction that matters — exposed or not — with every disagreement listed alongside both sides' reasoning. See CONTRIBUTING.md.

Growing it

Curating by working down a download list is brute force; most of what you read does not belong. research/suggest_map.py inverts that — it ranks the packages the map has no opinion about by how much that silence costs, and prints what each one is for so the decision takes seconds:

package-doctor scan . --json -o scan.json
python research/suggest_map.py --scan scan.json

python research/suggest_map.py --dataset data/pypi-top3000.jsonl --limit 30
  1. pytorch-lightning
     1 advisory never fixed
     "PyTorch Lightning is the lightweight PyTorch wrapper for ML researchers..."
     https://pypi.org/project/pytorch-lightning/

Unfixed advisories rank highest, because if such a package does turn out to sit at a trust boundary then the map is actively hiding something dangerous.

An advisory count is a reason to look, never the answer. num2words reached the top of that list with three unfixed advisories, and reading them showed a maintainer account compromise rather than anything the library does with input — so it belongs in [reviewed] not_exposed, with that noted.

A note on tone

Being listed here is not an accusation. Most unmaintained packages are the work of volunteers who gave what they could, and "no releases since 2021, repository archived" is a fact that helps a user without indicting anyone. Findings are worded that way on purpose. If you find output that reads as a judgement on a maintainer rather than a description of risk, that is a bug — please report it.

Contributing

The most useful contribution isn't code — it's arguing with exposure.toml. That file is ~710 judgement calls about which packages sit where an attacker can reach, all made by one person, and a wrong entry is worse than a missing one.

If you know a corner of Python well — Django, ML, crypto, packaging — twenty minutes reading the relevant category is worth more than a month of new entries.

git clone https://github.com/binuka200/package-doctor
cd package-doctor
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest

See CONTRIBUTING.md for what belongs in the map, what doesn't, and the four principles that are load-bearing enough to have tests guarding them.

Tests

pytest                 # the normal suite: offline, about 2s
pytest -m live         # contract tests against the real APIs

The default suite serves all HTTP from httpx.MockTransport, so no upstream outage can redden a build. That proves the code is self-consistent but nothing about upstream, so a second suite runs weekly against real PyPI, OSV, ecosyste.ms, CISA and FIRST endpoints to catch drift — a renamed field, a changed envelope, a tightened pagination cap.

Those tests draw one distinction deliberately: a service being unreachable skips, a service answering with a shape we do not expect fails. Somebody else's outage is not a defect here, and paging on it would train everyone to ignore the suite; drift silently corrupts findings and is exactly what the tests are for. There is a test for that distinction itself.

Accuracy

Measured, not asserted, on sixty open source repositories: web frameworks and the apps built on them, the pallets and pydantic families, LLM tooling, ML and data projects, infrastructure and CLI tools, from home-assistant at 1,247 packages down to ansible at 5. The list is research/eval-repos.txt and the harness is research/evaluate_repos.py; everything below reproduces from those two files.

repositories 60 (59 with a dependency file the tool reads)
packages assessed 12,973
distinct pinned (package, version) pairs 6,728
import sites reported 108,193

Advisory matching, against OSV's own version-scoped query — the authoritative answer to "is this pinned version affected?" — over all 6,728 pairs, after collapsing GHSA and PYSEC aliases to their CVE: 6,727 / 6,728 exact agreement. The one disagreement is langsmith 0.3.45, whose record carries a range typed SEMVER alongside its ECOSYSTEM range; OSV ignores the first for PyPI, this tool reads it, and reports the version affected. Earlier runs found and fixed three failures, each now a test: tornado 6.3.0 matched as a string rather than under PEP 440, scrapy 2.17.0 an open-ended range that a matcher waiting for a fixed event never reported, and click==8.* stored as a version rather than read as a range.

Against pip-audit (--no-deps -s osv) on the same pins, compared at the vulnerability level with identifiers canonicalised to CVE:

packages flagged, pip-audit / package-doctor 322 / 322
vulnerabilities found by both 1,672
found only by pip-audit 0
found only by package-doctor 1 (the langsmith range above)

Reachability: of 108,193 import sites reported, 108,162 open to an import of that package on that line. The 31 remaining are from _pytest... and import py attributed to pytest, which is correct — both modules ship in the pytest distribution — and only the checker's static table did not know it.

The honest limit: both tools read OSV, so this shows we read it correctly, not that OSV is complete. And it measures the data layer.

The verdicts, read

The verdict layer is a judgement, so the check is reading them. Of the 437 act on these verdicts across the sixty projects, 340 rest on the pinned version being affected by a published advisory, which the OSV agreement above makes right by construction; they are mostly old lockfiles. The other 97, on 28 distinct packages, rest on maintenance signals, and every one was read: libraries deprecated by their owners (adal, msrest, oauth2client, google-generativeai, redis-py-cluster), advisories the maintainers consider by-design and will not fix (nltk, keras, diskcache), and packages quiet for three to eight years. Two are arguable at the margin — python-pptx and requests-kerberos, each just over the two-year threshold — and that is a threshold choice rather than a misreading.

That result is recent. An earlier version of this tool recognised only a fixed event as closing an advisory's range, and on an eight-project sample 36 of its 64 act verdicts rested on advisories that were in fact closed by last_affected, Django 6 among them for a 2011 CSRF bug. Reading ranges as OSV does took the sixty-project run from 207 "never fixed" advisory records to 15.

The exposure map, measured

The map is the judgement layer, and it is a different kind of claim. Tested against 3,000 packages, asking whether its call predicts anything real — whether a package it marks exposed is more likely to have security advisories:

bucket packages have advisories
curated, exposed 635 31.3%
curated, not exposed 124 12.1%
no opinion 2,236 6.8%

Packages the map calls exposed carry advisories at ~2.6× the rate of packages it reviewed and cleared, and that holds when controlling for popularity (43.7% against 16.2% within the top 500). The curated map carries real signal.

The inference fallback did not. Packages it guessed as exposed had advisories at 11.3% — indistinguishable from the 12.1% of packages reviewed and marked not exposed. A hand audit showed why: Topic :: Security was catching bandit, semgrep and pip-audit, which are security tools; Topic :: System :: Archiving caught setuptools and wheel; Framework :: Django caught pytest-django and factory-boy. Roughly three in five were wrong.

So inference was cut back to the handful of classifiers that genuinely imply untrusted input, taking it from 151 packages to 5, and an inferred category can no longer produce an actionable verdict — it can raise something to watch and say why, and that is all. Saying "not reviewed" beats guessing.

What is still unmeasured

Coverage, on the 12,973 packages above: 45% get a curated call, 55% get no opinion. The larger the sample, the longer the tail of transitive dependencies, and that tail is what "no opinion" is for. Those 55% are not assessed as safe — they surface as unknown, which is the honest answer, but it is a gap rather than a result.

And there is no external review. Every batch of repositories scanned so far has turned up at least one miscategorisation, the rate is falling, and it is not zero. Treat the advisory numbers as solid and the categories as a draft.

Use it alongside pip-audit, not instead of it

pip-audit is the PyPA tool and is better at what it does: telling you, on every commit, which pinned versions have known CVEs. Most of what lands in act on these here, it would also find.

What it does not do is tell you which of those to fix first, which of them your code actually imports, or which of your dependencies has nobody left to ship a patch at all. That is this tool's job, and it is a different cadence — a quarterly maintenance review rather than a per-commit gate.

Background

The reasoning behind this tool is set out in Rethinking Dependency Maintenance in the Age of AI Vulnerability Research.

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

package_doctor-0.6.1.tar.gz (228.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

package_doctor-0.6.1-py3-none-any.whl (113.5 kB view details)

Uploaded Python 3

File details

Details for the file package_doctor-0.6.1.tar.gz.

File metadata

  • Download URL: package_doctor-0.6.1.tar.gz
  • Upload date:
  • Size: 228.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for package_doctor-0.6.1.tar.gz
Algorithm Hash digest
SHA256 b3e333804caf0b5cb256768aa4ded43efbb5b191c775cdd3e05ec6ec2072b69b
MD5 9e42816f548a7145209bb438eb678c5a
BLAKE2b-256 c8478a548809016c796f63349dc2d89b98b5aa39d01328899820731bd203ee7c

See more details on using hashes here.

Provenance

The following attestation bundles were made for package_doctor-0.6.1.tar.gz:

Publisher: publish.yml on binuka200/package-doctor

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file package_doctor-0.6.1-py3-none-any.whl.

File metadata

  • Download URL: package_doctor-0.6.1-py3-none-any.whl
  • Upload date:
  • Size: 113.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for package_doctor-0.6.1-py3-none-any.whl
Algorithm Hash digest
SHA256 99373143adfa7dee0bb913a3d01adb0dba9c8387a556cd0a588003281514b546
MD5 a54113e8a63545d13915966b21badec3
BLAKE2b-256 127d3d94868a072706a1d5467e025c262224fdd3bb402e3cf39f2053cb4dd900

See more details on using hashes here.

Provenance

The following attestation bundles were made for package_doctor-0.6.1-py3-none-any.whl:

Publisher: publish.yml on binuka200/package-doctor

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.8.4

2 files

0.8.3

2 files

0.8.2

2 files

0.8.1

2 files

0.8.0

2 files

0.7.0

2 files

This release

0.6.1 This release

2 files

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page