Skip to main content

pipscout — say what you need, get the best PyPI package

PyPI Python 3.9+ CI Dependencies: just rich MIT license

Stop googling “best python library for X”.
Tell pipscout what you're building. It combs through PyPI and hands you the package everyone actually uses — with receipts.


pip install pipscout
pipscout "pdf extraction"

pipscout recommending pymupdf, pypdf, pdfminer.six and pdfplumber for 'pdf extraction'

That's it. 🔭 One command gets you the best pick, the runners-up, their monthly downloads, how fresh each release is, and the exact pip install line.

✨ Why you'll like it

🗣️ Plain English in "pdf extraction", "i want to parse yaml files", "excel spreadsheets" — it strips the filler, knows extraction ≈ extract ≈ parse, and gets to work.
🏆 The favourite out Ranks by what matters: is it on-topic, does the world actually use it (30-day downloads), is it still maintained. Abandoned and “Inactive” projects sink.
🧾 Receipts, not vibes Every pick shows why: which words matched, downloads per month, last release, relevance & freshness meters.
⚡ Fast ~2 s on a cold cache, near-instant after — metadata is cached for a week and fetched 16-wide in parallel.
🪶 Featherweight One dependency (rich, for the drip). Ships a 170 KB catalog of PyPI's top 5,000 projects so it's smart from the very first run.
📈 Download stats on tap pipscout downloads pypdf → yesterday / week / month, PyPI rank, 30-day trend and a weekly chart. Name several to race them head-to-head.
🤖 Script-friendly --json output and a one-function Python API.

🚀 Take it for a spin

pipscout "web scraping"                  # 🥇🥈🥉 top 5
pipscout "data validation" -n 10         # show me more
pipscout "http client" --json | jq       # pipe it anywhere
pipscout "pdf extraction" --deep         # also scan the names of all ~700k projects on PyPI
pipscout downloads pypdf pymupdf         # 📈 who's winning? downloads, rank, trend, chart
pipscout index                           # optional: pre-index the top 5,000 so READMEs are searchable too

pipscout recommending beautifulsoup4 for 'web scraping'

👀 beautifulsoup4 won “web scraping” even though neither word is in its name — pipscout reads summaries and keywords (and, after pipscout index, READMEs), not just names.

Real results, fresh install, no index

You type pipscout says
pdf extraction pymupdf → pypdf → pdfminer.six → pdfplumber
web scraping beautifulsoup4 → htmldate → trafilatura → firecrawl-py → Scrapy
parse yaml files PyYAML → ruamel.yaml
data validation pydantic → jsonschema → email-validator
http client aiohttp → httpx → httpcore → urllib3
excel spreadsheets xlrd → openpyxl → xlsxwriter → gspread
plotting charts matplotlib → pyecharts

📈 How many people use it?

Already have a package in mind? Get its download numbers straight away:

pipscout downloads pymupdf               # one package: the full card + a weekly chart
pipscout dl pypdf pymupdf pdfplumber     # several: a head-to-head table, 👑 for the leader
pipscout dl httpx --json                 # the raw numbers, incl. ~180 days of daily counts

One package gets a card:

  • downloads yesterday, over the last 7 days and over the last 30 days
  • its rank among PyPI's most-downloaded projects
  • a ▲/▼ 30-day trend (compared with the 30 days before)
  • a column chart of weekly downloads for the last ~6 months, plus the peak day

Several packages get a leaderboard with rank, week, month, share-of-the-leader bars, trend and a sparkline for each.

Daily numbers come from pypistats.org (mirrors excluded) and are cached for 12 hours. If pypistats can't be reached, pipscout still shows each package's 30-day total and rank from the top-15k download list.

🐍 From Python

from pipscout import recommend

for rec in recommend("pdf extraction", limit=2):
    print(f"{rec.name:<8} {rec.score:.2f}  {rec.downloads or 0:>12,} dl/mo")
    print("   ", rec.install, "|", "; ".join(rec.reasons))
pymupdf  0.46   101,779,570 dl/mo
    pip install pymupdf | summary matches 'pdf'; summary matches 'extract'; 101.8M downloads in the last 30 days
pypdf    0.30   151,277,866 dl/mo
    pip install pypdf | summary matches 'pdf'; description matches 'extract'; 151.3M downloads in the last 30 days

Download numbers are one call away too:

from pipscout import download_stats, download_stats_many

s = download_stats("pypdf")
s.last_day, s.last_week, s.last_month   # ints
s.rank                                   # position among the ~15k most-downloaded projects
s.trend                                  # last 30 days vs the 30 before, e.g. 0.12 == +12%
s.daily                                  # [("2026-04-01", 5123456), ...] ~180 days, oldest first
s.weekly                                 # the same, summed per week

for s in download_stats_many(["httpx", "aiohttp", "requests"]):
    print(s.name, s.last_month)

Each Recommendation has name, score, relevance, popularity, health, downloads, summary, version, last_release, homepage, install, reasons, warnings — plus .to_dict() for JSON.

🧠 How the magic works

PyPI's website search has no public API (and sits behind a bot wall), so pipscout brings its own brain:

  "i want pdf extraction"
          │
   1. 🗣️  understand ── drop filler · stem (extraction → extract) · add synonyms (extract ≈ parse ≈ read)
          │
   2. 🔎  shortlist ─── names of the ~15k most-downloaded projects
          │             + summaries & keywords from the bundled top-5k catalog
          │             + READMEs from your local index · + every name on PyPI with --deep
          │
   3. 📡  inspect ───── live metadata for the best 60 from PyPI's JSON API (parallel, cached 7 days)
          │
   4. 🏆  rank ──────── score = relevance² × popularity × √freshness
  • relevance (0–1) — how many of your concepts a package covers, each weighted by how rare the word is on PyPI (so pdf outweighs data). A hit in the name or summary counts more than one buried in the README; a synonym counts a bit less than the real word. Anything under 0.5 is dropped.
  • popularity — (monthly downloads ÷ 1B)^¼. A package used 10× more can beat a slightly more on-topic one… but never an off-topic one.
  • freshness — 1.0 if released in the last year, sliding to 0.2 at five years, with an extra haircut for Inactive or Alpha status.

Download counts come from the excellent top-pypi-packages dataset (PyPI's public BigQuery stats, refreshed monthly).

🎛️ CLI reference

Command / flag What it does
pipscout "<need>" Recommend packages (shorthand for pipscout search "<need>")
-n, --limit N How many results (default 5)
--json Machine-readable output
--deep Also match names across every project on PyPI (fetches the ~45 MB name index, cached for a day)
--min-relevance X Loosen or tighten the on-topic bar (0–1, default 0.5)
--max-fetch N Candidates to inspect in detail (default 60)
--refresh Ignore the cache and re-fetch
pipscout downloads PKG [PKG…] Download stats for one package, or a head-to-head for several (alias: pipscout dl)
--json / --no-history Raw numbers / skip the ~180-day daily history
pipscout index [--top N] Pre-fetch metadata for the top N packages (default 5,000, ~3 min)
pipscout clear-cache Wipe the cache (~/.cache/pipscout, or $PIPSCOUT_CACHE_DIR)

🤔 FAQ

Is it just sorting by downloads? Nope. Downloads only count once a package is on-topic — relevance is squared, so an off-topic giant (looking at you, requests) never wins a PDF query.

Why not just use PyPI search? It sorts by text match, not by what people actually use, has no API, and won't tell you a project died in 2017.

Does it phone home? Only to pypi.org, raw.githubusercontent.com (monthly download rankings) and, for pipscout downloads, pypistats.org. No telemetry, no accounts, no API keys.

It missed my favourite package! Try --deep, rephrase with the words the package would use to describe itself, or run pipscout index once so READMEs get searched too. PRs to the synonym list in text.py are very welcome.

🛠️ Hacking on it

git clone https://github.com/Meet2147/pythonLibraries && cd pythonLibraries/pipscout
pip install -e ".[dev]"
pytest                                                   # fully offline — PyPI is faked
python scripts/render_demo.py                            # re-shoot the terminal screenshots
pipscout index && python scripts/build_bundled_data.py   # refresh the bundled catalog

📜 License

MIT © Meet2147 — now go build something. 🔭

Metadata

Release files for pipscout 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pipscout 0.2.0
File Size Uploaded
pipscout-0.2.0.tar.gz 223.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pipscout 0.2.0
File Interpreter ABI Platform
pipscout-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 437.4 kB

Release files / pipscout-0.2.0.tar.gz

Download URL pipscout-0.2.0.tar.gz
Size 223.6 kB
Tags Source
SHA-256 checksum
How to use checksums
d61204cfd26ffdd221099e3af38b60288ef3c4a7a0f580fd041ce62d0dd05aa0
BLAKE2b-256 checksum
How to use checksums
e4d269d1a37dda2f7ccdbfb4847e0bf9496a2393ce86c33d91e8505c906e84a4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / pipscout-0.2.0-py3-none-any.whl

Download URL pipscout-0.2.0-py3-none-any.whl
Size 213.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e63d4a7e98af2f804803dfbc057b9fe0d2b163b826b0dbbcaa8e1faa7fe71b9a
BLAKE2b-256 checksum
How to use checksums
693d329a056f55a172de7f22c815b8c8b8f98af3a77ba627608d52cf1dad1a4a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page