Stop googling “best python library for X”.
Tell pipscout what you're building. It combs through PyPI and hands you the package everyone actually uses — with receipts.
pip install pipscout
pipscout "pdf extraction"
That's it. 🔭 One command gets you the best pick, the runners-up, their monthly downloads, how fresh each release is, and the exact pip install line.
✨ Why you'll like it
| 🗣️ Plain English in | "pdf extraction", "i want to parse yaml files", "excel spreadsheets" — it strips the filler, knows extraction ≈ extract ≈ parse, and gets to work. |
| 🏆 The favourite out | Ranks by what matters: is it on-topic, does the world actually use it (30-day downloads), is it still maintained. Abandoned and “Inactive” projects sink. |
| 🧾 Receipts, not vibes | Every pick shows why: which words matched, downloads per month, last release, relevance & freshness meters. |
| ⚡ Fast | ~2 s on a cold cache, near-instant after — metadata is cached for a week and fetched 16-wide in parallel. |
| 🪶 Featherweight | One dependency (rich, for the drip). Ships a 170 KB catalog of PyPI's top 5,000 projects so it's smart from the very first run. |
| 📈 Download stats on tap | pipscout downloads pypdf → yesterday / week / month, PyPI rank, 30-day trend and a weekly chart. Name several to race them head-to-head. |
| 🤖 Script-friendly | --json output and a one-function Python API. |
🚀 Take it for a spin
pipscout "web scraping" # 🥇🥈🥉 top 5
pipscout "data validation" -n 10 # show me more
pipscout "http client" --json | jq # pipe it anywhere
pipscout "pdf extraction" --deep # also scan the names of all ~700k projects on PyPI
pipscout downloads pypdf pymupdf # 📈 who's winning? downloads, rank, trend, chart
pipscout index # optional: pre-index the top 5,000 so READMEs are searchable too
👀
beautifulsoup4won “web scraping” even though neither word is in its name — pipscout reads summaries and keywords (and, afterpipscout index, READMEs), not just names.
Real results, fresh install, no index
| You type | pipscout says |
|---|---|
pdf extraction |
pymupdf → pypdf → pdfminer.six → pdfplumber |
web scraping |
beautifulsoup4 → htmldate → trafilatura → firecrawl-py → Scrapy |
parse yaml files |
PyYAML → ruamel.yaml |
data validation |
pydantic → jsonschema → email-validator |
http client |
aiohttp → httpx → httpcore → urllib3 |
excel spreadsheets |
xlrd → openpyxl → xlsxwriter → gspread |
plotting charts |
matplotlib → pyecharts |
📈 How many people use it?
Already have a package in mind? Get its download numbers straight away:
pipscout downloads pymupdf # one package: the full card + a weekly chart
pipscout dl pypdf pymupdf pdfplumber # several: a head-to-head table, 👑 for the leader
pipscout dl httpx --json # the raw numbers, incl. ~180 days of daily counts
One package gets a card:
- downloads yesterday, over the last 7 days and over the last 30 days
- its rank among PyPI's most-downloaded projects
- a ▲/▼ 30-day trend (compared with the 30 days before)
- a column chart of weekly downloads for the last ~6 months, plus the peak day
Several packages get a leaderboard with rank, week, month, share-of-the-leader bars, trend and a sparkline for each.
Daily numbers come from pypistats.org (mirrors excluded) and are cached for 12 hours. If pypistats can't be reached, pipscout still shows each package's 30-day total and rank from the top-15k download list.
🐍 From Python
from pipscout import recommend
for rec in recommend("pdf extraction", limit=2):
print(f"{rec.name:<8} {rec.score:.2f} {rec.downloads or 0:>12,} dl/mo")
print(" ", rec.install, "|", "; ".join(rec.reasons))
pymupdf 0.46 101,779,570 dl/mo
pip install pymupdf | summary matches 'pdf'; summary matches 'extract'; 101.8M downloads in the last 30 days
pypdf 0.30 151,277,866 dl/mo
pip install pypdf | summary matches 'pdf'; description matches 'extract'; 151.3M downloads in the last 30 days
Download numbers are one call away too:
from pipscout import download_stats, download_stats_many
s = download_stats("pypdf")
s.last_day, s.last_week, s.last_month # ints
s.rank # position among the ~15k most-downloaded projects
s.trend # last 30 days vs the 30 before, e.g. 0.12 == +12%
s.daily # [("2026-04-01", 5123456), ...] ~180 days, oldest first
s.weekly # the same, summed per week
for s in download_stats_many(["httpx", "aiohttp", "requests"]):
print(s.name, s.last_month)
Each Recommendation has name, score, relevance, popularity, health, downloads, summary, version, last_release, homepage, install, reasons, warnings — plus .to_dict() for JSON.
🧠 How the magic works
PyPI's website search has no public API (and sits behind a bot wall), so pipscout brings its own brain:
"i want pdf extraction"
│
1. 🗣️ understand ── drop filler · stem (extraction → extract) · add synonyms (extract ≈ parse ≈ read)
│
2. 🔎 shortlist ─── names of the ~15k most-downloaded projects
│ + summaries & keywords from the bundled top-5k catalog
│ + READMEs from your local index · + every name on PyPI with --deep
│
3. 📡 inspect ───── live metadata for the best 60 from PyPI's JSON API (parallel, cached 7 days)
│
4. 🏆 rank ──────── score = relevance² × popularity × √freshness
- relevance (0–1) — how many of your concepts a package covers, each weighted by how rare the word is on PyPI (so
pdfoutweighsdata). A hit in the name or summary counts more than one buried in the README; a synonym counts a bit less than the real word. Anything under 0.5 is dropped. - popularity —
(monthly downloads ÷ 1B)^¼. A package used 10× more can beat a slightly more on-topic one… but never an off-topic one. - freshness — 1.0 if released in the last year, sliding to 0.2 at five years, with an extra haircut for Inactive or Alpha status.
Download counts come from the excellent top-pypi-packages dataset (PyPI's public BigQuery stats, refreshed monthly).
🎛️ CLI reference
| Command / flag | What it does |
|---|---|
pipscout "<need>" |
Recommend packages (shorthand for pipscout search "<need>") |
-n, --limit N |
How many results (default 5) |
--json |
Machine-readable output |
--deep |
Also match names across every project on PyPI (fetches the ~45 MB name index, cached for a day) |
--min-relevance X |
Loosen or tighten the on-topic bar (0–1, default 0.5) |
--max-fetch N |
Candidates to inspect in detail (default 60) |
--refresh |
Ignore the cache and re-fetch |
pipscout downloads PKG [PKG…] |
Download stats for one package, or a head-to-head for several (alias: pipscout dl) |
--json / --no-history |
Raw numbers / skip the ~180-day daily history |
pipscout index [--top N] |
Pre-fetch metadata for the top N packages (default 5,000, ~3 min) |
pipscout clear-cache |
Wipe the cache (~/.cache/pipscout, or $PIPSCOUT_CACHE_DIR) |
🤔 FAQ
Is it just sorting by downloads? Nope. Downloads only count once a package is on-topic — relevance is squared, so an off-topic giant (looking at you, requests) never wins a PDF query.
Why not just use PyPI search? It sorts by text match, not by what people actually use, has no API, and won't tell you a project died in 2017.
Does it phone home? Only to pypi.org, raw.githubusercontent.com (monthly download rankings) and, for pipscout downloads, pypistats.org. No telemetry, no accounts, no API keys.
It missed my favourite package! Try --deep, rephrase with the words the package would use to describe itself, or run pipscout index once so READMEs get searched too. PRs to the synonym list in text.py are very welcome.
🛠️ Hacking on it
git clone https://github.com/Meet2147/pythonLibraries && cd pythonLibraries/pipscout
pip install -e ".[dev]"
pytest # fully offline — PyPI is faked
python scripts/render_demo.py # re-shoot the terminal screenshots
pipscout index && python scripts/build_bundled_data.py # refresh the bundled catalog
📜 License
MIT © Meet2147 — now go build something. 🔭
Metadata
Release files for pipscout 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pipscout-0.2.0.tar.gz | 223.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pipscout-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 437.4 kB
Release files / pipscout-0.2.0.tar.gz
| Download URL | pipscout-0.2.0.tar.gz |
|---|---|
| Size | 223.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d61204cfd26ffdd221099e3af38b60288ef3c4a7a0f580fd041ce62d0dd05aa0
|
|
BLAKE2b-256 checksum How to use checksums |
e4d269d1a37dda2f7ccdbfb4847e0bf9496a2393ce86c33d91e8505c906e84a4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / pipscout-0.2.0-py3-none-any.whl
| Download URL | pipscout-0.2.0-py3-none-any.whl |
|---|---|
| Size | 213.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e63d4a7e98af2f804803dfbc057b9fe0d2b163b826b0dbbcaa8e1faa7fe71b9a
|
|
BLAKE2b-256 checksum How to use checksums |
693d329a056f55a172de7f22c815b8c8b8f98af3a77ba627608d52cf1dad1a4a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log