Stop googling “best python library for X”.
Tell pipscout what you're building. It combs through PyPI and hands you the package everyone actually uses — with receipts.
pip install pipscout
pipscout "pdf extraction"
That's it. 🔭 One command gets you the best pick, the runners-up, their monthly downloads, how fresh each release is, and the exact pip install line.
✨ Why you'll like it
| 🗣️ Plain English in | "pdf extraction", "i want to parse yaml files", "excel spreadsheets" — it strips the filler, knows extraction ≈ extract ≈ parse, and gets to work. |
| 🏆 The favourite out | Ranks by what matters: is it on-topic, does the world actually use it (30-day downloads), is it still maintained. Abandoned and “Inactive” projects sink. |
| 🧾 Receipts, not vibes | Every pick shows why: which words matched, downloads per month, last release, relevance & freshness meters. |
| ⚡ Fast | ~2 s on a cold cache, near-instant after — metadata is cached for a week and fetched 16-wide in parallel. |
| 🪶 Featherweight | One dependency (rich, for the drip). Ships a 170 KB catalog of PyPI's top 5,000 projects so it's smart from the very first run. |
| 🤖 Script-friendly | --json output and a one-function Python API. |
🚀 Take it for a spin
pipscout "web scraping" # 🥇🥈🥉 top 5
pipscout "data validation" -n 10 # show me more
pipscout "http client" --json | jq # pipe it anywhere
pipscout "pdf extraction" --deep # also scan the names of all ~700k projects on PyPI
pipscout index # optional: pre-index the top 5,000 so READMEs are searchable too
👀
beautifulsoup4won “web scraping” even though neither word is in its name — pipscout reads summaries and keywords (and, afterpipscout index, READMEs), not just names.
Real results, fresh install, no index
| You type | pipscout says |
|---|---|
pdf extraction |
pymupdf → pypdf → pdfminer.six → pdfplumber |
web scraping |
beautifulsoup4 → htmldate → trafilatura → firecrawl-py → Scrapy |
parse yaml files |
PyYAML → ruamel.yaml |
data validation |
pydantic → jsonschema → email-validator |
http client |
aiohttp → httpx → httpcore → urllib3 |
excel spreadsheets |
xlrd → openpyxl → xlsxwriter → gspread |
plotting charts |
matplotlib → pyecharts |
🐍 From Python
from pipscout import recommend
for rec in recommend("pdf extraction", limit=2):
print(f"{rec.name:<8} {rec.score:.2f} {rec.downloads or 0:>12,} dl/mo")
print(" ", rec.install, "|", "; ".join(rec.reasons))
pymupdf 0.46 101,779,570 dl/mo
pip install pymupdf | summary matches 'pdf'; summary matches 'extract'; 101.8M downloads in the last 30 days
pypdf 0.30 151,277,866 dl/mo
pip install pypdf | summary matches 'pdf'; description matches 'extract'; 151.3M downloads in the last 30 days
Each Recommendation has name, score, relevance, popularity, health, downloads, summary, version, last_release, homepage, install, reasons, warnings — plus .to_dict() for JSON.
🧠 How the magic works
PyPI's website search has no public API (and sits behind a bot wall), so pipscout brings its own brain:
"i want pdf extraction"
│
1. 🗣️ understand ── drop filler · stem (extraction → extract) · add synonyms (extract ≈ parse ≈ read)
│
2. 🔎 shortlist ─── names of the ~15k most-downloaded projects
│ + summaries & keywords from the bundled top-5k catalog
│ + READMEs from your local index · + every name on PyPI with --deep
│
3. 📡 inspect ───── live metadata for the best 60 from PyPI's JSON API (parallel, cached 7 days)
│
4. 🏆 rank ──────── score = relevance² × popularity × √freshness
- relevance (0–1) — how many of your concepts a package covers, each weighted by how rare the word is on PyPI (so
pdfoutweighsdata). A hit in the name or summary counts more than one buried in the README; a synonym counts a bit less than the real word. Anything under 0.5 is dropped. - popularity —
(monthly downloads ÷ 1B)^¼. A package used 10× more can beat a slightly more on-topic one… but never an off-topic one. - freshness — 1.0 if released in the last year, sliding to 0.2 at five years, with an extra haircut for Inactive or Alpha status.
Download counts come from the excellent top-pypi-packages dataset (PyPI's public BigQuery stats, refreshed monthly).
🎛️ CLI reference
| Command / flag | What it does |
|---|---|
pipscout "<need>" |
Recommend packages (shorthand for pipscout search "<need>") |
-n, --limit N |
How many results (default 5) |
--json |
Machine-readable output |
--deep |
Also match names across every project on PyPI (fetches the ~45 MB name index, cached for a day) |
--min-relevance X |
Loosen or tighten the on-topic bar (0–1, default 0.5) |
--max-fetch N |
Candidates to inspect in detail (default 60) |
--refresh |
Ignore the cache and re-fetch |
pipscout index [--top N] |
Pre-fetch metadata for the top N packages (default 5,000, ~3 min) |
pipscout clear-cache |
Wipe the cache (~/.cache/pipscout, or $PIPSCOUT_CACHE_DIR) |
🤔 FAQ
Is it just sorting by downloads? Nope. Downloads only count once a package is on-topic — relevance is squared, so an off-topic giant (looking at you, requests) never wins a PDF query.
Why not just use PyPI search? It sorts by text match, not by what people actually use, has no API, and won't tell you a project died in 2017.
Does it phone home? Only to pypi.org and raw.githubusercontent.com (for the download stats). No telemetry, no accounts, no API keys.
It missed my favourite package! Try --deep, rephrase with the words the package would use to describe itself, or run pipscout index once so READMEs get searched too. PRs to the synonym list in text.py are very welcome.
🛠️ Hacking on it
git clone https://github.com/Meet2147/pythonLibraries && cd pythonLibraries/pipscout
pip install -e ".[dev]"
pytest # fully offline — PyPI is faked
python scripts/render_demo.py # re-shoot the terminal screenshots
pipscout index && python scripts/build_bundled_data.py # refresh the bundled catalog
📜 License
MIT © Meet2147 — now go build something. 🔭
Metadata
Release files for pipscout 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pipscout-0.1.0.tar.gz | 218.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pipscout-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 427.4 kB
Release files / pipscout-0.1.0.tar.gz
| Download URL | pipscout-0.1.0.tar.gz |
|---|---|
| Size | 218.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b17d73e8c6831018d05573cdef2e0949990063decced1ec653c2c5d02505b363
|
|
BLAKE2b-256 checksum How to use checksums |
f931cdbd8cbcf53daf6dbf9944fede3f95f873fca5f316b3bf02eeeb2e02a6e2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / pipscout-0.1.0-py3-none-any.whl
| Download URL | pipscout-0.1.0-py3-none-any.whl |
|---|---|
| Size | 208.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
41da3ad20f45bb19b54ba410c78b22b658dca97942c765caa7bff93c255f32da
|
|
BLAKE2b-256 checksum How to use checksums |
1d2f4adadbc3ec5504ec5a67e195e61acf0013f50d3851cb7227ea24c49a1e5c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log