pdf-hunter
Search for and download PDF file links from a webpage.
Installation
pip install pdf_hunter
CLI Usage
After installing, the pdf-hunter executable is made available in your path.
1. List PDF links on a webpage
By default, running pdf-hunter with a webpage URL will print all discovered absolute PDF links to standard output:
pdf-hunter "https://example.com/books-list"
Output:
https://example.com/books/guide-to-python.pdf
https://example.com/books/advanced-algorithms.pdf
2. Download PDFs to current directory
Pass the -d (or --download) flag to download all discovered PDFs to your current directory:
pdf-hunter "https://example.com/books-list" -d
3. Download PDFs to a custom directory
Use the -o (or --output-dir) option to specify a target directory for the downloaded files:
pdf-hunter "https://example.com/books-list" -d -o /path/to/downloads
Python Library Usage
You can also use pdf-hunter programmatically in your Python scripts.
import pdf_hunter
url = "https://github.com/EbookFoundation/free-programming-books/blob/main/books/free-programming-books-langs.md"
Get a list of PDF URLs found on a page
pdf_urls = pdf_hunter.get_pdf_urls(url)
print(pdf_urls[:3])
Output:
[
"https://www.cs.uni.edu/~mccormic/4740/guide-c2ada.pdf",
"http://www.adapower.com/pdfs/AdaDistilled07-27-2003.pdf",
"https://www.adacore.com/uploads/books/pdf/Ada_for_the_C_or_Java_Developer-cc.pdf",
]
Download a single PDF file
import os
pdf_url = pdf_urls[0]
file_name = pdf_hunter.get_pdf_name(pdf_url)
# Download to a specific directory
pdf_hunter.download_file(pdf_url, folder_path=os.getcwd())
print(os.path.isfile(file_name)) # True
Download all PDF files from a page
pdf_hunter.download_pdf_files(url, folder_path=os.getcwd())
Release files for pdf-hunter 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pdf_hunter-1.1.0.tar.gz | 3.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pdf_hunter-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 7.7 kB
Release files / pdf_hunter-1.1.0.tar.gz
| Download URL | pdf_hunter-1.1.0.tar.gz |
|---|---|
| Size | 3.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f67f5f2d7125cf7efb7be7daaef73df56f1a2264c4ed5dc34f0828a2b826c8ce
|
|
BLAKE2b-256 checksum How to use checksums |
8042ab7ce40599ef1d69a65d1039f1d9380c943b9dd67873873529399621090b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.2 {"installer":{"name":"uv","version":"0.12.2","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / pdf_hunter-1.1.0-py3-none-any.whl
| Download URL | pdf_hunter-1.1.0-py3-none-any.whl |
|---|---|
| Size | 4.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
828b10f880e91eea1b9bcd31e8818721b5112b0473f673e2731d73584b4ae16b
|
|
BLAKE2b-256 checksum How to use checksums |
8d153bce65c192d8a33920f02dd2070048294244b67e685d0475630c73d2f676
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.2 {"installer":{"name":"uv","version":"0.12.2","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|