Skip to main content

Web Scraper to Markdown 🌐✍️

This Python-based web scraper fetches content from URLs and exports it into Markdown and JSON formats, specifically designed for simplicity, extensibility, and for uploading JSON files to GPT models. It is ideal for those looking to leverage web content for AI training or analysis. 🤖💡

🚀 Quick Start

(Or even better, use Docker! 🐳)

Recommended installation using pipx (isolated environment)

pipx install crawler-to-md

Alternatively, install with pip

pip install crawler-to-md

Then run the scraper:

crawler-to-md --url https://www.example.com

🌟 Features

  • Scrapes web pages for content and metadata. 📄
  • Filters links by base URL. 🔍
  • Excludes URLs containing certain strings. ❌
  • Automatically finds links or can use a file of URLs to scrape. 🔗
  • Rate limiting and delay support. 🕘
  • Exports data to Markdown and JSON, ready for GPT uploads. 📤
  • Exports each page as an individual Markdown file if --export-individual is used. 📝
  • Uses SQLite for efficient data management. 📊
  • Configurable via command-line arguments. ⚙️
  • Include or exclude specific HTML elements using CSS-like selectors (#id, .class, tag) during Markdown conversion. 🧩
  • Docker support. 🐳

📋 Requirements

Python 3.10 or higher is required.

Project dependencies are managed with pyproject.toml. Install them with:

pip install .

🛠 Usage

Start scraping with the following command:

crawler-to-md --url <URL> [--output-folder ./output] [--cache-folder ./cache] [--overwrite-cache|-w] [--base-url <BASE_URL>] [--exclude-url <KEYWORD_IN_URL>] [--title <TITLE>] [--urls-file <URLS_FILE>] [-p <PROXY_URL>]

Options:

  • --url, -u: The starting URL. 🌍
  • --urls-file: Path to a file containing URLs to scrape, one URL per line. If '-', read from stdin. 📁
  • --output-folder, -o: Where to save Markdown files (default: ./output). 📂
  • --cache-folder, -c: Where to store the database (default: ./cache). 💾
  • --overwrite-cache, -w: Overwrite existing cache database before scraping. 🧹
  • --base-url, -b: Filter links by base URL (default: URL's base). 🔎
  • --title, -t: Final title of the markdown file. Defaults to the URL. 🏷️
  • --exclude-url, -e: Exclude URLs containing this string (repeatable). ❌
  • --export-individual, -ei: Export each page as an individual Markdown file. 📝
  • --rate-limit, -rl: Maximum number of requests per minute (default: 0, no rate limit). ⏱️
  • --delay, -d: Delay between requests in seconds (default: 0, no delay). 🕒
  • --proxy, -p: Proxy URL for HTTP or SOCKS requests. 🌐
  • --include, -i: CSS-like selector (#id, .class, tag) to include before Markdown conversion (repeatable). ✅
  • --exclude, -x: CSS-like selector (#id, .class, tag) to exclude before Markdown conversion (repeatable). 🚫

One of the --url or --urls-file options is required.

📚 Log level

By default, the WARN level is used. You can change it with the LOG_LEVEL environment variable.

🐳 Docker Support

Run with Docker:

docker run --rm \
  -v $(pwd)/output:/app/output \
  -v cache:/home/app/.cache/crawler-to-md \
  ghcr.io/obeone/crawler-to-md --url <URL>

Build from source:

docker build -t crawler-to-md .

docker run --rm \
  -v $(pwd)/output:/app/output \
  crawler-to-md --url <URL>

🤝 Contributing

Contributions are welcome! Feel free to submit pull requests or open issues. 🌟

Metadata

Release files for crawler-to-md 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for crawler-to-md 0.6.0
File Size Uploaded
crawler_to_md-0.6.0.tar.gz 23.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for crawler-to-md 0.6.0
File Interpreter ABI Platform
crawler_to_md-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 39.1 kB

Release files / crawler_to_md-0.6.0.tar.gz

Download URL crawler_to_md-0.6.0.tar.gz
Size 23.3 kB
Tags Source
SHA-256 checksum
How to use checksums
a980247e7af2d76cb767d8ba36fbae8b1e45b1015474d80dc2f5af0f4e53bf7d
BLAKE2b-256 checksum
How to use checksums
74518ea06aa9be1d6f6b00bff25dc5a08437f146c67ff885e130a41a03f41e26
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.0

Release files / crawler_to_md-0.6.0-py3-none-any.whl

Download URL crawler_to_md-0.6.0-py3-none-any.whl
Size 15.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
797ec37e3b4d46c6c4af627a53345c14728d776a51cea95f649b187466ace3de
BLAKE2b-256 checksum
How to use checksums
096652291bfd1b5b5c76850a9544f40d6b5ff9cb118b5f8aa8a8e2e763dd5487
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.0

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page