Skip to main content

Web2LLM

CI/CD Pipeline

A command-line tool to scrape web pages, GitHub repos, local folders, and PDFs into clean, aggregated Markdown suitable for Large Language Models.

Description

This tool provides a unified interface to process various sources—from live websites and code repositories to local directories and PDF files—and convert them into a structured Markdown format. The clean, token-efficient output is ideal for use as context in prompts for Large Language Models, for Retrieval-Augmented Generation (RAG) pipelines, or for documentation archiving.

Installation

For standard scraping of static websites, local files, and GitHub repositories, install the base package:

pip install web2llm

To enable JavaScript rendering for Single-Page Applications (SPAs) and other dynamic websites, you must install the [js] extra, which includes Playwright:

pip install "web2llm[js]"

After installing the js extra, you must also download the necessary browser binaries for Playwright to function:

playwright install

Usage

Command-Line Interface

The tool is run from the command line with the following structure:

web2llm <SOURCE> -o <OUTPUT_NAME> [OPTIONS]
  • <SOURCE>: The URL or local path to scrape.
  • -o, --output: The base name for the output folder and the .md and .json files created inside it.

All scraped content is saved to a new directory at output/<OUTPUT_NAME>/.

General Options:

  • --debug: Enable debug mode for verbose, step-by-step output to stderr.

Web Scraper Options (For URLs):

  • --render-js: Render JavaScript using a headless browser. Slower but necessary for SPAs. Requires installation with the [js] extra.
  • --check-content-type: Force a network request to check the page's Content-Type header. Use for URLs that serve PDFs without a .pdf extension.

Filesystem Options (For GitHub & Local Folders):

When scraping a local folder or a GitHub repository, web2llm will automatically find and respect the rules in the project's .gitignore file. This ensures that the scrape accurately reflects the intended source code of the project.

  • --exclude <PATTERN>: A .gitignore-style pattern for files/directories to exclude. Can be used multiple times.
  • --include <PATTERN>: A pattern to re-include a file that would otherwise be ignored by default or by an --exclude rule. Can be used multiple times.
  • --include-all: Disables all default, project-level, and .gitignore ignore patterns, providing a complete scrape of all text-based files. Explicit --exclude flags are still respected.

Configuration

web2llm uses a hierarchical configuration system that gives you precise control over the scraping process:

  1. Default Config: The tool comes with a built-in default_config.yaml containing a robust set of ignore patterns for common development files and selectors for web scraping.
  2. Project-Specific Config: You can create a .web2llm.yaml file in the root of your project to override or extend the default settings. This is the recommended way to manage project-specific rules.
  3. CLI Arguments: Command-line flags provide the final layer of control, overriding any settings from the configuration files for a single run.

Examples

1. Scrape a specific directory within a GitHub repo:

web2llm 'https://github.com/tiangolo/fastapi' -o fastapi-src --include 'fastapi/'

2. Scrape a local project, excluding test and documentation folders:

web2llm '~/dev/my-project' -o my-project-code --exclude 'tests/' --exclude 'docs/'

3. Scrape a local project but re-include the LICENSE file, which is ignored by default:

web2llm '.' -o my-project-with-license --include '!LICENSE'

4. Scrape everything in a project, including files normally ignored by .gitignore:

web2llm . -o my-project-full --include-all --exclude '.git/'

5. Scrape just the "Installation" section from a webpage:

web2llm 'https://fastapi.tiangolo.com/#installation' -o fastapi-install

6. Scrape a PDF from an arXiv URL:

web2llm 'https://arxiv.org/pdf/1706.03762.pdf' -o attention-is-all-you-need

Contributing

Contributions are welcome. Please refer to the project's issue tracker and CONTRIBUTING.md file for information on how to participate.

License

This project is licensed under the MIT License. See the LICENSE file for details.

Metadata

Release files for web2llm 0.5.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for web2llm 0.5.3
File Size Uploaded
web2llm-0.5.3.tar.gz 25.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for web2llm 0.5.3
File Interpreter ABI Platform
web2llm-0.5.3-py3-none-any.whl Python 3 none any Details

Total release size: 48.8 kB

Release files / web2llm-0.5.3.tar.gz

Download URL web2llm-0.5.3.tar.gz
Size 25.2 kB
Tags Source
SHA-256 checksum
How to use checksums
9131942f873c293e927cb5f97ea9d7e1baa863fd7aeaf1642e815af8744bd00e
BLAKE2b-256 checksum
How to use checksums
56a179297ccdd15b36115f78850358124e9e6be98baa34854f7ae109dd65bdcf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 3, 2025.

Transparency log

Release files / web2llm-0.5.3-py3-none-any.whl

Download URL web2llm-0.5.3-py3-none-any.whl
Size 23.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
09d88546efc3a7ed4334c2ac0e9e55a400dce761edc06bed68974b5f0b908a5f
BLAKE2b-256 checksum
How to use checksums
7fb79280588a794b78b13db7aaea5185b117ca5eabe789bf36d187cdd4ab5e7a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 3, 2025.

Transparency log

Release history Release notifications | RSS feed

This release

0.5.3 This release

2 release files

0.5.1

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page