Skip to main content

PyPI version Python Versions Ask DeepWiki

scrape cli

It's a command-line tool to extract HTML elements using an XPath query or CSS3 selector.

It's based on the great and simple scraping tool written by Jeroen Janssens.

Installation

You can install scrape-cli using several methods:

Using pipx (recommended for CLI tools)

pipx install scrape-cli

Using uv (modern Python package manager)

# Install as a global CLI tool (recommended)
uv tool install scrape-cli

# Or install with uv pip
uv pip install scrape-cli

# Or run temporarily without installing
uvx --from scrape-cli scrape --help

Using pip

pip install scrape-cli

Or install from source:

git clone https://github.com/aborruso/scrape-cli
cd scrape-cli
pip install -e .

Requirements

  • Python >=3.6
  • requests
  • lxml
  • cssselect

How does it work?

Using the Test HTML File

In the resources directory you'll find a test.html file that you can use to test various scraping scenarios.

Note: You can also test directly from the URL without cloning the repository:

scrape -e "h1" https://raw.githubusercontent.com/aborruso/scrape-cli/refs/heads/master/resources/test.html

Here are some examples:

  1. Extract all table data:
# CSS
scrape -e "table.data-table td" resources/test.html
# XPath
scrape -e "//table[contains(@class, 'data-table')]//td" resources/test.html
  1. Get all list items:
# CSS
scrape -e "ul.items-list li" resources/test.html
# XPath
scrape -e "//ul[contains(@class, 'items-list')]/li" resources/test.html
  1. Extract specific attributes:
# CSS
scrape -e "a.external-link" -a href resources/test.html
# XPath
scrape -e "//a[contains(@class, 'external-link')]/@href" resources/test.html
  1. Check if an element exists:
# CSS
scrape -e "#main-title" --check-existence resources/test.html
# XPath
scrape -e "//h1[@id='main-title']" --check-existence resources/test.html
  1. Extract nested elements:
# CSS
scrape -e ".nested-elements p" resources/test.html
# XPath
scrape -e "//div[contains(@class, 'nested-elements')]//p" resources/test.html
  1. Get elements with specific attributes:
# CSS
scrape -e "[data-test]" resources/test.html
# XPath
scrape -e "//*[@data-test]" resources/test.html
  1. Additional XPath examples:
# Get all links with href attribute
scrape -e "//a[@href]" resources/test.html

# Get checked input elements
scrape -e "//input[@checked]" resources/test.html

# Get elements with multiple classes
scrape -e "//div[contains(@class, 'class1') and contains(@class, 'class2')]" resources/test.html

# Get text content of specific element
scrape -e "//h1[@id='main-title']/text()" resources/test.html

General Usage Examples

A CSS selector query like this

curl -L 'https://en.wikipedia.org/wiki/List_of_sovereign_states' -s \
| scrape -be 'table.wikitable > tbody > tr > td > b > a'

Note: When using both -b and -e options together, they must be specified in the order -be (body first, then expression). Using -eb will not work correctly.

or an XPATH query like this one:

curl -L 'https://en.wikipedia.org/wiki/List_of_sovereign_states' -s \
| scrape -be "//table[contains(@class, 'wikitable')]/tbody/tr/td/b/a"

gives you back:

<html>
 <head>
 </head>
 <body>
  <a href="/wiki/Afghanistan" title="Afghanistan">
   Afghanistan
  </a>
  <a href="/wiki/Albania" title="Albania">
   Albania
  </a>
  <a href="/wiki/Algeria" title="Algeria">
   Algeria
  </a>
  <a href="/wiki/Andorra" title="Andorra">
   Andorra
  </a>
  <a href="/wiki/Angola" title="Angola">
   Angola
  </a>
  <a href="/wiki/Antigua_and_Barbuda" title="Antigua and Barbuda">
   Antigua and Barbuda
  </a>
  <a href="/wiki/Argentina" title="Argentina">
   Argentina
  </a>
  <a href="/wiki/Armenia" title="Armenia">
   Armenia
  </a>
...
...
 </body>
</html>

Text Extraction

You can extract only the text content (without HTML tags) using the -t option, which is particularly useful for LLMs and text processing:

# Extract all text content from a page
curl -L 'https://en.wikipedia.org/wiki/List_of_sovereign_states' -s \
| scrape -t

# Extract text from specific elements
curl -L 'https://en.wikipedia.org/wiki/List_of_sovereign_states' -s \
| scrape -te 'table.wikitable td'

# Extract text from headings only
scrape -te 'h1, h2, h3' resources/test.html

The -t option automatically excludes text from <script> and <style> tags and cleans up whitespace for better readability.

JSON Output

Use the -j/--json flag to get structured JSON natively, with no external tools:

scrape -je "a.external-link" resources/test.html

Output:

{
  "html": {
    "body": {
      "a": {
        "@href": "https://example.com",
        "@class": "external-link",
        "#text": "Example Link"
      }
    }
  }
}

Table extraction example:

scrape -je "table.data-table td" resources/test.html

Output:

{
  "html": {
    "body": {
      "td": [
        "Italy",
        "Rome",
        "59",
        "France",
        "Paris",
        "68"
      ]
    }
  }
}

-j automatically wraps the result in <html>/<body> before conversion, so you don't need to pass -b. The output is the same you would get from scrape -be ... | xq . (the underlying converter is xmltodict, the same library used by xq).

-j is mutually exclusive with -t, -x (--check-existence) and -a (--argument).

Useful for JSON-based pipelines, APIs, databases, and processing with jq/DuckDB.

Some notes on the commands:

  • -e to set the query
  • -b to add <html>, <head> and <body> tags to the HTML output
  • -j to output structured JSON (built-in)
  • -t to extract only text content (useful for LLMs and text processing)

License

MIT

Metadata

Release files for scrape-cli 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scrape-cli 1.3.0
File Size Uploaded
scrape_cli-1.3.0.tar.gz 10.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for scrape-cli 1.3.0
File Interpreter ABI Platform
scrape_cli-1.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 18.7 kB

Release files / scrape_cli-1.3.0.tar.gz

Download URL scrape_cli-1.3.0.tar.gz
Size 10.2 kB
Tags Source
SHA-256 checksum
How to use checksums
0e9da5d97322d804f2921cbebfa3075cb994af41238ea013abdac4f3a2c906ac
BLAKE2b-256 checksum
How to use checksums
72c988cd0ef75d2b6a659a12ceae179b218d57a7623a91fad623c9fc70db6bea
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.7

Release files / scrape_cli-1.3.0-py3-none-any.whl

Download URL scrape_cli-1.3.0-py3-none-any.whl
Size 8.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f7b80df14b73a25c1418c7b0865c859b523e45304a34a722fd0a20ce326fc250
BLAKE2b-256 checksum
How to use checksums
6a9019af82da643ac5de4d9c0ad701b5fc30d2d7c79b440162286051ddee95f5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.7

Release history Release notifications | RSS feed

This release

1.3.0 This release

2 release files

1.2.4

2 release files

1.2.3

2 release files

1.2.2

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.9

2 release files

1.1.8

2 release files

1.1.7

2 release files

1.1.6

2 release files

1.1.5

2 release files

1.1.4

2 release files

1.1.3

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.1

2 release files

0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page