Skip to main content
License License Version Version
Github Actions Github Actions Coverage CodeCov
Supported versions Python Versions Wheel Wheel
Status Status Downloads Downloads
All Contributors All Contributors

dude uncomplicated data extraction

Dude is a very simple framework for writing web scrapers using Python decorators. The design, inspired by Flask, was to easily build a web scraper in just a few lines of code. Dude has an easy-to-learn syntax.

🚨 Dude is currently in Pre-Alpha. Please expect breaking changes.

Installation

To install, simply run the following from terminal.

pip install pydude
playwright install  # Install playwright binaries for Chrome, Firefox and Webkit.

Minimal web scraper

The simplest web scraper will look like this:

from dude import select


@select(css="a")
def get_link(element):
    return {"url": element.get_attribute("href")}

The example above will get all the hyperlink elements in a page and calls the handler function get_link() for each element.

How to run the scraper

You can run your scraper from terminal/shell/command-line by supplying URLs, the output filename of your choice and the paths to your python scripts to dude scrape command.

dude scrape --url "<url>" --output data.json path/to/script.py

The output in data.json should contain the actual URL and the metadata prepended with underscore.

[
  {
    "_page_number": 1,
    "_page_url": "https://dude.ron.sh/",
    "_group_id": 4502003824,
    "_group_index": 0,
    "_element_index": 0,
    "url": "/url-1.html"
  },
  {
    "_page_number": 1,
    "_page_url": "https://dude.ron.sh/",
    "_group_id": 4502003824,
    "_group_index": 0,
    "_element_index": 1,
    "url": "/url-2.html"
  },
  {
    "_page_number": 1,
    "_page_url": "https://dude.ron.sh/",
    "_group_id": 4502003824,
    "_group_index": 0,
    "_element_index": 2,
    "url": "/url-3.html"
  }
]

Changing the output to --output data.csv should result in the following CSV content.

data.csv

Features

  • Simple Flask-inspired design - build a scraper with decorators.
  • Uses Playwright API - run your scraper in Chrome, Firefox and Webkit and leverage Playwright's powerful selector engine supporting CSS, XPath, text, regex, etc.
  • Data grouping - group related results.
  • URL pattern matching - run functions on matched URLs.
  • Priority - reorder functions based on priority.
  • Setup function - enable setup steps (clicking dialogs or login).
  • Navigate function - enable navigation steps to move to other pages.
  • Custom storage - option to save data to other formats or database.
  • Async support - write async handlers.
  • Option to use other parser backends aside from Playwright.
  • Option to follow all links indefinitely (Crawler/Spider).
  • Events - attach functions to startup, pre-setup, post-setup and shutdown events.
  • Option to save data on every page.

Supported Parser Backends

By default, Dude uses Playwright but gives you an option to use parser backends that you are familiar with. It is possible to use parser backends like BeautifulSoup4, Parsel, lxml, and Selenium.

Here is the summary of features supported by each parser backend.

Parser Backend Supports
Sync?
Supports
Async?
Selectors Setup
Handler
Navigate
Handler
Comments
CSS XPath Text Regex
Playwright ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
BeautifulSoup4 ✅ ✅ ✅ 🚫 🚫 🚫 🚫 🚫
Parsel ✅ ✅ ✅ ✅ ✅ ✅ 🚫 🚫
lxml ✅ ✅ ✅ ✅ ✅ ✅ 🚫 🚫
Pyppeteer 🚫 ✅ ✅ ✅ ✅ 🚫 ✅ ✅ Not supported from 0.23.0
Selenium ✅ ✅ ✅ ✅ ✅ 🚫 ✅ ✅

Using the Docker image

Pull the docker image using the following command.

docker pull roniemartinez/dude

Assuming that script.py exist in the current directory, run Dude using the following command.

docker run -it --rm -v "$PWD":/code roniemartinez/dude dude scrape --url <url> script.py

Documentation

Read the complete documentation at https://roniemartinez.github.io/dude/. All the advanced and useful features are documented there.

Requirements

  • ✅ Any dude should know how to work with selectors (CSS or XPath).
  • ✅ Familiarity with any backends that you love (see Supported Parser Backends)
  • ✅ Python decorators... you'll live, dude!

Why name this project "dude"?

  • ✅ A Recursive acronym looks nice.
  • ✅ Adding "uncomplicated" (like ufw) into the name says it is a very simple framework.
  • ✅ Puns! I also think that if you want to do web scraping, there's probably some random dude around the corner who can make it very easy for you to start with it. 😊

Author

Ronie Martinez

Contributors ✨

Thanks goes to these wonderful people (emoji key):


Ronie Martinez

🚧 💻 📖 🚇

This project follows the all-contributors specification. Contributions of any kind welcome!

Metadata

Release files for pydude 0.28.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pydude 0.28.0
File Size Uploaded
pydude-0.28.0.tar.gz 33.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pydude 0.28.0
File Interpreter ABI Platform
pydude-0.28.0-py3-none-any.whl Python 3 none any Details

Total release size: 74.4 kB

Release files / pydude-0.28.0.tar.gz

Download URL pydude-0.28.0.tar.gz
Size 33.4 kB
Tags Source
SHA-256 checksum
How to use checksums
bfb175902c8580ae83ffa8c92ddefebda58287b5f4ff05f9c3c80e92ab05c4df
BLAKE2b-256 checksum
How to use checksums
8fff892125ab96c970c9a841c783d35a6e287f9db895c290a34ea3a585530aa9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.8.1 CPython/3.11.8 Linux/6.5.0-1015-azure

Release files / pydude-0.28.0-py3-none-any.whl

Download URL pydude-0.28.0-py3-none-any.whl
Size 41.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b37e9362b3da8dddd5f12230a43177c23f5e7fc28771e1ed314c41b8e6f83dae
BLAKE2b-256 checksum
How to use checksums
0c353658f627d850ab051290e2e203164c69f8ac1d2b6b8385f81cf11cbc40a2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.8.1 CPython/3.11.8 Linux/6.5.0-1015-azure

Release history Release notifications | RSS feed

This release

0.28.0 This release

2 release files

0.27.0

2 release files

0.26.0

2 release files

0.25.1

2 release files

0.25.0

2 release files

0.20.2

2 release files

0.20.1

2 release files

0.20.0

2 release files

0.18.0

2 release files

0.15.1

2 release files

0.15.0

2 release files

0.14.0

2 release files

0.13.0

2 release files

0.12.2

2 release files

0.12.1

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page