Skip to main content

a Rust-based web scraping library for Python

Project description

scrapr-rs

scrapr-rs is a library for scraping HTML from the web.


Functions

  • Fetch HTML from an HTTP or HTTPS site
  • Extract specific tags (<ul>, <li>, <div>, etc.)
  • Requests with headers, cookies, and query strings

Examples

use scrapr::{fetch_url, fetch_url_with_options, RequestOptions, extract_tag};

fn main() {
    // Basic GET request
    let html = fetch_url("http://localhost:8000/demo.html");
    println!("Raw HTML:\n{}", html);

    // Extract the first <ul> tag and its contents
    let tag = extract_tag(&html, "ul");
    println!("First <ul> tag:\n{}", tag);

    // Custom headers, cookies, and query parameters
    let opts = RequestOptions {
        headers: Some({
            let mut h = std::collections::HashMap::new();
            h.insert("User-Agent".into(), "scrapr/0.1".into());
            h
        }),
        cookies: Some({
            let mut c = std::collections::HashMap::new();
            c.insert("sessionid".into(), "abc123".into());
            c
        }),
        query: Some({
            let mut q = std::collections::HashMap::new();
            q.insert("q".into(), "Rust programming".into());
            q
        }),
    };

    let response = fetch_url_with_options("https://www.wikipedia.org", opts);
    println!("Wikipedia page HTML:\n{}", response);
}
import scrapr_rs

opts = scrapr.RequestOptions(
    headers={"User-Agent": "XYZ/1.0"},
    cookies={"sessionid": "abc123"},
    query={"q": "Shrek"}
)

text = scrapr.fetch_url_with_options("https://html.duckduckgo.com/html", opts)

print(text)

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scraprr-0.1.7.tar.gz (10.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scraprr-0.1.7-cp313-cp313-manylinux_2_34_x86_64.whl (5.0 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.34+ x86-64

File details

Details for the file scraprr-0.1.7.tar.gz.

File metadata

  • Download URL: scraprr-0.1.7.tar.gz
  • Upload date:
  • Size: 10.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: maturin/1.9.1

File hashes

Hashes for scraprr-0.1.7.tar.gz
Algorithm Hash digest
SHA256 4b91aa389d1805709c5ed78313fef728c0ea5730fd2dd0370689aa0765595aae
MD5 4faf393d664f2a5dc24d9aa03ce2ed1d
BLAKE2b-256 7b1825c55c4a31735a0f3ae9754218729cd2e29914212ebc0628e3ec30be5ef0

See more details on using hashes here.

File details

Details for the file scraprr-0.1.7-cp313-cp313-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for scraprr-0.1.7-cp313-cp313-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 76187e9bf7c33e021885307dbacab5ed31b49ff72b9b89ab04c684697cb0a7ad
MD5 1c4dc18c6a7bb6c9ffee882593caee0a
BLAKE2b-256 42dde31c99a0b345dce0630697ef3c491c93f1fc2b998232bfc6a271aee16452

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page