Skip to main content

a Rust-based web scraping library for Python

Project description

scrapr-rs

scrapr-rs is a library for scraping HTML from the web.


Functions

  • Fetch HTML from an HTTP or HTTPS site
  • Extract specific tags (<ul>, <li>, <div>, etc.)
  • Requests with headers, cookies, and query strings

Examples

use scrapr::{fetch_url, fetch_url_with_options, RequestOptions, extract_tag};

fn main() {
    // Basic GET request
    let html = fetch_url("http://localhost:8000/demo.html");
    println!("Raw HTML:\n{}", html);

    // Extract the first <ul> tag and its contents
    let tag = extract_tag(&html, "ul");
    println!("First <ul> tag:\n{}", tag);

    // Custom headers, cookies, and query parameters
    let opts = RequestOptions {
        headers: Some({
            let mut h = std::collections::HashMap::new();
            h.insert("User-Agent".into(), "scrapr/0.1".into());
            h
        }),
        cookies: Some({
            let mut c = std::collections::HashMap::new();
            c.insert("sessionid".into(), "abc123".into());
            c
        }),
        query: Some({
            let mut q = std::collections::HashMap::new();
            q.insert("q".into(), "Rust programming".into());
            q
        }),
    };

    let response = fetch_url_with_options("https://www.wikipedia.org", opts);
    println!("Wikipedia page HTML:\n{}", response);
}
import scrapr_rs

opts = scrapr.RequestOptions(
    headers={"User-Agent": "XYZ/1.0"},
    cookies={"sessionid": "abc123"},
    query={"q": "Shrek"}
)

text = scrapr.fetch_url_with_options("https://html.duckduckgo.com/html", opts)

print(text)

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scraprr-0.1.12.tar.gz (10.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scraprr-0.1.12-cp313-cp313-macosx_11_0_arm64.whl (241.6 kB view details)

Uploaded CPython 3.13macOS 11.0+ ARM64

File details

Details for the file scraprr-0.1.12.tar.gz.

File metadata

  • Download URL: scraprr-0.1.12.tar.gz
  • Upload date:
  • Size: 10.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: maturin/1.9.1

File hashes

Hashes for scraprr-0.1.12.tar.gz
Algorithm Hash digest
SHA256 2d529b6d017dcb623fdf405010083e9295a3f9aea4652372be25a4a31ce34abc
MD5 c8775c5deaa075ced722d99b01327883
BLAKE2b-256 bbee2a37bafe5a26a52870a5644127b69880139725d019ee1580340dd65d7727

See more details on using hashes here.

File details

Details for the file scraprr-0.1.12-cp313-cp313-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for scraprr-0.1.12-cp313-cp313-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 91296a77b80ff577609ed289f21ac561b079b28472aedc689035d2fbd3ee6d54
MD5 6e2cc97306d467c2870b50008591a054
BLAKE2b-256 f32db5e29aab89c69bee51d18638b606f95d328bfe362803733334f9b96a2763

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page