Skip to main content

a Rust-based web scraping library for Python

Project description

scrapr-rs

scrapr-rs is a library for scraping HTML from the web.


Functions

  • Fetch HTML from an HTTP or HTTPS site
  • Extract specific tags (<ul>, <li>, <div>, etc.)
  • Requests with headers, cookies, and query strings

Examples

use scrapr::{fetch_url, fetch_url_with_options, RequestOptions, extract_tag};

fn main() {
    // Basic GET request
    let html = fetch_url("http://localhost:8000/demo.html");
    println!("Raw HTML:\n{}", html);

    // Extract the first <ul> tag and its contents
    let tag = extract_tag(&html, "ul");
    println!("First <ul> tag:\n{}", tag);

    // Custom headers, cookies, and query parameters
    let opts = RequestOptions {
        headers: Some({
            let mut h = std::collections::HashMap::new();
            h.insert("User-Agent".into(), "scrapr/0.1".into());
            h
        }),
        cookies: Some({
            let mut c = std::collections::HashMap::new();
            c.insert("sessionid".into(), "abc123".into());
            c
        }),
        query: Some({
            let mut q = std::collections::HashMap::new();
            q.insert("q".into(), "Rust programming".into());
            q
        }),
    };

    let response = fetch_url_with_options("https://www.wikipedia.org", opts);
    println!("Wikipedia page HTML:\n{}", response);
}
import scrapr_rs

opts = scrapr.RequestOptions(
    headers={"User-Agent": "XYZ/1.0"},
    cookies={"sessionid": "abc123"},
    query={"q": "Shrek"}
)

text = scrapr.fetch_url_with_options("https://html.duckduckgo.com/html", opts)

print(text)

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scraprr-0.1.4.tar.gz (10.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scraprr-0.1.4-cp313-cp313-manylinux_2_34_x86_64.whl (5.0 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.34+ x86-64

File details

Details for the file scraprr-0.1.4.tar.gz.

File metadata

  • Download URL: scraprr-0.1.4.tar.gz
  • Upload date:
  • Size: 10.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: maturin/1.9.1

File hashes

Hashes for scraprr-0.1.4.tar.gz
Algorithm Hash digest
SHA256 35766c5efb692e9587e6ee68be2e3440d5d61ccc39fb31052f2bcee5f07c0271
MD5 79e867c54a20d873870b58d8b72ef4bf
BLAKE2b-256 d500a05f7853a13ed4f92128ffd711bb9b4954f7e50010669e190f711604e743

See more details on using hashes here.

File details

Details for the file scraprr-0.1.4-cp313-cp313-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for scraprr-0.1.4-cp313-cp313-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 d1ea923f4251be1ee5b9f3e1f74c9c251f23f6e7eaa85f526668047a0eadc82d
MD5 4c328abd6df14f99251d67e44171cbfb
BLAKE2b-256 592baa971448c09e4aee091a01a7e3699bd438f09b70964ba58d49b29da7a4f9

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page