Skip to main content

A pluggable SDK for extracting author information from academic publisher websites

Project description

Crawler SDK

A pluggable SDK for extracting author information from academic publisher websites.

Features

  • Pluggable Architecture: Create custom extraction strategies for any publisher
  • Dual Mode Support: Browser-based (Playwright) and API-based extraction
  • Chain of Responsibility: Flexible email extraction with fallback mechanisms
  • Type Safe: Full type hints support
  • Minimal Dependencies: Core SDK has no required dependencies

Installation

# Core SDK only (no dependencies)
pip install crawler-sdk

# With browser automation support (Playwright)
pip install crawler-sdk[browser]

# With HTTP request support (curl_cffi)
pip install crawler-sdk[http]

# All features
pip install crawler-sdk[all]

Quick Start

Creating a Custom Strategy

from crawler_sdk import ParseStrategy, AuthorInfo
from crawler_sdk.extractors import ExtractorChain, ButtonClickExtractor

class MyPublisherStrategy(ParseStrategy):
    """Custom strategy for extracting authors from MyPublisher"""

    def _build_extractor_chain(self):
        chain = ExtractorChain()
        chain.add(ButtonClickExtractor())
        return chain

    async def extract_authors(self, page):
        # Use the extractor chain
        emails = await self._extract_emails_with_chain(page)

        # Build author list
        authors = []
        for i, email in enumerate(emails, 1):
            authors.append(AuthorInfo(
                author_index=i,
                name=f"Author {i}",
                emails=[email]
            ))
        return authors

    def get_selectors(self):
        return {
            'email_button': 'button.show-email',
            'email_content': '.author-email',
            'close_modal': 'button.close'
        }

Using API-based Extraction

from crawler_sdk.extractors import APIExtractor, APIExtractorChain
from crawler_sdk import AuthorInfo

class MyAPIExtractor(APIExtractor):
    """Extract authors via direct API calls"""

    async def extract(self, article_data):
        doi = article_data.get('doi')
        # Make API request and parse response
        # ...
        return [AuthorInfo(author_index=1, name="John Doe", emails=["john@example.com"])]

# Use the extractor
chain = APIExtractorChain()
chain.add(MyAPIExtractor())
authors = await chain.execute({'doi': '10.1234/example'})

Using Validators

from crawler_sdk import validate_email, validate_doi, validate_url

# Validate email addresses
assert validate_email("user@example.com") == True
assert validate_email("invalid") == False

# Validate DOIs
assert validate_doi("10.1016/j.example.2023.001") == True

# Validate URLs
assert validate_url("https://www.example.com") == True

Using Decorators

from crawler_sdk import retry, rate_limit, timeout

@retry(max_attempts=3, delay=1.0, backoff=2.0)
@rate_limit(min_delay=0.5, max_delay=2.0)
@timeout(seconds=30)
def fetch_article(url):
    # Your fetch logic here
    pass

Data Models

AuthorInfo

from crawler_sdk import AuthorInfo

author = AuthorInfo(
    author_index=1,
    name="John Smith",
    emails=["john.smith@example.edu"],
    affiliations=["Harvard University"],
    confidence=0.95
)

# Access properties
print(author.primary_email)  # "john.smith@example.edu"
print(author.has_email())    # True

ArticleInfo

from crawler_sdk import ArticleInfo

article = ArticleInfo(
    title="Machine Learning in Healthcare",
    doi="10.1016/j.example.2024.001",
    url="https://example.com/article",
    authors_raw=["John Smith", "Jane Doe"],
    publisher="Elsevier",
    year=2024
)

Available Extractors

Browser-based (require Playwright)

  • ButtonClickExtractor - Click buttons to reveal hidden emails
  • DirectTextExtractor - Extract emails from visible page text
  • MetadataExtractor - Extract from meta tags and JSON-LD
  • JavaScriptExtractor - Execute JS to decode obfuscated emails

API-based

  • APIExtractor - Base class for custom API extractors
  • APIExtractorChain - Chain multiple API extractors

Architecture

┌─────────────────────────────────────────────────────────┐
│                    Your Application                      │
├─────────────────────────────────────────────────────────┤
│                     Crawler SDK                          │
│  ┌─────────────┐  ┌─────────────┐  ┌─────────────────┐  │
│  │ Strategies  │  │ Extractors  │  │     Types       │  │
│  │             │  │             │  │                 │  │
│  │ ParseStrategy│ │ EmailExtract│  │ AuthorInfo     │  │
│  │             │  │ ExtractChain│  │ ArticleInfo    │  │
│  └─────────────┘  │ APIExtractor│  │ CrawlResult    │  │
│                   └─────────────┘  └─────────────────┘  │
├─────────────────────────────────────────────────────────┤
│                  Optional Dependencies                   │
│  ┌─────────────┐                 ┌───────────────────┐  │
│  │ Playwright  │                 │    curl_cffi      │  │
│  │ (browser)   │                 │    (API)          │  │
│  └─────────────┘                 └───────────────────┘  │
└─────────────────────────────────────────────────────────┘

License

MIT License

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

crawler_sdk-1.3.0.tar.gz (33.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

crawler_sdk-1.3.0-py3-none-any.whl (46.1 kB view details)

Uploaded Python 3

File details

Details for the file crawler_sdk-1.3.0.tar.gz.

File metadata

  • Download URL: crawler_sdk-1.3.0.tar.gz
  • Upload date:
  • Size: 33.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for crawler_sdk-1.3.0.tar.gz
Algorithm Hash digest
SHA256 224e2cf4fd7fc7b5e809183ea6b1590e5f18c032dee2e61c8795384aa2d19d9d
MD5 872f38a42083a37a80e83da7a9f030d9
BLAKE2b-256 32d111553cc15c8b5ad5aaef646879fb68d637c64d5a6130f7aeb68c822e1c02

See more details on using hashes here.

File details

Details for the file crawler_sdk-1.3.0-py3-none-any.whl.

File metadata

  • Download URL: crawler_sdk-1.3.0-py3-none-any.whl
  • Upload date:
  • Size: 46.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for crawler_sdk-1.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 3b5b6960028cc2919cb65bb76b4871cc44cd5c3d8f2f7193ec52dadfd83bdd41
MD5 31d35ecf9d6238a85f124b27fc1d11fd
BLAKE2b-256 ca36abfe5a79e86f37020afb24db7945dcd90c31339d830ab41e0596e34b4180

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page