Skip to main content

A pluggable SDK for extracting author information from academic publisher websites

Project description

Crawler SDK

A pluggable SDK for extracting author information from academic publisher websites.

Features

  • Pluggable Architecture: Create custom extraction strategies for any publisher
  • Dual Mode Support: Browser-based (Playwright) and API-based extraction
  • Chain of Responsibility: Flexible email extraction with fallback mechanisms
  • Type Safe: Full type hints support
  • Minimal Dependencies: Core SDK has no required dependencies

Installation

# Core SDK only (no dependencies)
pip install crawler-sdk

# With browser automation support (Playwright)
pip install crawler-sdk[browser]

# With HTTP request support (curl_cffi)
pip install crawler-sdk[http]

# All features
pip install crawler-sdk[all]

Quick Start

Creating a Custom Strategy

from crawler_sdk import ParseStrategy, AuthorInfo
from crawler_sdk.extractors import ExtractorChain, ButtonClickExtractor

class MyPublisherStrategy(ParseStrategy):
    """Custom strategy for extracting authors from MyPublisher"""

    def _build_extractor_chain(self):
        chain = ExtractorChain()
        chain.add(ButtonClickExtractor())
        return chain

    async def extract_authors(self, page):
        # Use the extractor chain
        emails = await self._extract_emails_with_chain(page)

        # Build author list
        authors = []
        for i, email in enumerate(emails, 1):
            authors.append(AuthorInfo(
                author_index=i,
                name=f"Author {i}",
                emails=[email]
            ))
        return authors

    def get_selectors(self):
        return {
            'email_button': 'button.show-email',
            'email_content': '.author-email',
            'close_modal': 'button.close'
        }

Using API-based Extraction

from crawler_sdk.extractors import APIExtractor, APIExtractorChain
from crawler_sdk import AuthorInfo

class MyAPIExtractor(APIExtractor):
    """Extract authors via direct API calls"""

    async def extract(self, article_data):
        doi = article_data.get('doi')
        # Make API request and parse response
        # ...
        return [AuthorInfo(author_index=1, name="John Doe", emails=["john@example.com"])]

# Use the extractor
chain = APIExtractorChain()
chain.add(MyAPIExtractor())
authors = await chain.execute({'doi': '10.1234/example'})

Using Validators

from crawler_sdk import validate_email, validate_doi, validate_url

# Validate email addresses
assert validate_email("user@example.com") == True
assert validate_email("invalid") == False

# Validate DOIs
assert validate_doi("10.1016/j.example.2023.001") == True

# Validate URLs
assert validate_url("https://www.example.com") == True

Using Decorators

from crawler_sdk import retry, rate_limit, timeout

@retry(max_attempts=3, delay=1.0, backoff=2.0)
@rate_limit(min_delay=0.5, max_delay=2.0)
@timeout(seconds=30)
def fetch_article(url):
    # Your fetch logic here
    pass

Data Models

AuthorInfo

from crawler_sdk import AuthorInfo

author = AuthorInfo(
    author_index=1,
    name="John Smith",
    emails=["john.smith@example.edu"],
    affiliations=["Harvard University"],
    confidence=0.95
)

# Access properties
print(author.primary_email)  # "john.smith@example.edu"
print(author.has_email())    # True

ArticleInfo

from crawler_sdk import ArticleInfo

article = ArticleInfo(
    title="Machine Learning in Healthcare",
    doi="10.1016/j.example.2024.001",
    url="https://example.com/article",
    authors_raw=["John Smith", "Jane Doe"],
    publisher="Elsevier",
    year=2024
)

Available Extractors

Browser-based (require Playwright)

  • ButtonClickExtractor - Click buttons to reveal hidden emails
  • DirectTextExtractor - Extract emails from visible page text
  • MetadataExtractor - Extract from meta tags and JSON-LD
  • JavaScriptExtractor - Execute JS to decode obfuscated emails

API-based

  • APIExtractor - Base class for custom API extractors
  • APIExtractorChain - Chain multiple API extractors

Architecture

┌─────────────────────────────────────────────────────────┐
│                    Your Application                      │
├─────────────────────────────────────────────────────────┤
│                     Crawler SDK                          │
│  ┌─────────────┐  ┌─────────────┐  ┌─────────────────┐  │
│  │ Strategies  │  │ Extractors  │  │     Types       │  │
│  │             │  │             │  │                 │  │
│  │ ParseStrategy│ │ EmailExtract│  │ AuthorInfo     │  │
│  │             │  │ ExtractChain│  │ ArticleInfo    │  │
│  └─────────────┘  │ APIExtractor│  │ CrawlResult    │  │
│                   └─────────────┘  └─────────────────┘  │
├─────────────────────────────────────────────────────────┤
│                  Optional Dependencies                   │
│  ┌─────────────┐                 ┌───────────────────┐  │
│  │ Playwright  │                 │    curl_cffi      │  │
│  │ (browser)   │                 │    (API)          │  │
│  └─────────────┘                 └───────────────────┘  │
└─────────────────────────────────────────────────────────┘

License

MIT License

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

crawler_sdk-1.0.0.tar.gz (23.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

crawler_sdk-1.0.0-py3-none-any.whl (33.5 kB view details)

Uploaded Python 3

File details

Details for the file crawler_sdk-1.0.0.tar.gz.

File metadata

  • Download URL: crawler_sdk-1.0.0.tar.gz
  • Upload date:
  • Size: 23.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for crawler_sdk-1.0.0.tar.gz
Algorithm Hash digest
SHA256 679c9ec412165cba2e5397dc8bf71dc1119c5dbb9a5e0b42b7a763a2b59e5743
MD5 e633665b41d5995abc59424df7bdd8b3
BLAKE2b-256 264c1f38a20fa86f4fe9326bf22be07730eb5bd920828ef2077b35384856362d

See more details on using hashes here.

File details

Details for the file crawler_sdk-1.0.0-py3-none-any.whl.

File metadata

  • Download URL: crawler_sdk-1.0.0-py3-none-any.whl
  • Upload date:
  • Size: 33.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for crawler_sdk-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e4730f7a9563197c47dd1ae7bcd61d98af7f90789b2c58860922c8e3455e43f4
MD5 e447c2f91074559bfd56001cbc0c2cc9
BLAKE2b-256 3b0abe06698bf6c15d27cfefdb0ed51701aa4163e9c2cbc904494c64d38335ab

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page