Skip to main content

A pluggable SDK for extracting author information from academic publisher websites

Project description

Crawler SDK

A pluggable SDK for extracting author information from academic publisher websites.

Features

  • Pluggable Architecture: Create custom extraction strategies for any publisher
  • Dual Mode Support: Browser-based (Playwright) and API-based extraction
  • Chain of Responsibility: Flexible email extraction with fallback mechanisms
  • Type Safe: Full type hints support
  • Minimal Dependencies: Core SDK has no required dependencies

Installation

# Core SDK only (no dependencies)
pip install crawler-sdk

# With browser automation support (Playwright)
pip install crawler-sdk[browser]

# With HTTP request support (curl_cffi)
pip install crawler-sdk[http]

# All features
pip install crawler-sdk[all]

Quick Start

Creating a Custom Strategy

from crawler_sdk import ParseStrategy, AuthorInfo
from crawler_sdk.extractors import ExtractorChain, ButtonClickExtractor

class MyPublisherStrategy(ParseStrategy):
    """Custom strategy for extracting authors from MyPublisher"""

    def _build_extractor_chain(self):
        chain = ExtractorChain()
        chain.add(ButtonClickExtractor())
        return chain

    async def extract_authors(self, page):
        # Use the extractor chain
        emails = await self._extract_emails_with_chain(page)

        # Build author list
        authors = []
        for i, email in enumerate(emails, 1):
            authors.append(AuthorInfo(
                author_index=i,
                name=f"Author {i}",
                emails=[email]
            ))
        return authors

    def get_selectors(self):
        return {
            'email_button': 'button.show-email',
            'email_content': '.author-email',
            'close_modal': 'button.close'
        }

Using API-based Extraction

from crawler_sdk.extractors import APIExtractor, APIExtractorChain
from crawler_sdk import AuthorInfo

class MyAPIExtractor(APIExtractor):
    """Extract authors via direct API calls"""

    async def extract(self, article_data):
        doi = article_data.get('doi')
        # Make API request and parse response
        # ...
        return [AuthorInfo(author_index=1, name="John Doe", emails=["john@example.com"])]

# Use the extractor
chain = APIExtractorChain()
chain.add(MyAPIExtractor())
authors = await chain.execute({'doi': '10.1234/example'})

Using Validators

from crawler_sdk import validate_email, validate_doi, validate_url

# Validate email addresses
assert validate_email("user@example.com") == True
assert validate_email("invalid") == False

# Validate DOIs
assert validate_doi("10.1016/j.example.2023.001") == True

# Validate URLs
assert validate_url("https://www.example.com") == True

Using Decorators

from crawler_sdk import retry, rate_limit, timeout

@retry(max_attempts=3, delay=1.0, backoff=2.0)
@rate_limit(min_delay=0.5, max_delay=2.0)
@timeout(seconds=30)
def fetch_article(url):
    # Your fetch logic here
    pass

Data Models

AuthorInfo

from crawler_sdk import AuthorInfo

author = AuthorInfo(
    author_index=1,
    name="John Smith",
    emails=["john.smith@example.edu"],
    affiliations=["Harvard University"],
    confidence=0.95
)

# Access properties
print(author.primary_email)  # "john.smith@example.edu"
print(author.has_email())    # True

ArticleInfo

from crawler_sdk import ArticleInfo

article = ArticleInfo(
    title="Machine Learning in Healthcare",
    doi="10.1016/j.example.2024.001",
    url="https://example.com/article",
    authors_raw=["John Smith", "Jane Doe"],
    publisher="Elsevier",
    year=2024
)

Available Extractors

Browser-based (require Playwright)

  • ButtonClickExtractor - Click buttons to reveal hidden emails
  • DirectTextExtractor - Extract emails from visible page text
  • MetadataExtractor - Extract from meta tags and JSON-LD
  • JavaScriptExtractor - Execute JS to decode obfuscated emails

API-based

  • APIExtractor - Base class for custom API extractors
  • APIExtractorChain - Chain multiple API extractors

Architecture

┌─────────────────────────────────────────────────────────┐
│                    Your Application                      │
├─────────────────────────────────────────────────────────┤
│                     Crawler SDK                          │
│  ┌─────────────┐  ┌─────────────┐  ┌─────────────────┐  │
│  │ Strategies  │  │ Extractors  │  │     Types       │  │
│  │             │  │             │  │                 │  │
│  │ ParseStrategy│ │ EmailExtract│  │ AuthorInfo     │  │
│  │             │  │ ExtractChain│  │ ArticleInfo    │  │
│  └─────────────┘  │ APIExtractor│  │ CrawlResult    │  │
│                   └─────────────┘  └─────────────────┘  │
├─────────────────────────────────────────────────────────┤
│                  Optional Dependencies                   │
│  ┌─────────────┐                 ┌───────────────────┐  │
│  │ Playwright  │                 │    curl_cffi      │  │
│  │ (browser)   │                 │    (API)          │  │
│  └─────────────┘                 └───────────────────┘  │
└─────────────────────────────────────────────────────────┘

License

MIT License

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

crawler_sdk-1.2.0.tar.gz (33.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

crawler_sdk-1.2.0-py3-none-any.whl (46.1 kB view details)

Uploaded Python 3

File details

Details for the file crawler_sdk-1.2.0.tar.gz.

File metadata

  • Download URL: crawler_sdk-1.2.0.tar.gz
  • Upload date:
  • Size: 33.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for crawler_sdk-1.2.0.tar.gz
Algorithm Hash digest
SHA256 113dbc6d3e7cef3719b9bf0c4c84b3675443123ce2e2234dbd91ea7267aa0eb2
MD5 6dcfd93ffaa77416658d7b806e2dfb89
BLAKE2b-256 9e363cd9a8d8a6ffa4498e63ceaf868c165e52ddaf038a3ab0a1a18b6ece3845

See more details on using hashes here.

File details

Details for the file crawler_sdk-1.2.0-py3-none-any.whl.

File metadata

  • Download URL: crawler_sdk-1.2.0-py3-none-any.whl
  • Upload date:
  • Size: 46.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for crawler_sdk-1.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b86a3ad1746110608e1f55aa49169cd8bfa70ed1329ef64a2b4368317686df53
MD5 9292d3a784de33636f379fdbb84fa855
BLAKE2b-256 c9356a64440502bec3d4d5c86ef213ee497a599a611d670d03609f2a7aae8054

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page