Skip to main content

A pluggable SDK for extracting author information from academic publisher websites

Project description

Crawler SDK

A pluggable SDK for extracting author information from academic publisher websites.

Features

  • Pluggable Architecture: Create custom extraction strategies for any publisher
  • Dual Mode Support: Browser-based (Playwright) and API-based extraction
  • Chain of Responsibility: Flexible email extraction with fallback mechanisms
  • Type Safe: Full type hints support
  • Minimal Dependencies: Core SDK has no required dependencies

Installation

# Core SDK only (no dependencies)
pip install crawler-sdk

# With browser automation support (Playwright)
pip install crawler-sdk[browser]

# With HTTP request support (curl_cffi)
pip install crawler-sdk[http]

# All features
pip install crawler-sdk[all]

Quick Start

Creating a Custom Strategy

from crawler_sdk import ParseStrategy, AuthorInfo
from crawler_sdk.extractors import ExtractorChain, ButtonClickExtractor

class MyPublisherStrategy(ParseStrategy):
    """Custom strategy for extracting authors from MyPublisher"""

    def _build_extractor_chain(self):
        chain = ExtractorChain()
        chain.add(ButtonClickExtractor())
        return chain

    async def extract_authors(self, page):
        # Use the extractor chain
        emails = await self._extract_emails_with_chain(page)

        # Build author list
        authors = []
        for i, email in enumerate(emails, 1):
            authors.append(AuthorInfo(
                author_index=i,
                name=f"Author {i}",
                emails=[email]
            ))
        return authors

    def get_selectors(self):
        return {
            'email_button': 'button.show-email',
            'email_content': '.author-email',
            'close_modal': 'button.close'
        }

Using API-based Extraction

from crawler_sdk.extractors import APIExtractor, APIExtractorChain
from crawler_sdk import AuthorInfo

class MyAPIExtractor(APIExtractor):
    """Extract authors via direct API calls"""

    async def extract(self, article_data):
        doi = article_data.get('doi')
        # Make API request and parse response
        # ...
        return [AuthorInfo(author_index=1, name="John Doe", emails=["john@example.com"])]

# Use the extractor
chain = APIExtractorChain()
chain.add(MyAPIExtractor())
authors = await chain.execute({'doi': '10.1234/example'})

Using Validators

from crawler_sdk import validate_email, validate_doi, validate_url

# Validate email addresses
assert validate_email("user@example.com") == True
assert validate_email("invalid") == False

# Validate DOIs
assert validate_doi("10.1016/j.example.2023.001") == True

# Validate URLs
assert validate_url("https://www.example.com") == True

Using Decorators

from crawler_sdk import retry, rate_limit, timeout

@retry(max_attempts=3, delay=1.0, backoff=2.0)
@rate_limit(min_delay=0.5, max_delay=2.0)
@timeout(seconds=30)
def fetch_article(url):
    # Your fetch logic here
    pass

Data Models

AuthorInfo

from crawler_sdk import AuthorInfo

author = AuthorInfo(
    author_index=1,
    name="John Smith",
    emails=["john.smith@example.edu"],
    affiliations=["Harvard University"],
    confidence=0.95
)

# Access properties
print(author.primary_email)  # "john.smith@example.edu"
print(author.has_email())    # True

ArticleInfo

from crawler_sdk import ArticleInfo

article = ArticleInfo(
    title="Machine Learning in Healthcare",
    doi="10.1016/j.example.2024.001",
    url="https://example.com/article",
    authors_raw=["John Smith", "Jane Doe"],
    publisher="Elsevier",
    year=2024
)

Available Extractors

Browser-based (require Playwright)

  • ButtonClickExtractor - Click buttons to reveal hidden emails
  • DirectTextExtractor - Extract emails from visible page text
  • MetadataExtractor - Extract from meta tags and JSON-LD
  • JavaScriptExtractor - Execute JS to decode obfuscated emails

API-based

  • APIExtractor - Base class for custom API extractors
  • APIExtractorChain - Chain multiple API extractors

Architecture

┌─────────────────────────────────────────────────────────┐
│                    Your Application                      │
├─────────────────────────────────────────────────────────┤
│                     Crawler SDK                          │
│  ┌─────────────┐  ┌─────────────┐  ┌─────────────────┐  │
│  │ Strategies  │  │ Extractors  │  │     Types       │  │
│  │             │  │             │  │                 │  │
│  │ ParseStrategy│ │ EmailExtract│  │ AuthorInfo     │  │
│  │             │  │ ExtractChain│  │ ArticleInfo    │  │
│  └─────────────┘  │ APIExtractor│  │ CrawlResult    │  │
│                   └─────────────┘  └─────────────────┘  │
├─────────────────────────────────────────────────────────┤
│                  Optional Dependencies                   │
│  ┌─────────────┐                 ┌───────────────────┐  │
│  │ Playwright  │                 │    curl_cffi      │  │
│  │ (browser)   │                 │    (API)          │  │
│  └─────────────┘                 └───────────────────┘  │
└─────────────────────────────────────────────────────────┘

License

MIT License

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

crawler_sdk-1.0.1.tar.gz (23.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

crawler_sdk-1.0.1-py3-none-any.whl (33.5 kB view details)

Uploaded Python 3

File details

Details for the file crawler_sdk-1.0.1.tar.gz.

File metadata

  • Download URL: crawler_sdk-1.0.1.tar.gz
  • Upload date:
  • Size: 23.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for crawler_sdk-1.0.1.tar.gz
Algorithm Hash digest
SHA256 36779f1e54f41054a7bed80636c4c69e3cc79735128be299b2de5d7e954965d5
MD5 8698c90cd3b58812afb847cad3758647
BLAKE2b-256 10b93dc331b2054eb5059aa7593cd0c48a39b53e7671465928c5189e72212adb

See more details on using hashes here.

File details

Details for the file crawler_sdk-1.0.1-py3-none-any.whl.

File metadata

  • Download URL: crawler_sdk-1.0.1-py3-none-any.whl
  • Upload date:
  • Size: 33.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for crawler_sdk-1.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 f665e916cb2bba4c7363d7832f83efb03d68114c6be6b058d0de07c3bc75a43d
MD5 a383cccc13054d7c663d47f4873441ce
BLAKE2b-256 669fe42b3f67098ebfcd7d9db2e188617f60bcc2da29e630d0de1c77fe095a99

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page