Skip to main content

Python SDK for WebCrawler API

Project description

WebCrawler API Python SDK

A Python SDK for interacting with the WebCrawlerAPI.

In order to us API you have to get an API key from WebCrawlerAPI

Installation

pip install webcrawlerapi

Usage

Crawling

from webcrawlerapi import WebCrawlerAPI

# Initialize the client
crawler = WebCrawlerAPI(api_key="your_api_key")

# Synchronous crawling (blocks until completion)
job = crawler.crawl(
    url="https://example.com",
    scrape_type="markdown",
    items_limit=10,
    webhook_url="https://yourserver.com/webhook",
    max_polls=100  # Optional: maximum number of status checks. Use higher for bigger websites
 )

print(f"Job completed with status: {job.status}")

# Access job items and their content
for item in job.job_items:
    print(f"Page title: {item.title}")
    print(f"Original URL: {item.original_url}")
    print(f"Item status: {item.status}")
    
    # Get the content based on job's scrape_type
    # Returns None if item is not in "done" status
    content = item.content
    if content:
        print(f"Content length: {len(content)}")
        print(f"Content preview: {content[:200]}...")
    else:
        print("Content not available or item not done")

# Access job items and their parent job
for item in job.job_items:
    print(f"Item URL: {item.original_url}")
    print(f"Parent job status: {item.job.status}")
    print(f"Parent job URL: {item.job.url}")

# Or use asynchronous crawling
response = crawler.crawl_async(
    url="https://example.com",
    scrape_type="markdown",
    items_limit=10,
    webhook_url="https://yourserver.com/webhook"
)

# Get the job ID from the response
job_id = response.id
print(f"Crawling job started with ID: {job_id}")

# Check job status and get results
job = crawler.get_job(job_id)
print(f"Job status: {job.status}")

# Access job details
print(f"Crawled URL: {job.url}")
print(f"Created at: {job.created_at}")
print(f"Number of items: {len(job.job_items)}")

# Cancel a running job if needed
cancel_response = crawler.cancel_job(job_id)
print(f"Cancellation response: {cancel_response['message']}")

Scraping

Check a working code example of scraping and scraping with a prompt

# Returns structured data directly
response = crawler.scrape(
    url="https://webcrawlerapi.com"
)
if response.success:
    print(response.markdown)
else:
    print(f"Code: {response.error_code} Error: {response.error_message}")

API Methods

crawl()

Starts a new crawling job and waits for its completion. This method will continuously poll the job status until:

  • The job reaches a terminal state (done, error, or cancelled)
  • The maximum number of polls is reached (default: 100)
  • The polling interval is determined by the server's recommended_pull_delay_ms or defaults to 5 seconds

crawl_async()

Starts a new crawling job and returns immediately with a job ID. Use this when you want to handle polling and status checks yourself, or when using webhooks.

Use crawl_raw_markdown() when you need the combined /job/{id}/markdown output after a crawl finishes.

get_job()

Retrieves the current status and details of a specific job.

cancel_job()

Cancels a running job. Any items that are not in progress or already completed will be marked as canceled and will not be charged.

scrape()

Scrapes a single URL and returns the markdown, cleaned or raw content, page status code and page title.

Scrape Params

Read more in API Docs

  • url (required): The URL to scrape.
  • output_format (required): The format of the output. Can be "markdown", "cleaned" or "raw"

Parameters

Crawl Methods (crawl and crawl_async)

  • url (required): The seed URL where the crawler starts. Can be any valid URL.
  • scrape_type (default: "html"): The type of scraping you want to perform. Can be "html", "cleaned", or "markdown".
  • items_limit (default: 10): Crawler will stop when it reaches this limit of pages for this job.
  • webhook_url (optional): The URL where the server will send a POST request once the task is completed.
  • whitelist_regexp (optional): A regular expression to whitelist URLs. Only URLs that match the pattern will be crawled.
  • blacklist_regexp (optional): A regular expression to blacklist URLs. URLs that match the pattern will be skipped.
  • max_polls (optional, crawl only): Maximum number of status checks before returning (default: 100)

Responses

CrawlAsync Response

The crawl_async() method returns a CrawlResponse object with:

  • id: The unique identifier of the created job

Job Response

The Job object contains detailed information about the crawling job:

  • id: The unique identifier of the job
  • org_id: Your organization identifier
  • url: The seed URL where the crawler started
  • status: The status of the job (new, in_progress, done, error)
  • scrape_type: The type of scraping performed
  • created_at: The date when the job was created
  • finished_at: The date when the job was finished (if completed)
  • webhook_url: The webhook URL for notifications
  • webhook_status: The status of the webhook request
  • webhook_error: Any error message if the webhook request failed
  • job_items: List of JobItem objects representing crawled pages
  • recommended_pull_delay_ms: Server-recommended delay between status checks

JobItem Properties

Each JobItem object represents a crawled page and contains:

  • id: The unique identifier of the item
  • job_id: The parent job identifier
  • job: Reference to the parent Job object
  • original_url: The URL of the page
  • page_status_code: The HTTP status code of the page request
  • status: The status of the item (new, in_progress, done, error)
  • title: The page title
  • created_at: The date when the item was created
  • cost: The cost of the item in $
  • referred_url: The URL where the page was referred from
  • last_error: Any error message if the item failed
  • error_code: The error code if the item failed (if available)
  • content: The page content based on the job's scrape_type (html, cleaned, or markdown). Returns None if the item's status is not "done" or if content is not available. Content is automatically fetched and cached when accessed.
  • raw_content_url: URL to the raw content (if available)
  • cleaned_content_url: URL to the cleaned content (if scrape_type is "cleaned")
  • markdown_content_url: URL to the markdown content (if scrape_type is "markdown")

Requirements

  • Python 3.6+
  • requests>=2.25.0

License

MIT License

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

webcrawlerapi-2.0.11.tar.gz (16.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

webcrawlerapi-2.0.11-py3-none-any.whl (15.9 kB view details)

Uploaded Python 3

File details

Details for the file webcrawlerapi-2.0.11.tar.gz.

File metadata

  • Download URL: webcrawlerapi-2.0.11.tar.gz
  • Upload date:
  • Size: 16.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.2

File hashes

Hashes for webcrawlerapi-2.0.11.tar.gz
Algorithm Hash digest
SHA256 e5164a80082e876c500ebbdc63a9a8a0a5bd4cfe4f790b439ea9c7d1ad6d95e3
MD5 fa12e0ab828536c7de14803fe44d74c3
BLAKE2b-256 c3c850201fb641955564c75080eae9af9c180b758249fce9215db5860cae5ab3

See more details on using hashes here.

File details

Details for the file webcrawlerapi-2.0.11-py3-none-any.whl.

File metadata

  • Download URL: webcrawlerapi-2.0.11-py3-none-any.whl
  • Upload date:
  • Size: 15.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.2

File hashes

Hashes for webcrawlerapi-2.0.11-py3-none-any.whl
Algorithm Hash digest
SHA256 1f1bf04d4a8e31df3768f9b5e4b7d52700d44419b75ec0892a6f9b9d1a412d17
MD5 d9156e2148649419d7544f2d6b3ec5df
BLAKE2b-256 ad984fccc539c1f4ee3d75deabc33cde99b9df3080860b4635fc47563a546909

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page