Skip to main content

ScrapeGraph Python SDK for API

Project description

🌐 ScrapeGraph Python SDK

PyPI version Python Support License Code style: black Documentation Status

Official Python SDK for the ScrapeGraph AI API - Smart web scraping powered by AI.

🚀 Features

  • ✨ Smart web scraping with AI
  • 🔄 Both sync and async clients
  • 📊 Structured output with Pydantic schemas
  • 🔍 Detailed logging with emojis
  • ⚡ Automatic retries and error handling
  • 🔐 Secure API authentication

📦 Installation

Using pip

pip install scrapegraph-py

Using uv

We recommend using uv to install the dependencies and pre-commit hooks.

# Install uv if you haven't already
pip install uv

# Install dependencies
uv sync

# Install pre-commit hooks
uv run pre-commit install

🔧 Quick Start

[!NOTE] If you prefer, you can use the environment variables to configure the API key and load them using load_dotenv()

from scrapegraph_py import SyncClient
from scrapegraph_py.logger import get_logger

# Enable debug logging
logger = get_logger(level="DEBUG")

# Initialize client
sgai_client = SyncClient(api_key="your-api-key-here")

# Make a request
response = sgai_client.smartscraper(
    website_url="https://example.com",
    user_prompt="Extract the main heading and description"
)

print(response["result"])

🎯 Examples

Async Usage

import asyncio
from scrapegraph_py import AsyncClient

async def main():
    async with AsyncClient(api_key="your-api-key-here") as sgai_client:
        response = await sgai_client.smartscraper(
            website_url="https://example.com",
            user_prompt="Summarize the main content"
        )
        print(response["result"])

asyncio.run(main())
With Output Schema
from pydantic import BaseModel, Field
from scrapegraph_py import SyncClient

class WebsiteData(BaseModel):
    title: str = Field(description="The page title")
    description: str = Field(description="The meta description")

sgai_client = SyncClient(api_key="your-api-key-here")
response = sgai_client.smartscraper(
    website_url="https://example.com",
    user_prompt="Extract the title and description",
    output_schema=WebsiteData
)

print(response["result"])

📚 Documentation

For detailed documentation, visit docs.scrapegraphai.com

🛠️ Development

Setup

  1. Clone the repository:
git clone https://github.com/ScrapeGraphAI/scrapegraph-sdk.git
cd scrapegraph-sdk/scrapegraph-py
  1. Install dependencies:
uv sync
  1. Install pre-commit hooks:
uv run pre-commit install

Running Tests

# Run all tests
uv run pytest

# Run specific test file
poetry run pytest tests/test_client.py

📝 License

This project is licensed under the MIT License - see the LICENSE file for details.

🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request. For major changes, please open an issue first to discuss what you would like to change.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/AmazingFeature)
  3. Commit your changes (git commit -m 'Add some AmazingFeature')
  4. Push to the branch (git push origin feature/AmazingFeature)
  5. Open a Pull Request

🔗 Links

💬 Support


Made with ❤️ by ScrapeGraph AI

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scrapegraph_py-1.4.1.tar.gz (105.2 kB view details)

Uploaded Source

Built Distribution

scrapegraph_py-1.4.1-py3-none-any.whl (11.9 kB view details)

Uploaded Python 3

File details

Details for the file scrapegraph_py-1.4.1.tar.gz.

File metadata

  • Download URL: scrapegraph_py-1.4.1.tar.gz
  • Upload date:
  • Size: 105.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.10.12

File hashes

Hashes for scrapegraph_py-1.4.1.tar.gz
Algorithm Hash digest
SHA256 1c122042ece4b3527bfb83e06bcd548cedd96e7ecbca8eff9aa1d550f6f33eba
MD5 a409ee52372f86fd8dbc0c0fa779d91e
BLAKE2b-256 b671911505c0b1722a46e0de4ea344134ecf14c124562a8f3167c74c1fa46827

See more details on using hashes here.

File details

Details for the file scrapegraph_py-1.4.1-py3-none-any.whl.

File metadata

File hashes

Hashes for scrapegraph_py-1.4.1-py3-none-any.whl
Algorithm Hash digest
SHA256 ceddf24b1a2813239207e6ac65edb87c4437a2e48601bd06b4b04a65e90c0c20
MD5 715039537c65c1d4a5e25e5d01c29611
BLAKE2b-256 3706c48f5861709325b690ebc22007e9c2e82891769aac9e836ec8d07468bd7b

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page