aioscraper
High-performance asynchronous Python framework for large-scale API data collection.
API-first. aioscraper orchestrates thousands of concurrent JSON/REST calls with adaptive rate limiting, retries, priority queues and item pipelines. Selectors and a crawling engine are not part of the core - plug in your own parser if you need one (see examples/quotes.py).
Beta notice: APIs and behavior may change; expect sharp edges while things settle.
Table of Contents
- What is aioscraper?
- Key Features
- Installation
- Quick Start
- Examples
- Why aioscraper?
- Use Cases
- Performance
- Documentation
- Changelog
- Contributing
What is aioscraper?
aioscraper is an async Python framework designed for mass data collection from APIs and external services at scale.
Built for:
- Fetching data from hundreds/thousands of REST API endpoints concurrently
- Integrating multiple external services (payment gateways, analytics APIs, etc.)
- Building data aggregation pipelines from heterogeneous API sources
- Queue-based scraping workers consuming tasks from Redis/RabbitMQ
- Microservice fan-out requests with automatic rate limiting and retries
NOT built for:
- Parsing HTML/CSS (but nothing stops you from using BeautifulSoup if you want - see examples/quotes.py)
- Single API requests (use httpx or aiohttp directly)
- GraphQL or WebSocket scraping (different paradigm)
Think: "I need to fetch data from 10,000 product API endpoints" or "I need to poll 50 microservices every minute" → aioscraper is for you.
Key Features
- Async-first core with pluggable HTTP backends (
aiohttp/httpx/httpx2) andaiojobsscheduling - Declarative flow: requests → callbacks → pipelines, with middleware hooks at each stage
- Priority queueing with backpressure, a global concurrency limit and per-group rate limits
- Adaptive rate limiting with EWMA + AIMD algorithm - automatically backs off on server overload
- Small, explicit API that is easy to test and compose with existing async applications
Installation
Choose your HTTP backend:
# Option 1: Use aiohttp (recommended for most cases)
pip install "aioscraper[aiohttp]"
# Option 2: Use httpx (if you prefer httpx ecosystem)
pip install "aioscraper[httpx]"
# Option 3: Use httpx2, the Pydantic-maintained fork of httpx
pip install "aioscraper[httpx2]"
# Option 4: Install several backends for flexibility
pip install "aioscraper[aiohttp,httpx,httpx2]"
Quick Start
Create scraper.py:
import logging
from aioscraper import AIOScraper, Request, Response, ScheduleRequest, Pipeline
from dataclasses import dataclass
logger = logging.getLogger("github_repos")
scraper = AIOScraper()
@dataclass(slots=True)
class RepoStats:
name: str
stars: int
language: str
# registers the pipeline that handles RepoStats items
@scraper.pipeline(RepoStats)
class StatsPipeline:
def __init__(self):
self.total_stars = 0
async def put_item(self, item: RepoStats) -> RepoStats:
# runs once per extracted item: store it, queue it, validate it, or aggregate as here
self.total_stars += item.stars
logger.info("✓ %s: ⭐ %s (%s)", item.name, item.stars, item.language)
return item
async def close(self):
# runs once when the scraper stops: flush buffers, close connections, report totals
logger.info("Total stars collected: %s", self.total_stars)
# registers an entry point; schedule_request is injected by parameter name
@scraper
async def get_repos(schedule_request: ScheduleRequest):
repos = (
"django/django",
"fastapi/fastapi",
"pallets/flask",
"encode/httpx",
"aio-libs/aiohttp",
)
for repo in repos:
await schedule_request(
Request(
url=f"https://api.github.com/repos/{repo}",
callback=parse_repo, # runs on a response with a status below 400
errback=on_failure, # runs on anything else: 4xx/5xx, timeouts, connection failures
cb_kwargs={"repo": repo}, # extra arguments for both of them
headers={"Accept": "application/vnd.github+json"}, # required by the GitHub API
)
)
async def parse_repo(response: Response, pipeline: Pipeline):
# the body has to be read here: the connection is released when the callback returns
data = await response.json()
await pipeline(
RepoStats(
name=data["full_name"],
stars=data["stargazers_count"],
language=data.get("language", "Unknown"),
)
)
async def on_failure(exc: Exception, repo: str):
logger.error("%s: cannot parse response: %s", repo, exc)
Run it:
aioscraper scraper
What's happening?
@scraperregisters the entry point;@scraper.pipelineregisters a pipeline forRepoStatsschedule_request()queues a request and returns; the framework dispatches it when a slot frees up- Requests run concurrently up to the limit, so responses arrive in no particular order
parse_repohandles each response,on_failurehandles each failure that was not retriedStatsPipeline.close()runs once at the end, after every request has finished
Retries are on by default. Rate limiting is not - turn it on, along with concurrency and timeouts, through environment variables before running this against a real API.
Examples
Runnable, commented scrapers live in examples/.
Why aioscraper?
vs Scrapy:
- Scrapy is built for HTML scraping with CSS/XPath selectors and website crawling
- aioscraper is optimized for API data collection (JSON, REST, microservices)
- Native asyncio (no Twisted), modern type hints, minimal footprint
- Easily embeds into existing async applications
vs httpx/aiohttp directly:
- Manual approach: you handle rate limiting, retries, queuing, concurrency, backpressure
- aioscraper: adaptive rate limits, priority queues, pipelines, middleware out of the box
- Declarative Request → callback → pipeline instead of imperative control flow
vs building custom async workers:
- Less boilerplate: focus on business logic, not infrastructure
- Production-ready components: EWMA+AIMD rate limiting, graceful shutdown, dependency injection
- Testable: explicit dependencies, no global state, easy mocking
When to use aioscraper:
- Collecting data from 100+ API endpoints
- Fan-out calls to microservices for data enrichment
- Queue consumers processing API scraping tasks
- API aggregation/monitoring pipelines
- High-throughput data collection jobs
Use Cases
1. E-commerce price monitoring
Poll 10,000 product API endpoints across multiple marketplaces:
- Adaptive rate limiting prevents bans
- Priority queue for trending products
- Pipeline aggregates prices → saves to DB → sends alerts on changes
2. Cryptocurrency data aggregation
Collect real-time prices from 20+ exchange APIs:
- Concurrent requests with per-exchange rate limits
- Built-in retry for transient failures
- Pipeline normalizes data formats → writes to time-series DB
3. Microservice data hydration
Your FastAPI app needs data from 50 internal services:
- Embed aioscraper in your async application
- Fan-out concurrent requests with backpressure control
- Middleware for auth, logging, circuit breaking
4. Queue-based scraping workers
Distributed architecture with Redis/RabbitMQ/SQS:
- Message queue publishes scraping tasks (URLs + params)
- aioscraper workers consume queue → fetch data → process
- Pipeline acknowledges messages after successful processing
5. Social media API aggregation
Aggregate user stats from Twitter, LinkedIn, GitHub APIs:
- Different rate limits per platform (adaptive throttling)
- Error callbacks for quota exceeded / auth failures
- Pipeline deduplicates → enriches → stores to database
6. Multi-source data snapshots
Collect point-in-time data from 500+ API sources simultaneously:
- Health monitoring: poll status endpoints of distributed services every minute
- Market data: snapshot prices from 200+ suppliers at exact intervals
- Analytics aggregation: fetch metrics from dozens of analytics APIs on schedule
- Concurrent execution with precise timing and automatic retries for failed sources
Performance
Benchmarks show stable throughput across CPython 3.11–3.14 (see benchmarks)
Documentation
Full documentation at aioscraper.readthedocs.io
Changelog
See CHANGELOG.md for version history and release notes.
Contributing
Please see the Contributing guide for workflow, tooling, and review expectations.
License
MIT License
Copyright (c) 2025 darkstussy
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file aioscraper-0.16.0.tar.gz.
File metadata
- Download URL: aioscraper-0.16.0.tar.gz
- Upload date:
- Size: 910.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b1d010473e0ce731ec22da50b5481e649ada92f48300f82712a61ae9c8b19efd
|
|
| MD5 |
a7004d7ae64b37a5c43e249a7d460e26
|
|
| BLAKE2b-256 |
26132efd7a79c1136cff68503fa63b08bf831a69bfd33853704ab5b3912aa08f
|
Provenance
The following attestation bundles were made for aioscraper-0.16.0.tar.gz:
Publisher:
release.yml on DarkStussy/aioscraper
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
aioscraper-0.16.0.tar.gz -
Subject digest:
b1d010473e0ce731ec22da50b5481e649ada92f48300f82712a61ae9c8b19efd - Sigstore transparency entry: 2582781687
- Sigstore integration time:
-
Permalink:
DarkStussy/aioscraper@807858a90556c33a27d21f9dd0ae7c4273df615f -
Branch / Tag:
refs/tags/0.16.0 - Owner: https://github.com/DarkStussy
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@807858a90556c33a27d21f9dd0ae7c4273df615f -
Trigger Event:
release
-
Statement type:
File details
Details for the file aioscraper-0.16.0-py3-none-any.whl.
File metadata
- Download URL: aioscraper-0.16.0-py3-none-any.whl
- Upload date:
- Size: 79.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
719dc668acae6391d54c3e811f8e1fac3f774527a690d1f149feac5f6da3b06c
|
|
| MD5 |
53a2a6b49460554f6b247e091ed2d2b4
|
|
| BLAKE2b-256 |
43b2275b6e82b3e4921703766ada38ce30ecf0f2b6ddc2fa0544575455b12351
|
Provenance
The following attestation bundles were made for aioscraper-0.16.0-py3-none-any.whl:
Publisher:
release.yml on DarkStussy/aioscraper
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
aioscraper-0.16.0-py3-none-any.whl -
Subject digest:
719dc668acae6391d54c3e811f8e1fac3f774527a690d1f149feac5f6da3b06c - Sigstore transparency entry: 2582781717
- Sigstore integration time:
-
Permalink:
DarkStussy/aioscraper@807858a90556c33a27d21f9dd0ae7c4273df615f -
Branch / Tag:
refs/tags/0.16.0 - Owner: https://github.com/DarkStussy
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@807858a90556c33a27d21f9dd0ae7c4273df615f -
Trigger Event:
release
-
Statement type: