Apify Public Data Scrapers & Extractors
A curated collection of reliable, production-ready scrapers and public-data extractors hosted on the Apify Store.
Each actor is built with strict schema validation, deterministic field mapping, self-healing DOM selectors, and pay-per-event pricing starting at $0.0002 / start.
Quick Navigation
- Available Extractors & Store Listings
- Python Quickstart
- Node.js Quickstart
- No-Code & Automation Workflows (n8n, Sheets, Slack)
- Pre-Built Example Tasks (Zero Code)
- Free Sample Datasets
- AI Agent & MCP Integration (Claude Desktop, Cursor, Custom Agent)
- In-Depth Engineering Guides
- Repository Structure
- Contributing & Author
Available Extractors & Store Listings
| Tool | Store Link | Key Output Fields | Best For |
|---|---|---|---|
| Google Maps Business Leads | captainhandsome/google-maps-business-search |
Name, phone, website, rating, reviews, address, coordinates, hours | B2B lead generation, local agency prospecting |
| Glassdoor Jobs & Salaries | captainhandsome/glassdoor-jobs-scraper |
Title, company, salary estimate, rating, location, job URL, posting date | Hiring intelligence, compensation benchmarking |
| Airbnb Vacation Rentals | captainhandsome/airbnb-listings-search |
Title, room type, nightly price, rating, reviews count, listing URL | Real estate research, market rate tracking |
| SEC EDGAR Corporate Filings | captainhandsome/sec-edgar-filings-search |
Ticker, CIK, form (10-K, 10-Q, 8-K), filing date, primary document URL | Financial diligence, equity research, compliance |
| USAspending Federal Awards | captainhandsome/usaspending-federal-awards |
Recipient vendor, award amount, awarding agency, description, dates | Government contracting, procurement intel |
| LinkedIn Public Jobs | captainhandsome/linkedin-public-jobs-search |
Job title, employer, location, direct apply URL, posting age | Recruitment, tech talent monitoring |
| Google Play App Reviews | captainhandsome/google-play-reviews-scraper |
Review text, star score, thumbs up, date, reviewer name | App store sentiment, competitor feedback |
| YouTube Video Search | captainhandsome/youtube-search-scraper |
Title, video URL, channel, views count, duration, publish date | Content tracking, creator outreach |
| Twitch Live Streams | captainhandsome/twitch-live-streams-scraper |
Streamer username, title, viewer count, language, category | Esports analytics, live stream monitoring |
| US Contractor Licenses | captainhandsome/us-contractor-license-search |
Contractor name, license number, classification, status, state | Trades verification, subcontractor diligence |
| US Business Entity Registries | captainhandsome/us-business-entity-search |
Legal entity name, filing number, jurisdiction, status | Legal due diligence, corporate registration checks |
Python Quickstart
1. Install dependencies
pip install apify-client pandas python-dotenv
2. Export 50 Google Maps Leads to CSV
import os
from apify_client import ApifyClient
import pandas as pd
# Get your API token from https://console.apify.com/account/integrations
client = ApifyClient(os.getenv("APIFY_TOKEN"))
# Run the actor
run = client.actor("captainhandsome/google-maps-business-search").call(run_input={
"search_query": "commercial electricians",
"location": "Dallas, Texas",
"max_items": 50,
"include_details": True,
})
# Fetch dataset items and export to CSV
items = list(client.dataset(run["defaultDatasetId"]).iterate_items())
df = pd.DataFrame(items)
df.to_csv("dallas_electricians.csv", index=False)
print(f"Exported {len(df)} leads to dallas_electricians.csv")
See examples/google_maps_leads_to_csv.py for the full script.
Node.js Quickstart
1. Install dependencies
npm install apify-client
2. Query SEC EDGAR Filings
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('captainhandsome/sec-edgar-filings-search').call({
companies: ['AAPL', 'NVDA', 'MSFT'],
forms: ['10-K'],
max_items: 15,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach(filing => {
console.log(`[${filing.ticker}] ${filing.form} (${filing.filing_date}): ${filing.primary_document_url}`);
});
See examples/sec_filings.js for the full script.
No-Code & Automation Workflows
If you automate via n8n, Make, Zapier, or Google Sheets, ready-to-import blueprints are included in workflows/:
- Google Maps Leads to Google Sheets (n8n): Daily automated cron scrape piping HVAC/trade leads directly into Google Sheets with deduplication.
- SEC EDGAR 10-K & 8-K Alerts to Slack (n8n): Hourly monitor alerting Slack or Discord when watchlisted public companies drop new filings.
Pre-Built Example Tasks (Zero Code)
If you prefer runnable web UI tasks without writing any code, each actor includes pre-configured tasks published on Apify Store:
Google Maps Leads
- Phoenix HVAC Company Leads
- Dallas Commercial Electrician Leads
- Chicago Italian Restaurants & Reviews
Glassdoor Jobs
Airbnb Rentals
YouTube & Google Play
Free Sample Datasets
Looking for clean data to benchmark, analyze, or train models? Verified sample bundles with metadata schemas are available in datasets/ and hosted publicly on Hugging Face Datasets:
- Phoenix HVAC Contractor Leads:
datasets/phoenix_hvac_leads/| Hugging Face Hub (20 verified HVAC contractor profiles with ratings, addresses, and phone numbers). - California Licensed Contractors:
datasets/california_solar_contractors/| Hugging Face Hub (Active C-46 and B licensed solar installers with state verification numbers). - Austin Software Engineer Postings:
datasets/austin_software_jobs/| Hugging Face Hub (Normalized job listings with estimated posting dates and salary ranges).
AI Agent & MCP Integration
All actors in this repository conform to OpenAPI and JSON Schema standards, making them directly callable by AI agents via the Model Context Protocol (MCP):
Option A: Hosted Apify MCP Server (Claude Desktop / Cursor)
Add this to your claude_desktop_config.json or Cursor MCP settings:
{
\"mcpServers\": {
\"apify\": {
\"command\": \"npx\",
\"args\": [\"-y\", \"@apify/mcp-server\"],
\"env\": {
\"APIFY_TOKEN\": \"YOUR_APIFY_API_TOKEN\"
}
}
}
}
Option B: Local Lightweight Python MCP Server
For local agent workflows without Node.js dependencies, a direct Python MCP server is included:
export APIFY_TOKEN=\"your_token_here\"
python mcp_server.py
Inspect tools and capabilities via mcp.json.
Agent Prompts That Work Out-of-the-Box:
- "Search Google Maps for 50 commercial roofers in Atlanta with phone numbers and websites."
- "Retrieve Apple and Microsoft Form 10-K filings from SEC EDGAR for the last 2 years."
- "Find the 30 newest reviews for Duolingo on Google Play and analyze negative feedback."
In-Depth Engineering Guides
Technical case studies and problem-solution writeups are located in articles/:
- Bypassing Playwright Headless Pagination Hurdles on Airbnb: How to solve sticky overlay modal interruptions and viewport boundary clipping in large headless browser crawls.
- Extracting & Normalizing Clean Job Posting Dates from Glassdoor: Overcoming relative timestamp drift ("24h", "3d", "30d+") with deterministic parsing and ISO-8601 boundary tracking.
Repository Structure
apify-scrapers/
README.md # Documentation and quickstart
LICENSE # MIT License
requirements.txt # Python client dependencies
package.json # Node.js dependencies
mcp.json # MCP tool registry specification
mcp_server.py # Native Python stdio MCP server
articles/ # In-depth engineering case studies
airbnb_playwright_pagination_guide.md
glassdoor_posting_dates_guide.md
reddit_community_responses.md # Reference technical answers for forums
datasets/ # Sample benchmark datasets
phoenix_hvac_leads/
california_solar_contractors/
austin_software_jobs/
workflows/ # No-code automation templates
n8n_google_maps_to_sheets.json
n8n_sec_edgar_to_slack.json
README.md
examples/ # Standalone developer scripts
google_maps_leads_to_csv.py
sec_edgar_filings_downloader.py
glassdoor_jobs_tracker.py
airbnb_market_scraper.py
usaspending_defense_awards.py
twitch_live_stream_monitor.py
google_maps_leads.js
sec_filings.js
Author & Support
Maintained by Joseph McRell.
- Apify Store: https://apify.com/captainhandsome
- GitHub: @jlucasmcrell
- Hugging Face: @joeygambino
- Issues & Requests: Please open an issue on this repository or submit a ticket on the respective Apify Actor Store page.
License
This project is licensed under the MIT License - see the LICENSE file for details.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file apify_data_scrapers-1.0.0.tar.gz.
File metadata
- Download URL: apify_data_scrapers-1.0.0.tar.gz
- Upload date:
- Size: 9.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
77be4f736772a91ff8010ff00800439b453b2cc857c415721aa313a3b8c32faa
|
|
| MD5 |
f01a814fa747161f36e052af74b643ae
|
|
| BLAKE2b-256 |
700a7b7676b0666950fff4c829ec508e243300be6d43a26a68a72ddd06d20d6f
|
File details
Details for the file apify_data_scrapers-1.0.0-py3-none-any.whl.
File metadata
- Download URL: apify_data_scrapers-1.0.0-py3-none-any.whl
- Upload date:
- Size: 10.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
01ca0c9f5cdd6ccb18f8c9570806034a048a36c0bab16a2abf4d0a3810e0c13b
|
|
| MD5 |
89b9ca8f7985ba3214a1747cdc0f3ed3
|
|
| BLAKE2b-256 |
b84e4c77243ca95fc091e607a3b00092925829f0da5e39de0c976b09fd84ab50
|