Skip to main content

Job Scraper Tool

CI License Docker Python

Self-hosted job search that scores every listing against your CV — AI optional. Paste your CV and a few keywords, hit Start, and it searches multiple job sites and ranks results 0-10 by how well they actually match you, instead of you skimming ten tabs every morning.

Built it for myself while job hunting. It found the job. Now I'm making it good enough for other people to use too.

Job Scraper Tool dashboard

Why this one

  • AI scoring is optional, not required. Bring a Groq, Anthropic, or Gemini key for AI-scored relevance, or run entirely free in Lite Mode (keyword matching, no key, no signup).
  • Your data stays yours. Everything lands in a local SQLite file you own — nothing is sent anywhere unless you explicitly configure an AI key or a notification webhook.
  • One command to run the whole thing. docker compose up --build and you have a dashboard, a REST API, and a database. No local Python setup, no dependency hell.
  • Built to be scriptable. Everything the UI does is also a REST endpoint (see API for power users) and a CLI command, so you can wire it into your own automation.

What you get

  • Paste your CV and keywords into the web UI
  • AI scores every job 0-10 for relevance (or Lite Mode keyword matching, no key needed)
  • Pause, resume, or restart runs from the dashboard
  • Export results as JSON or CSV
  • Optional webhook/email alerts when a high-scoring job shows up
  • All data lives in a SQLite database you actually own

Sidebar: sites, keywords, CV, and run controls

How to run it

You need Docker. That's it.

1. Set your environment

cp .env.example .env

Edit .env and drop in any AI keys you have (Groq, Anthropic, or Gemini). If you don't have any, Lite Mode works fine with keyword matching.

2. Spin it up

docker compose up --build

This builds two containers:

  • scraper at http://localhost:8000 (the brain)
  • UI at http://localhost:8501 (your dashboard)

The optional n8n automation engine lives under a separate profile if you want it later:

docker compose --profile automation up

3. Open the UI

Go to http://localhost:8501, paste your CV, add some keywords like "senior python remote", pick your AI provider (or stay in Lite Mode), and hit Start. Watch the progress bar fill up. High-scoring jobs bubble to the top.

4. Export when done

curl http://localhost:8000/export/csv > jobs.csv

Docker is the only way

This app is designed to run inside Docker containers on a Linux VM. Do not try to run it natively on Windows or macOS. The scraper uses Playwright, the UI needs Streamlit, and the database expects a Unix path structure. Docker handles all of that for you.

Requirements:

  • Docker Engine 24+ or Docker Desktop
  • A Linux VM (WSL2 on Windows, OrbStack or Docker Desktop on Mac, any Linux host)
  • 2GB RAM minimum, 4GB recommended

Environment variables

Variable What it does Default
GROQ_API_KEY Groq AI scoring empty
ANTHROPIC_API_KEY Claude AI scoring empty
GEMINI_API_KEY Google AI scoring empty
DATA_DIR Where SQLite and logs live ./data
REQUEST_DELAY_SECONDS Politeness between searches 2.0
RETRY_MAX_ATTEMPTS How many times to retry a failed search 5
MIN_SCORE_NOTIFY Minimum score (0-10) to trigger a notification 7
NOTIFICATION_WEBHOOK Webhook URL for high-score job alerts empty
EMAIL_HOST / EMAIL_PORT / EMAIL_USER / EMAIL_PASS / EMAIL_TO SMTP settings for email alerts empty

API for power users

The scraper exposes a FastAPI server. The UI talks to it, but you can too.

Start a run:

curl -X POST http://localhost:8000/run \
  -H "Content-Type: application/json" \
  -d '{"provider":"groq","lite_mode":true,"sites":["example.com"],"keywords":["python"],"cv_text":"developer"}'

Check status:

curl http://localhost:8000/status

Pause a running job:

curl -X POST http://localhost:8000/pause

Resume:

curl -X POST http://localhost:8000/resume

Kill it:

curl -X POST http://localhost:8000/stop

Makefile shortcuts

make build
make up
make down
make logs

Keeping your keys safe

Never commit .env. It is gitignored by default. If you accidentally pushed a key, rotate it immediately.

Contributing

Issues and PRs are welcome — see CONTRIBUTING.md. If you're using this and hit something, opening an issue is the single most useful thing you can do; this is a solo project so far and every report helps.

If you find it useful, a star helps other people find it too.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

job_scraper02-0.3.0.tar.gz (17.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

job_scraper02-0.3.0-py3-none-any.whl (16.9 kB view details)

Uploaded Python 3

File details

Details for the file job_scraper02-0.3.0.tar.gz.

File metadata

  • Download URL: job_scraper02-0.3.0.tar.gz
  • Upload date:
  • Size: 17.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for job_scraper02-0.3.0.tar.gz
Algorithm Hash digest
SHA256 46f45e7b4fb81ca9379f48103095af618aeacfc99780e441bc72bc2e55dfeb79
MD5 eead0c40f2edf2df05d97c0d5ef7dbb9
BLAKE2b-256 b3a6f5a64edbaf067269f98ed98b853d89a28bb1d73f8a64a8b07fcf72d617d7

See more details on using hashes here.

Provenance

The following attestation bundles were made for job_scraper02-0.3.0.tar.gz:

Publisher: ci.yml on firaslamouchi21/Job-Scraper02

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file job_scraper02-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: job_scraper02-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 16.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for job_scraper02-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 eb916b02629edceda81c1b17d6aa8cda1fe125beec5c586161af0f3d4f6f380c
MD5 c5682d62cfb4f3b4b6b9eae6c6692a47
BLAKE2b-256 95b719920b81f5f916e9aa89bca2c897280616914505ebb43be7d5550fa59e4e

See more details on using hashes here.

Provenance

The following attestation bundles were made for job_scraper02-0.3.0-py3-none-any.whl:

Publisher: ci.yml on firaslamouchi21/Job-Scraper02

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 files

0.2.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page