Skip to main content

webget

Local search + scrape CLI. Zero API keys, unlimited usage.

webget is a web acquisition layer for agents and scripts: it routes every URL through a strategy ladder (HTTP fast path → Crawl4AI browser → optional Firecrawl), and reports provenance - where the content came from and whether the session that fetched it can be trusted.

CI PyPI Python License: MIT Stars


Why

Most scraping tools assume one engine. webget assumes the web is messy:

FETCH_AUTO
├── http       → fast path (httpx + trafilatura/html2text), no browser
├── crawl4ai   → Playwright browser, JS rendering, persistent auth sessions
└── firecrawl  → optional cloud fallback (needs WEBGET_FIRECRAWL_KEY)

Every fetch classifies what it hit - success, login_required, challenge, blocked, or error - and reports it in machine-readable JSON. webget never pretends an empty page is success, and it never solves CAPTCHAs or evades anti-bot systems; it tells you honestly what happened.

Install

Requires Python 3.11+.

PyPI (webget-cli)

The CLI command is webget; the PyPI package name is webget-cli (the bare webget name is taken by an unrelated package).

# pip - HTTP fast path + search only (no browser)
pip install webget-cli

# pip - full stack with Crawl4AI/Playwright browser fallback
pip install "webget-cli[browser]"

# uv tool - isolated executable on your PATH
uv tool install webget-cli --with "webget-cli[browser]"

After install, webget is available as a command:

webget --help

Browser runtime (optional)

Crawl4AI drives a Playwright Chromium. pip install "webget-cli[browser]" installs the Python packages; the browser binary itself is downloaded separately:

python -m playwright install chromium

Without the browser extra, webget still works for search and plain HTTP fetches. A fetch that needs the browser (JS rendering, --profile sessions, login) prints a clear warning telling you how to install it.

From source (development)

git clone https://github.com/DavidPandleton/webget
cd webget
uv pip install -e ".[dev,browser]"

Usage

webget s "rust async runtime"             # search DuckDuckGo (top 5)
webget u https://example.com              # scrape (auto: http -> crawl4ai)
webget su "llm inference" 5               # search + scrape top 5, parallel
cat urls.txt | webget u -                 # batch scrape, one browser instance
webget fetch https://example.com --json   # machine-readable result

Long aliases: search = s, fetch = u, search-fetch = su.

Options

Flag Meaning
-c, --cookies FILE Netscape-format cookie file
--profile NAME Persistent browser profile (auth session)
-H, --header "K: V" Extra header (repeatable)
-n, --max-chars N Max output chars (default: 10000 for u, 4000 for su)
--limit N Result count for s/su
-t, --timeout N Per-URL timeout seconds (default 20)
--fresh Bypass cache
--ttl N Cache TTL seconds (default 3600)
--strategy S auto | http | crawl4ai | firecrawl
--no-cache Don't read or write the disk cache (private fetch)
--json JSON output with metadata

Authenticated sessions (profiles)

# interactive login: browser opens, YOU log in manually, session persists
webget login https://campus.example --profile campus

# list profiles and their session status
webget profiles
webget profiles --json

# later fetches reuse the session - even on the HTTP fast path
webget fetch https://campus.example/dashboard --profile campus --json

# log out ONE domain, keep the rest of the profile
webget logout https://campus.example --profile campus

webget login never stores passwords and never fills forms. A visible browser opens, you authenticate yourself, then press Enter in the terminal and webget persists the session. Persistent profiles live in ~/.local/share/webget/profiles/<name>; session cookies are exported to storage_state.json inside the profile after each browser run, so the fast path can reuse them. Secrets are never printed.

JSON output

--json returns a dict keyed by URL, so batch results are easy to inspect:

{
  "https://campus.example/dashboard": {
    "status": "success",
    "method": "crawl4ai",
    "cached": false,
    "attempts": 1,
    "auth": {
      "profile": "campus",
      "authenticated": true,
      "state": "success"
    },
    "error": null
  }
}

Status values: success | login_required | challenge | blocked | error.

Status detection rules

Signal State
Valid content (≥100 chars) success
HTTP 401, login form, 403 + login markers login_required
Cloudflare / CAPTCHA / "verify you are human" challenge
HTTP 403 generic, 429, "access denied" blocked
DNS failure, timeout, unexpected exception error

Cache

Results are cached in ~/.cache/webget/ (sha1 of url + profile + options, TTL 1h, eviction at 500 files). The cache is content-level, not strategy-level, and isolated per profile - public, campus, and work fetches never collide. Failures are never cached.

Privacy note: cached content is plaintext JSON on disk. If you fetch authenticated/personal pages, use --no-cache.

Development

make dev        # install runtime + dev deps
make test       # pytest (pure logic, no network needed)
make lint       # ruff
  • Single-file Python (webget_cli.py), no build step, runs via uv run.
  • Lazy imports: --strategy http never pays the Crawl4AI import cost.
  • Crawl4AI 0.9.2's export_storage_state() is broken (wrong attribute); webget works around it by reaching into browser_manager directly.

Contributing

Found a bug or have an idea? Open an issue - we have templates. Pull requests welcome, see CONTRIBUTING.md.

License

MIT

Release files for webget-cli 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for webget-cli 0.6.0
File Size Uploaded
webget_cli-0.6.0.tar.gz 21.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for webget-cli 0.6.0
File Interpreter ABI Platform
webget_cli-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 38.2 kB

Release files / webget_cli-0.6.0.tar.gz

Download URL webget_cli-0.6.0.tar.gz
Size 21.9 kB
Tags Source
SHA-256 checksum
How to use checksums
0ac0501360a998f6ebf3677c207522a1d625ed9379f8e609e98f0332dac7d409
BLAKE2b-256 checksum
How to use checksums
c6d5ab10cfdd7cf957741f097ba6253ccc8a1e8ae2267e6e23d15a5a74ca6317
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.15 {"installer":{"name":"uv","version":"0.11.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"CachyOS Linux","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / webget_cli-0.6.0-py3-none-any.whl

Download URL webget_cli-0.6.0-py3-none-any.whl
Size 16.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
270e36d6294b6e60879f8ebe62c50fc693ebe2d60e674c416ec184ee3dd97c47
BLAKE2b-256 checksum
How to use checksums
1d4af4665910d2d2f56a3dbe7d0bad558c7aef9e3043d5bd6166340d2ccfb0c8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.15 {"installer":{"name":"uv","version":"0.11.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"CachyOS Linux","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.16.1

2 release files

0.16.0

2 release files

0.15.0

2 release files

0.14.0

2 release files

0.12.1

2 release files

0.12.0

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

This release

0.6.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page