snapshot — offline website snapshots
Take a fast offline copy of any public website, then serve it locally.
Install
Requires Python 3.10+.
pip install web-snapshot-cli
Then run:
snapshot https://example.com ./mirror
Command generator: snapshot.sushii.dev
Usage
Snapshot a single page
snapshot https://example.com ./mirror
Crawl an entire site
snapshot --crawl https://docs.example.com ./docs --max-pages 200 --depth 4
Save pages as Markdown
snapshot https://example.com ./mirror --lang md
Crawl with filters and politeness
snapshot --crawl --include '/docs/*' --exclude '/docs/drafts/*' \
--crawl-delay 1 --robots https://docs.example.com ./docs
Authenticated pages
snapshot --cookie session=abc123 --header "Authorization: Bearer TOKEN" \
https://app.example.com ./mirror
Sitemap-based crawl
snapshot --crawl --sitemap https://example.com ./mirror --max-pages 500
Gobuster-style path discovery
Brute-force common paths (including robots.txt, sitemap.xml, admin panels, etc.) using bundled wordlists (common + ~18k large entries). Add your own SecLists/gobuster wordlists with -w:
snapshot --crawl --gobuster https://example.com ./mirror
snapshot --crawl -w common -w large -w /path/to/SecLists/Discovery/Web-Content/common.txt URL ./out
snapshot --crawl --gobuster --wordlist-ext html,php,asp URL ./out
Resume an interrupted snapshot
snapshot --resume --crawl https://example.com ./mirror
Dry run (no writes)
snapshot --dry-run --verbose --crawl https://example.com ./mirror
Extra args (positional)
Any key=value pairs after the output directory are merged into options:
snapshot https://example.com ./mirror crawl=true max-pages=100 lang=html concurrency=32
Restore locally
snapshot -restore ./mirror
This starts a local HTTP server (default http://127.0.0.1:8080) and opens your browser.
snapshot -restore ./mirror --port 3000 --no-open
Options
| Flag | Description |
|---|---|
--crawl, -c |
Follow same-origin links and download all pages |
--lang, -l |
Output format: html (default) or md |
--max-pages |
Max pages when crawling (default: 50) |
--depth |
Max crawl depth (default: 3) |
--no-assets |
Skip CSS, JS, images, fonts |
--timeout |
HTTP timeout in seconds (default: 15) |
--concurrency |
Parallel downloads (default: 16) |
--same-origin / --no-same-origin |
Restrict crawl to same origin (default: on) |
--user-agent |
Custom User-Agent header |
--cookie |
Cookie as name=value (repeatable) |
--header |
Extra HTTP header (repeatable) |
--include |
Only fetch URLs matching glob (repeatable) |
--exclude |
Skip URLs matching glob (repeatable) |
--robots / --no-robots |
Respect robots.txt (default: on) |
--crawl-delay |
Seconds to wait after each request (default: 0) |
--sitemap |
Seed crawl from sitemap.xml |
--gobuster, -g |
Discover paths with bundled wordlists |
--wordlist, -w |
Custom or builtin wordlist (common, large) |
--wordlist-ext |
Extensions to append per word (html,php) |
--resume |
Skip pages/assets already saved |
--verbose, -v |
Detailed log output |
--dry-run |
Fetch without writing files |
-restore DIR |
Serve a saved snapshot from DIR |
--port |
Port for restore (default: 8080) |
--host |
Host for restore (default: 127.0.0.1) |
--no-open |
Don't open a browser on restore |
How it works
- snapshot fetches pages with async HTTP (httpx), rewrites links to local paths, and downloads linked assets in parallel.
- A
.snapshot.jsonmanifest is written to the output folder with metadata for restore. - snapshot -restore serves the saved files and maps
/back to the original root page.
Output layout
mirror/
.snapshot.json
example.com/
index.html
about/
index.html
_assets/
...
License
MIT
Metadata
Release files for web-snapshot-cli 1.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| web_snapshot_cli-1.3.0.tar.gz | 80.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| web_snapshot_cli-1.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 155.7 kB
Release files / web_snapshot_cli-1.3.0.tar.gz
| Download URL | web_snapshot_cli-1.3.0.tar.gz |
|---|---|
| Size | 80.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6191f0e3b8e1998efd40f1c555447faea2827292635c392200d156ba9e34b300
|
|
BLAKE2b-256 checksum How to use checksums |
6c51a8e7effb85e6ee03aa8746340d7a8e33dd5711cfca98ac585c80af492047
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 7, 2026.
Transparency logRelease files / web_snapshot_cli-1.3.0-py3-none-any.whl
| Download URL | web_snapshot_cli-1.3.0-py3-none-any.whl |
|---|---|
| Size | 75.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5bf362ce5e3a1fd4bad2ba8bb1f91c2747f6226b661a6a221f82d185120ecd1e
|
|
BLAKE2b-256 checksum How to use checksums |
bfe7057b548a9fc2eb76460d999984dc52f7534ec98d1926a97a2fe1d16be7f3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 7, 2026.
Transparency log