Skip to main content

snapshot — offline website snapshots

Take a fast offline copy of any public website, then serve it locally.

Install

Requires Python 3.10+.

pip install web-snapshot-cli

Then run:

snapshot https://example.com ./mirror

Command generator: snapshot.sushii.dev

Usage

Snapshot a single page

snapshot https://example.com ./mirror

Crawl an entire site

snapshot --crawl https://docs.example.com ./docs --max-pages 200 --depth 4

Save pages as Markdown

snapshot https://example.com ./mirror --lang md

Crawl with filters and politeness

snapshot --crawl --include '/docs/*' --exclude '/docs/drafts/*' \
  --crawl-delay 1 --robots https://docs.example.com ./docs

Authenticated pages

snapshot --cookie session=abc123 --header "Authorization: Bearer TOKEN" \
  https://app.example.com ./mirror

Sitemap-based crawl

snapshot --crawl --sitemap https://example.com ./mirror --max-pages 500

Gobuster-style path discovery

Brute-force common paths (including robots.txt, sitemap.xml, admin panels, etc.) using bundled wordlists (common + ~18k large entries). Add your own SecLists/gobuster wordlists with -w:

snapshot --crawl --gobuster https://example.com ./mirror
snapshot --crawl -w common -w large -w /path/to/SecLists/Discovery/Web-Content/common.txt URL ./out
snapshot --crawl --gobuster --wordlist-ext html,php,asp URL ./out

Resume an interrupted snapshot

snapshot --resume --crawl https://example.com ./mirror

Dry run (no writes)

snapshot --dry-run --verbose --crawl https://example.com ./mirror

Extra args (positional)

Any key=value pairs after the output directory are merged into options:

snapshot https://example.com ./mirror crawl=true max-pages=100 lang=html concurrency=32

Restore locally

snapshot -restore ./mirror

This starts a local HTTP server (default http://127.0.0.1:8080) and opens your browser.

snapshot -restore ./mirror --port 3000 --no-open

Options

Flag Description
--crawl, -c Follow same-origin links and download all pages
--lang, -l Output format: html (default) or md
--max-pages Max pages when crawling (default: 50)
--depth Max crawl depth (default: 3)
--no-assets Skip CSS, JS, images, fonts
--timeout HTTP timeout in seconds (default: 15)
--concurrency Parallel downloads (default: 16)
--same-origin / --no-same-origin Restrict crawl to same origin (default: on)
--user-agent Custom User-Agent header
--cookie Cookie as name=value (repeatable)
--header Extra HTTP header (repeatable)
--include Only fetch URLs matching glob (repeatable)
--exclude Skip URLs matching glob (repeatable)
--robots / --no-robots Respect robots.txt (default: on)
--crawl-delay Seconds to wait after each request (default: 0)
--sitemap Seed crawl from sitemap.xml
--gobuster, -g Discover paths with bundled wordlists
--wordlist, -w Custom or builtin wordlist (common, large)
--wordlist-ext Extensions to append per word (html,php)
--resume Skip pages/assets already saved
--verbose, -v Detailed log output
--dry-run Fetch without writing files
-restore DIR Serve a saved snapshot from DIR
--port Port for restore (default: 8080)
--host Host for restore (default: 127.0.0.1)
--no-open Don't open a browser on restore

How it works

  1. snapshot fetches pages with async HTTP (httpx), rewrites links to local paths, and downloads linked assets in parallel.
  2. A .snapshot.json manifest is written to the output folder with metadata for restore.
  3. snapshot -restore serves the saved files and maps / back to the original root page.

Output layout

mirror/
  .snapshot.json
  example.com/
    index.html
    about/
      index.html
    _assets/
      ...

License

MIT

Metadata

Release files for web-snapshot-cli 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for web-snapshot-cli 1.3.0
File Size Uploaded
web_snapshot_cli-1.3.0.tar.gz 80.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for web-snapshot-cli 1.3.0
File Interpreter ABI Platform
web_snapshot_cli-1.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 155.7 kB

Release files / web_snapshot_cli-1.3.0.tar.gz

Download URL web_snapshot_cli-1.3.0.tar.gz
Size 80.2 kB
Tags Source
SHA-256 checksum
How to use checksums
6191f0e3b8e1998efd40f1c555447faea2827292635c392200d156ba9e34b300
BLAKE2b-256 checksum
How to use checksums
6c51a8e7effb85e6ee03aa8746340d7a8e33dd5711cfca98ac585c80af492047
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 7, 2026.

Transparency log

Release files / web_snapshot_cli-1.3.0-py3-none-any.whl

Download URL web_snapshot_cli-1.3.0-py3-none-any.whl
Size 75.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5bf362ce5e3a1fd4bad2ba8bb1f91c2747f6226b661a6a221f82d185120ecd1e
BLAKE2b-256 checksum
How to use checksums
bfe7057b548a9fc2eb76460d999984dc52f7534ec98d1926a97a2fe1d16be7f3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 7, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.3.0 This release

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page