Skip to main content

snapshot — offline website snapshots

Take a fast offline copy of any public website, then serve it locally.

Install

Requires Python 3.10+.

pip install web-snapshot-cli

Then run:

snapshot https://example.com ./mirror

Command generator: snapshot.sushii.dev

Usage

Snapshot a single page

snapshot https://example.com ./mirror

Crawl an entire site

snapshot --crawl https://docs.example.com ./docs --max-pages 200 --depth 4

Save pages as Markdown

snapshot https://example.com ./mirror --lang md

Crawl with filters and politeness

snapshot --crawl --include '/docs/*' --exclude '/docs/drafts/*' \
  --crawl-delay 1 --robots https://docs.example.com ./docs

Authenticated pages

snapshot --cookie session=abc123 --header "Authorization: Bearer TOKEN" \
  https://app.example.com ./mirror

Sitemap-based crawl

snapshot --crawl --sitemap https://example.com ./mirror --max-pages 500

Gobuster-style path discovery

Brute-force common paths (including robots.txt, sitemap.xml, admin panels, etc.) using bundled wordlists (common + ~18k large entries). Add your own SecLists/gobuster wordlists with -w:

snapshot --crawl --gobuster https://example.com ./mirror
snapshot --crawl -w common -w large -w /path/to/SecLists/Discovery/Web-Content/common.txt URL ./out
snapshot --crawl --gobuster --wordlist-ext html,php,asp URL ./out

Resume an interrupted snapshot

snapshot --resume --crawl https://example.com ./mirror

Dry run (no writes)

snapshot --dry-run --verbose --crawl https://example.com ./mirror

Extra args (positional)

Any key=value pairs after the output directory are merged into options:

snapshot https://example.com ./mirror crawl=true max-pages=100 lang=html concurrency=32

Restore locally

snapshot -restore ./mirror

This starts a local HTTP server (default http://127.0.0.1:8080) and opens your browser.

snapshot -restore ./mirror --port 3000 --no-open

Options

Flag Description
--crawl, -c Follow same-origin links and download all pages
--lang, -l Output format: html (default) or md
--max-pages Max pages when crawling (default: 50)
--depth Max crawl depth (default: 3)
--no-assets Skip CSS, JS, images, fonts
--timeout HTTP timeout in seconds (default: 15)
--concurrency Parallel downloads (default: 16)
--same-origin / --no-same-origin Restrict crawl to same origin (default: on)
--user-agent Custom User-Agent header
--cookie Cookie as name=value (repeatable)
--header Extra HTTP header (repeatable)
--include Only fetch URLs matching glob (repeatable)
--exclude Skip URLs matching glob (repeatable)
--robots / --no-robots Respect robots.txt (default: on)
--crawl-delay Seconds to wait after each request (default: 0)
--sitemap Seed crawl from sitemap.xml
--gobuster, -g Discover paths with bundled wordlists
--wordlist, -w Custom or builtin wordlist (common, large)
--wordlist-ext Extensions to append per word (html,php)
--resume Skip pages/assets already saved
--verbose, -v Detailed log output
--dry-run Fetch without writing files
-restore DIR Serve a saved snapshot from DIR
--port Port for restore (default: 8080)
--host Host for restore (default: 127.0.0.1)
--no-open Don't open a browser on restore

How it works

  1. snapshot fetches pages with async HTTP (httpx), rewrites links to local paths, and downloads linked assets in parallel.
  2. A .snapshot.json manifest is written to the output folder with metadata for restore.
  3. snapshot -restore serves the saved files and maps / back to the original root page.

Output layout

mirror/
  .snapshot.json
  example.com/
    index.html
    about/
      index.html
    _assets/
      ...

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

web_snapshot_cli-1.3.0.tar.gz (80.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

web_snapshot_cli-1.3.0-py3-none-any.whl (75.5 kB view details)

Uploaded Python 3

File details

Details for the file web_snapshot_cli-1.3.0.tar.gz.

File metadata

  • Download URL: web_snapshot_cli-1.3.0.tar.gz
  • Upload date:
  • Size: 80.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for web_snapshot_cli-1.3.0.tar.gz
Algorithm Hash digest
SHA256 6191f0e3b8e1998efd40f1c555447faea2827292635c392200d156ba9e34b300
MD5 35d0d939a682635d998ef0815b774aa9
BLAKE2b-256 6c51a8e7effb85e6ee03aa8746340d7a8e33dd5711cfca98ac585c80af492047

See more details on using hashes here.

Provenance

The following attestation bundles were made for web_snapshot_cli-1.3.0.tar.gz:

Publisher: publish.yml on codingsushi79/Snapshot

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file web_snapshot_cli-1.3.0-py3-none-any.whl.

File metadata

File hashes

Hashes for web_snapshot_cli-1.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 5bf362ce5e3a1fd4bad2ba8bb1f91c2747f6226b661a6a221f82d185120ecd1e
MD5 e5e35daebe7db95dd1a3f1950144ddbb
BLAKE2b-256 bfe7057b548a9fc2eb76460d999984dc52f7534ec98d1926a97a2fe1d16be7f3

See more details on using hashes here.

Provenance

The following attestation bundles were made for web_snapshot_cli-1.3.0-py3-none-any.whl:

Publisher: publish.yml on codingsushi79/Snapshot

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page