Skip to main content

Wayback Machine Archiver

Wayback Machine Archiver (Archiver for short) is a command-line utility written in Python to back up web pages using the Internet Archive.

Installation

Install Archiver with your preferred Python package manager:

pip install wayback-machine-archiver

Or, if you use uv or pipx:

uv tool install wayback-machine-archiver
pipx install wayback-machine-archiver

This will give you access to the script simply by calling:

archiver --help

You can also install it directly from a local clone of this repository:

git clone https://github.com/agude/wayback-machine-archiver.git
cd wayback-machine-archiver
pip install .

All dependencies are handled automatically. Archiver supports Python 3.8+.

Usage

The archiver is simple to use from the command line.

Command-Line Examples

Archive a single page:

archiver https://alexgude.com

Archive all pages from a sitemap:

archiver --sitemaps https://alexgude.com/sitemap.xml

Archive from a local sitemap file: (Note the file:// prefix is required)

archiver --sitemaps file://sitemap.xml

Archive from a text file of URLs: (The file should contain one URL per line)

archiver --file urls.txt

Combine multiple sources:

archiver https://radiokeysmusic.com --sitemaps https://charles.uno/sitemap.xml

Use advanced API options: (Capture a screenshot and skip if archived in the last 10 days)

archiver https://alexgude.com --capture-screenshot --if-not-archived-within 10d

Archive the sitemap URL itself:

archiver --sitemaps https://alexgude.com/sitemaps.xml --archive-sitemap-also

Authentication (Required)

As of version 3.0.0, this tool requires authentication with the Internet Archive's SPN2 API. This change was made to ensure all archiving jobs are reliable and their final success or failure status can be confirmed. The previous, less reliable method for unauthenticated users has been removed.

If you run the script without credentials, it will exit with an error message.

To set up authentication:

  1. Get your S3-style API keys from your Internet Archive account settings: https://archive.org/account/s3.php

  2. Create a .env file in the directory where you run the archiver command. Add your keys to it:

    INTERNET_ARCHIVE_ACCESS_KEY="YOUR_ACCESS_KEY_HERE"
    INTERNET_ARCHIVE_SECRET_KEY="YOUR_SECRET_KEY_HERE"
    

The script will automatically detect this file (or the equivalent environment variables) and use the authenticated API.

Help

For a full list of command-line flags, Archiver has built-in help displayed with archiver --help:

usage: archiver [-h] [--version] [--file FILE]
                [--sitemaps SITEMAPS [SITEMAPS ...]]
                [--log {DEBUG,INFO,WARNING,ERROR,CRITICAL}]
                [--log-to-file LOG_FILE]
                [--archive-sitemap-also]
                [--rate-limit-wait RATE_LIMIT_IN_SEC]
                [--random-order] [--capture-all]
                [--capture-outlinks] [--capture-screenshot]
                [--delay-wb-availability] [--force-get]
                [--skip-first-archive] [--email-result]
                [--if-not-archived-within <timedelta>]
                [--js-behavior-timeout <seconds>]
                [--capture-cookie <cookie>]
                [--user-agent <string>]
                [urls ...]

A script to backup a web pages with Internet Archive

positional arguments:
  urls                  Specifies the URLs of the pages to archive.

options:
  -h, --help            show this help message and exit
  --version             show program's version number and exit
  --file FILE           Specifies the path to a file containing URLs to save,
                        one per line.
  --sitemaps SITEMAPS [SITEMAPS ...]
                        Specifies one or more URIs to sitemaps listing pages
                        to archive. Local paths must be prefixed with
                        'file://'.
  --log {DEBUG,INFO,WARNING,ERROR,CRITICAL}
                        Sets the logging level. Defaults to WARNING
                        (case-insensitive).
  --log-to-file LOG_FILE
                        Redirects logs to a specified file instead of the
                        console.
  --archive-sitemap-also
                        Submits the URL of the sitemap itself to be archived.
  --rate-limit-wait RATE_LIMIT_IN_SEC
                        Specifies the number of seconds to wait between
                        submissions. A minimum of 5 seconds is enforced for
                        authenticated users. Defaults to 15.
  --random-order        Randomizes the order of pages before archiving.

SPN2 API Options:
  Control the behavior of the Internet Archive capture API.

  --capture-all         Captures a web page even if it returns an error (e.g.,
                        404, 500).
  --capture-outlinks    Captures web page outlinks automatically. Note: this
                        can significantly increase the total number of
                        captures and runtime.
  --capture-screenshot  Captures a full page screenshot.
  --delay-wb-availability
                        Reduces load on Internet Archive systems by making the
                        capture publicly available after ~12 hours instead of
                        immediately.
  --force-get           Bypasses the headless browser check, which can speed
                        up captures for non-HTML content (e.g., PDFs, images).
  --skip-first-archive  Speeds up captures by skipping the check for whether
                        this is the first time a URL has been archived.
  --email-result        Sends an email report of the captured URLs to the
                        user's registered email.
  --if-not-archived-within <timedelta>
                        Captures only if the latest capture is older than
                        <timedelta> (e.g., '3d 5h').
  --js-behavior-timeout <seconds>
                        Runs JS code for <N> seconds after page load to
                        trigger dynamic content. Defaults to 5, max is 30. Use
                        0 to disable for static pages.
  --capture-cookie <cookie>
                        Uses an extra HTTP Cookie value when capturing the
                        target page.
  --user-agent <string>
                        Uses a custom HTTP User-Agent value when capturing the
                        target page.

Setting Up a Sitemap.xml for Github Pages

It is easy to automatically generate a sitemap for a Github Pages Jekyll site. Simply use jekyll/jekyll-sitemap.

Setup instructions can be found on the above site; they require changing just a single line of your site's _config.yml.

Release files for wayback-machine-archiver 4.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for wayback-machine-archiver 4.1.0
File Size Uploaded
wayback_machine_archiver-4.1.0.tar.gz 32.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for wayback-machine-archiver 4.1.0
File Interpreter ABI Platform
wayback_machine_archiver-4.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 50.1 kB

Release files / wayback_machine_archiver-4.1.0.tar.gz

Download URL wayback_machine_archiver-4.1.0.tar.gz
Size 32.0 kB
Tags Source
SHA-256 checksum
How to use checksums
61d55c4b55f767ff32965a1c9fd4e5e8d01264959d5905170e6f08b0e6677d09
BLAKE2b-256 checksum
How to use checksums
fadba3aeb4d0844981a2f744f29a9b67c921575ce770759877cbad096d7a1a0d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 24, 2026.

Transparency log

Release files / wayback_machine_archiver-4.1.0-py3-none-any.whl

Download URL wayback_machine_archiver-4.1.0-py3-none-any.whl
Size 18.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
17452dd5198cee8aaef09c0bed5edde96306f33ac9c0fccb14bce704813fc956
BLAKE2b-256 checksum
How to use checksums
d94f7a71f4cc362fedcc396b2c7bfe23e13b6dad7399e82727bb917eadf2cab1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 24, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

4.1.0 This release

2 release files

4.0.0

2 release files

3.5.0

2 release files

3.4.0

2 release files

3.3.1

2 release files

3.3.0

2 release files

3.2.0

2 release files

3.1.0

2 release files

3.0.0

2 release files

2.2.0

2 release files

2.1.0

2 release files

2.0.0

2 release files

1.10.0

2 release files

1.9.2

2 release files

1.9.1

2 release files

1.9.0

2 release files

1.8.1

2 release files

1.8.0

2 release files

1.7.3

2 release files

1.7.1

2 release files

1.7.0

2 release files

1.6.0

2 release files

1.5.1

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.2

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page