Skip to main content

scrapy-time-machine

PyPI PyPI - Python Version GitHub Workflow Status

Run your spider with a previously crawled request chain.

Install

pip install scrapy-time-machine

Why?

Lets say your spider crawls some page everyday and after some time you notice that an important information was added and you want to start saving it.

You may modify your spider and extract this information from now on, but what if you want the historical value of this data, since it was first introduced to the site?

With this extension you can save a snapshot of the site at every run to be used in the future (as long as you don't change the request chain).

Enabling

To enable this middlware, add this information to your projects's settings.py:

DOWNLOADER_MIDDLEWARES = {
    "scrapy_time_machine.timemachine.TimeMachineMiddleware": 901
}

TIME_MACHINE_ENABLED = True
TIME_MACHINE_STORAGE = "scrapy_time_machine.storages.DbmTimeMachineStorage"

Using

Store a snapshot of the current state of the site

scrapy crawl sample -s TIME_MACHINE_SNAPSHOT=true -s TIME_MACHINE_URI="/tmp/%(name)s-%(time)s.db"

This will save a snapshot at /tmp/sample-YYYY-MM-DDThh-mm-ss.db

Retrieve a snapshot from a previously saved state of the site

scrapy crawl sample -s TIME_MACHINE_RETRIEVE=true -s TIME_MACHINE_URI=/tmp/sample-YYYY-MM-DDThh-mm-ss.db

If no change was made to the spider between the current version and the version that produced the snapshot, the extracted items should be the same.

Sample project

There is a sample Scrapy project available at the examples directory.

Metadata

Release files for scrapy-time-machine 1.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scrapy-time-machine 1.1.1
File Size Uploaded
scrapy-time-machine-1.1.1.tar.gz 6.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for scrapy-time-machine 1.1.1
File Interpreter ABI Platform
scrapy_time_machine-1.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 12.7 kB

Release files / scrapy-time-machine-1.1.1.tar.gz

Download URL scrapy-time-machine-1.1.1.tar.gz
Size 6.0 kB
Tags Source
SHA-256 checksum
How to use checksums
72aabb16986c74abff8635a166e18f80f544f9fc0966b9e858ecc7afc8acfe4e
BLAKE2b-256 checksum
How to use checksums
7355fec84d80b58cc35bf9ec17fe9bca83eeda7093fe4fce35c5f3967873573c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.10.13

Release files / scrapy_time_machine-1.1.1-py3-none-any.whl

Download URL scrapy_time_machine-1.1.1-py3-none-any.whl
Size 6.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
199eea5eca5133e1978686689cb86b23a548a37feadca1bd1ccbbb861306c7bc
BLAKE2b-256 checksum
How to use checksums
931fcb2a210198495652aef03257d560c4d6ced6dd894d2447c5e4190415b44f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.10.13

Release history Release notifications | RSS feed

This release

1.1.1 This release

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page