WARC file format library
Project description
PyWarc
Python library for WebArchive's WARC file format manipulation.
How to read a warc file:
from pywarc import WarcReader
warc = WarcReader("my_archive.warc")
# you can also provide your own fp
# warc = WarcReader(sys.stdin.buffer)
blk = warc.get_next_block() # read a block
print(blk.content_length) # get content length
print(blk.headers) # read headers
print(blk.read(10)) # read x bytes
print(blk.read()) # read everything
blk = warc.get_next_block() # read next block
# note that you won't be able to read the previous block anymore
# if the file is not seekable.
print(blk.content_length) # get content length
print(blk.headers) # read headers
# you can also get the content as a stream
stream = blk.get_as_stream()
print(stream.readline()) # for easier manipulation
# or you can even use a for loop to iterate on each blocks
for blk in warc:
print(blk.headers["WARC-Record-ID"])
How to write a warc file:
from pywarc import WarcWriter
from datetime import datetime
import os
# you can start to write as easy as:
# warc = WarcWriter("my_archive.warc")
# but you have more options:
warc = WarcWriter(
"my_archive.warc",
truncate=True, # by default it will appends, but you can choose to truncate it.
software_name="my_program",
software_version="1.0.0",
warc_meta={
"My-Super-Meta": "Hello, World!", # you can set meta to the warcinfo's block.
"conformsTo": None}) # or delete default ones with None.
warc.write_block(
"response",
b"HTTP/1.1 200 OK\r\nHost: example.com\r\nContent-Type: 0\r\n\r\n",
# optional fields:
record_date=datetime.fromisocalendar(2000, 1, 1), # by default it will be set to the current UTC time.
record_id="urn:custom:i_dont_know", # your record identifier, NOTE: it _must_ be a valid URI. (default: uuid.uuid4().urn)
# You can add custom headers to your block,
# Please, try to conform to the IANA's assignments:
# https://www.iana.org/assignments/message-headers/message-headers.xhtml
# https://www.iana.org/assignments/http-fields/http-fields.xhtml
record_headers={"Content-Type": "application/http;msgtype=response"}) # custom meta (please try to respect IANA)
# if your object is too big to be in memory
# you can write it in several calls:
with open("hypothetical_file.txt", "rb") as fp:
size = fp.seek(0, os.SEEK_END)
fp.seek(0, os.SEEK_SET)
# please note that this implies that if an error occurs at this point,
# the archive will be truncated and so if you reopen this file for writing,
# all subsequent data will be corrupted.
warc.start_block("my_custom_type", size)
for l in fp:
warc.write_block_body(l)
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distributions
No source distribution files available for this release.See tutorial on generating distribution archives.
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
pywarc-0.1.3-py3-none-any.whl
(21.3 kB
view details)
File details
Details for the file pywarc-0.1.3-py3-none-any.whl.
File metadata
- Download URL: pywarc-0.1.3-py3-none-any.whl
- Upload date:
- Size: 21.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/4.0.2 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0d6e6c4895dd631f110c557e146ae029e0b9e32e6aa5f2c00bbb8bcb758d520f
|
|
| MD5 |
1b63333cf57e7f9789eb5845eb5eb43f
|
|
| BLAKE2b-256 |
adf56bc93ae6973a068a1963156f5a06f2f2cdc843b0bfcfcd91e8c7d5cff5a6
|