Skip to main content
This is a pre-production deployment of Warehouse. Changes made here affect the production instance of PyPI (
Help us improve Python packaging - Donate today!

Python 3 library to work with ARC and WARC files

Project Description

Note: This is a fork of the original (now dead) warc repository.

WARC (Web ARChive) is a file format for storing web crawls.

This warc library makes it very easy to work with WARC files.:

import warc
with"test.warc") as f:
    for record in f:
        print record['WARC-Target-URI'], record['Content-Length']


The documentation of the warc library is available at

Apart from the install from pip, which will not work for this warc3 version, the interface as described there is unchanged.


This software is licensed under GPL v2. See LICENSE file for details.


Original Python2 Versions:

  • Anand Chitipothu
  • Noufal Ibrahim

Python3 Port:

  • Ryan Chartier
  • Jan Pieter Bruins Slot
  • Almer S. Tigelaar

Release History

This version
History Node


Download Files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Filename, Size & Hash SHA256 Hash Help File Type Python Version Upload Date
(17.2 kB) Copy SHA256 Hash SHA256
Source None May 28, 2017

Supported By

Elastic Elastic Search Pingdom Pingdom Monitoring Dyn Dyn DNS Sentry Sentry Error Logging CloudAMQP CloudAMQP RabbitMQ Kabu Creative Kabu Creative UX & Design Fastly Fastly CDN DigiCert DigiCert EV Certificate Google Google Cloud Servers