Skip to main content

commoncrawl-fsspec

An fsspec filesystem plugin that turns the Common Crawl web archive into a browseable virtual filesystem.

Python >=3.10 License: MIT

✨ Features

  • Browse the Common Crawl archive like a filesystem — crawls, segments, and WARC files
  • Read individual WARC records via HTTP byte-range requests
  • Download files to disk with standard fsspec operations
  • Interactive CLI with a rich terminal browser for exploration

📦 Installation

pip install commoncrawl-fsspec

Optional extras

Extra Description
examples Interactive CLI (click, rich)
pip install commoncrawl-fsspec[examples]

🚀 Quick Start

Python API

import fsspec

# Create the filesystem
fs = fsspec.filesystem("cc")

# List top-level directories
fs.ls("/")
# ['/crawls']

# Browse available crawls
crawls = fs.ls("/crawls")
# ['/crawls/CC-MAIN-2026-25', '/crawls/CC-MAIN-2026-21', ...]

# Navigate into a crawl
segments = fs.ls("/crawls/CC-MAIN-2026-25/segments")
# ['/crawls/CC-MAIN-2026-25/segments/1780687572080.85', ...]

# List WARC files in a segment
files = fs.ls("/crawls/CC-MAIN-2026-25/segments/1780687572080.85/warc")
# ['/crawls/CC-MAIN-2026-25/segments/1780687572080.85/warc/CC-MAIN-20260605214811-00000.warc.gz', ...]

# Get file metadata
info = fs.info("/crawls/CC-MAIN-2026-25/segments/1780687572080.85/warc/CC-MAIN-20260605214811-00000.warc.gz")
# {'name': '...', 'type': 'file', 'size': 940246254, 'mtime': ...}

Interactive CLI

# Launch the interactive browser
python -m examples.cc interactive

# List a path
python -m examples.cc ls /crawls

📂 Virtual Filesystem

The plugin exposes a virtual filesystem for navigating the archive:

Archive Browser (/crawls)

Navigate the Common Crawl archive structure:

/
└── /crawls/                          → List of crawls (from collinfo.json)
    └── /crawls/{crawl-id}/
        └── /crawls/{crawl-id}/segments/
            └── /crawls/{crawl-id}/segments/{segment-id}/
                ├── /crawls/.../warc/   → WARC files
                ├── /crawls/.../wet/    → WET files
                └── /crawls/.../wat/    → WAT files

📖 API Reference

CommonCrawlFileSystem

Method Description
ls(path, detail=True) List directory contents
info(path) Get file/directory metadata
glob(pattern, detail=True) Glob for files matching pattern
open(path, mode="rb") Open a file for reading
cat_file(path, start, end) Read bytes from a file
get_file(rpath, lpath) Download a file to disk

Constructor Options

Parameter Default Description
cache_ttl 86400 Cache TTL in seconds (24 hours)

🏗️ Architecture

commoncrawl-fsspec/
├── filesystem.py          # CommonCrawlFileSystem (fsspec orchestrator)
├── paths.py               # Virtual path parsing and building
├── models.py              # Data models (CrawlInfo, WarcFileInfo)
├── caching.py             # TTL cache for crawl list
├── clients/
│   ├── http_client.py     # HTTP client with retries
│   ├── crawl_index_client.py  # Crawl discovery (collinfo.json)
│   ├── s3_listing_client.py   # Crawl manifest browsing
│   └── warc_fetcher.py    # WARC record byte-range fetching

🧪 Development

# Install dependencies
uv sync

# Run tests
uv run pytest

# Type checking
uv run mypy src

# Linting
uv run ruff check src

📝 License

MIT

Metadata

Release files for commoncrawl-fsspec 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for commoncrawl-fsspec 0.1.0
File Size Uploaded
commoncrawl_fsspec-0.1.0.tar.gz 8.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for commoncrawl-fsspec 0.1.0
File Interpreter ABI Platform
commoncrawl_fsspec-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 21.3 kB

Release files / commoncrawl_fsspec-0.1.0.tar.gz

Download URL commoncrawl_fsspec-0.1.0.tar.gz
Size 8.4 kB
Tags Source
SHA-256 checksum
How to use checksums
f7c67b00b6747aea8225122dcd4750af444d2a624ea802a00065dc8f131472a4
BLAKE2b-256 checksum
How to use checksums
211bc3d1b5e3e2e3c47389dd5ef3e76d765dc04ee2066d6ec06cdfee847dab9a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.

Transparency log

Release files / commoncrawl_fsspec-0.1.0-py3-none-any.whl

Download URL commoncrawl_fsspec-0.1.0-py3-none-any.whl
Size 12.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d87820db6d822d2f20f5a0c06d94e1cce28896166364fb70620514de45d53090
BLAKE2b-256 checksum
How to use checksums
3d724c7b4d14d4f19d75930f0d148bbcea3af52e225d4cec57caa6e19a6014b7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.2

2 release files

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page