Skip to main content

commoncrawl-fsspec

An fsspec filesystem plugin that turns the Common Crawl web archive into a browseable virtual filesystem.

Python >=3.10 License: MIT

✨ Features

  • Browse the Common Crawl archive like a filesystem — crawls, segments, and WARC files
  • Read individual WARC records via HTTP byte-range requests
  • Download files to disk with standard fsspec operations
  • Interactive CLI with a rich terminal browser for exploration

📦 Installation

pip install commoncrawl-fsspec

Optional extras

Extra Description
examples Interactive CLI (click, rich)
pip install commoncrawl-fsspec[examples]

🚀 Quick Start

Python API

import fsspec

# Create the filesystem
fs = fsspec.filesystem("cc")

# List top-level directories
fs.ls("/")
# ['/crawls']

# Browse available crawls
crawls = fs.ls("/crawls")
# ['/crawls/CC-MAIN-2026-25', '/crawls/CC-MAIN-2026-21', ...]

# Navigate into a crawl
segments = fs.ls("/crawls/CC-MAIN-2026-25/segments")
# ['/crawls/CC-MAIN-2026-25/segments/1780687572080.85', ...]

# List WARC files in a segment
files = fs.ls("/crawls/CC-MAIN-2026-25/segments/1780687572080.85/warc")
# ['/crawls/CC-MAIN-2026-25/segments/1780687572080.85/warc/CC-MAIN-20260605214811-00000.warc.gz', ...]

# Get file metadata
info = fs.info("/crawls/CC-MAIN-2026-25/segments/1780687572080.85/warc/CC-MAIN-20260605214811-00000.warc.gz")
# {'name': '...', 'type': 'file', 'size': 940246254, 'mtime': ...}

Interactive CLI

# Launch the interactive browser
python -m examples.cc interactive

# List a path
python -m examples.cc ls /crawls

📂 Virtual Filesystem

The plugin exposes a virtual filesystem for navigating the archive:

Archive Browser (/crawls)

Navigate the Common Crawl archive structure:

/
└── /crawls/                          → List of crawls (from collinfo.json)
    └── /crawls/{crawl-id}/
        └── /crawls/{crawl-id}/segments/
            └── /crawls/{crawl-id}/segments/{segment-id}/
                ├── /crawls/.../warc/   → WARC files
                ├── /crawls/.../wet/    → WET files
                └── /crawls/.../wat/    → WAT files

📖 API Reference

CommonCrawlFileSystem

Method Description
ls(path, detail=True) List directory contents
info(path) Get file/directory metadata
glob(pattern, detail=True) Glob for files matching pattern
open(path, mode="rb") Open a file for reading
cat_file(path, start, end) Read bytes from a file
get_file(rpath, lpath) Download a file to disk

Constructor Options

Parameter Default Description
cache_ttl 86400 Cache TTL in seconds (24 hours)

🏗️ Architecture

commoncrawl-fsspec/
├── filesystem.py          # CommonCrawlFileSystem (fsspec orchestrator)
├── paths.py               # Virtual path parsing and building
├── models.py              # Data models (CrawlInfo, WarcFileInfo)
├── caching.py             # TTL cache for crawl list
├── clients/
│   ├── http_client.py     # HTTP client with retries
│   ├── crawl_index_client.py  # Crawl discovery (collinfo.json)
│   ├── s3_listing_client.py   # Crawl manifest browsing
│   └── warc_fetcher.py    # WARC record byte-range fetching

🧪 Development

# Install dependencies
uv sync

# Run tests
uv run pytest

# Type checking
uv run mypy src

# Linting
uv run ruff check src

📝 License

MIT

Metadata

Release files for commoncrawl-fsspec 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for commoncrawl-fsspec 0.1.1
File Size Uploaded
commoncrawl_fsspec-0.1.1.tar.gz 8.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for commoncrawl-fsspec 0.1.1
File Interpreter ABI Platform
commoncrawl_fsspec-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 21.3 kB

Release files / commoncrawl_fsspec-0.1.1.tar.gz

Download URL commoncrawl_fsspec-0.1.1.tar.gz
Size 8.4 kB
Tags Source
SHA-256 checksum
How to use checksums
d8837a7e215dcbbd7626319fd8c474be6d101d0ba43d289c575eb2fa6e30958d
BLAKE2b-256 checksum
How to use checksums
5717fc8b012e7bbf767beb6ce4c4de9ba9c45009ecef22035637fff5780434e1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.

Transparency log

Release files / commoncrawl_fsspec-0.1.1-py3-none-any.whl

Download URL commoncrawl_fsspec-0.1.1-py3-none-any.whl
Size 12.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
534d1ef78e84dfcff07b97a6b5d457611ec9face1486be31708e3c11940348a4
BLAKE2b-256 checksum
How to use checksums
bdb822bfa75557749a3028cd664c23a92d0d520074ea2e29e45df67d731f7821
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page