commoncrawl-fsspec
An
fsspecfilesystem plugin that turns the Common Crawl web archive into a browseable virtual filesystem.
✨ Features
- Browse the Common Crawl archive like a filesystem — crawls, segments, and WARC files
- Read individual WARC records via HTTP byte-range requests
- Download files to disk with standard
fsspecoperations - Interactive CLI with a rich terminal browser for exploration
📦 Installation
pip install commoncrawl-fsspec
Optional extras
| Extra | Description |
|---|---|
examples |
Interactive CLI (click, rich) |
pip install commoncrawl-fsspec[examples]
🚀 Quick Start
Python API
import fsspec
# Create the filesystem
fs = fsspec.filesystem("cc")
# List top-level directories
fs.ls("/")
# ['/crawls']
# Browse available crawls
crawls = fs.ls("/crawls")
# ['/crawls/CC-MAIN-2026-25', '/crawls/CC-MAIN-2026-21', ...]
# Navigate into a crawl
segments = fs.ls("/crawls/CC-MAIN-2026-25/segments")
# ['/crawls/CC-MAIN-2026-25/segments/1780687572080.85', ...]
# List WARC files in a segment
files = fs.ls("/crawls/CC-MAIN-2026-25/segments/1780687572080.85/warc")
# ['/crawls/CC-MAIN-2026-25/segments/1780687572080.85/warc/CC-MAIN-20260605214811-00000.warc.gz', ...]
# Get file metadata
info = fs.info("/crawls/CC-MAIN-2026-25/segments/1780687572080.85/warc/CC-MAIN-20260605214811-00000.warc.gz")
# {'name': '...', 'type': 'file', 'size': 940246254, 'mtime': ...}
Interactive CLI
# Launch the interactive browser
python -m examples.cc interactive
# List a path
python -m examples.cc ls /crawls
📂 Virtual Filesystem
The plugin exposes a virtual filesystem for navigating the archive:
Archive Browser (/crawls)
Navigate the Common Crawl archive structure:
/
└── /crawls/ → List of crawls (from collinfo.json)
└── /crawls/{crawl-id}/
└── /crawls/{crawl-id}/segments/
└── /crawls/{crawl-id}/segments/{segment-id}/
├── /crawls/.../warc/ → WARC files
├── /crawls/.../wet/ → WET files
└── /crawls/.../wat/ → WAT files
📖 API Reference
CommonCrawlFileSystem
| Method | Description |
|---|---|
ls(path, detail=True) |
List directory contents |
info(path) |
Get file/directory metadata |
glob(pattern, detail=True) |
Glob for files matching pattern |
open(path, mode="rb") |
Open a file for reading |
cat_file(path, start, end) |
Read bytes from a file |
get_file(rpath, lpath) |
Download a file to disk |
Constructor Options
| Parameter | Default | Description |
|---|---|---|
cache_ttl |
86400 |
Cache TTL in seconds (24 hours) |
🏗️ Architecture
commoncrawl-fsspec/
├── filesystem.py # CommonCrawlFileSystem (fsspec orchestrator)
├── paths.py # Virtual path parsing and building
├── models.py # Data models (CrawlInfo, WarcFileInfo)
├── caching.py # TTL cache for crawl list
├── clients/
│ ├── http_client.py # HTTP client with retries
│ ├── crawl_index_client.py # Crawl discovery (collinfo.json)
│ ├── s3_listing_client.py # Crawl manifest browsing
│ └── warc_fetcher.py # WARC record byte-range fetching
🧪 Development
# Install dependencies
uv sync
# Run tests
uv run pytest
# Type checking
uv run mypy src
# Linting
uv run ruff check src
📝 License
MIT
Metadata
Release files for commoncrawl-fsspec 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| commoncrawl_fsspec-0.1.0.tar.gz | 8.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| commoncrawl_fsspec-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 21.3 kB
Release files / commoncrawl_fsspec-0.1.0.tar.gz
| Download URL | commoncrawl_fsspec-0.1.0.tar.gz |
|---|---|
| Size | 8.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f7c67b00b6747aea8225122dcd4750af444d2a624ea802a00065dc8f131472a4
|
|
BLAKE2b-256 checksum How to use checksums |
211bc3d1b5e3e2e3c47389dd5ef3e76d765dc04ee2066d6ec06cdfee847dab9a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency logRelease files / commoncrawl_fsspec-0.1.0-py3-none-any.whl
| Download URL | commoncrawl_fsspec-0.1.0-py3-none-any.whl |
|---|---|
| Size | 12.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d87820db6d822d2f20f5a0c06d94e1cce28896166364fb70620514de45d53090
|
|
BLAKE2b-256 checksum How to use checksums |
3d724c7b4d14d4f19d75930f0d148bbcea3af52e225d4cec57caa6e19a6014b7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency log