Skip to main content

pyccwebgraph: Python Interface to CommonCrawl Webgraph

PyPI version Python 3.8+ License: MIT

Discover related domains using link topology from CommonCrawl's webgraph.

Installation

Prerequisites:

pip install pyccwebgraph

First use downloads graph data:

from pyccwebgraph import CCWebgraph, get_available_versions

# List available versions
versions = get_available_versions()
print(versions[:3])  # ['cc-main-2024-nov-dec-jan', 'cc-main-2024-feb-apr-may', ...]

webgraph = CCWebgraph.setup(
    webgraph_dir="/data/my-webgraph",
    version="cc-main-2024-feb-apr-may"
)

# Find domains that link TO seeds (backlinks)
results = webgraph.discover_backlinks(
    seeds=["cnn.com", "bbc.com", "nytimes.com"],
    min_connections=3  # Must link to all seeds
)

print(f"Found {len(results['nodes'])} domains")
print(f"Top result: {results['nodes'][0]}")
# {'domain': 'news-aggregator.com', 'connections': 15, 'percentage': 50.0}

Working with NetworkX

# Get results as NetworkX graph
G = webgraph.discover_backlinks(
    seeds=["cnn.com", "bbc.com"],
    min_connections=2,
    format='networkx'  # Returns nx.DiGraph
)

# Run standard NetworkX algorithms
import networkx as nx

# Centrality analysis
pr = nx.pagerank(G)
bc = nx.betweenness_centrality(G)

# Community detection
from cdlib import algorithms
communities = algorithms.louvain(G)

# Visualization
from pyvis.network import Network
net = Network(notebook=True)
net.from_nx(G)
net.show("network.html")

Performance: Large Graphs with NetworKit

For large discovered subgraphs (>100K nodes), use NetworKit instead of NetworkX:

# Discover large subgraph
G_nk, name_map = webgraph.discover_backlinks(
    seeds=seed_list,
    min_connections=2,
    format='networkit'  # Returns NetworKit graph
)

CC-Webgraph mapping

# Check if domain exists in graph
vid = webgraph.domain_to_id("example.com")
if vid is not None:
    print(f"Found at vertex ID {vid}")

# Get all domains this domain links to
outlinks = webgraph.get_successors("cnn.com")
print(f"CNN links to {len(outlinks)} domains")

# Get all domains linking to this domain  
backlinks = webgraph.get_predecessors("cnn.com")
print(f"{len(backlinks)} domains link to CNN")

# Validate seeds before discovery
found, missing = webgraph.validate_seeds(["cnn.com", "fake-site.xyz"])
print(f"Found: {found}")
print(f"Missing: {missing}")

Metadata

Release files for pyccwebgraph 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyccwebgraph 0.3.1
File Size Uploaded
pyccwebgraph-0.3.1.tar.gz 17.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyccwebgraph 0.3.1
File Interpreter ABI Platform
pyccwebgraph-0.3.1-py3-none-any.whl Python 3 none any Details

Total release size: 34.0 kB

Release files / pyccwebgraph-0.3.1.tar.gz

Download URL pyccwebgraph-0.3.1.tar.gz
Size 17.9 kB
Tags Source
SHA-256 checksum
How to use checksums
b6a576c436de8c41d30e60763f10a5b5133fc9dbf0cf8f6dcd1cba2ab7be20e5
BLAKE2b-256 checksum
How to use checksums
b554690f20af0208fcebd54abbdbe8ee8ce33c71dca83ea9b0f87227f5d47841
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.7

Release files / pyccwebgraph-0.3.1-py3-none-any.whl

Download URL pyccwebgraph-0.3.1-py3-none-any.whl
Size 16.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7507f4f80a59f4a02c84f29651a4fee31c8968c786357c57f8392a1fd9d6dec1
BLAKE2b-256 checksum
How to use checksums
fc5361a6b61b5b2a5b5982d0f6246032d4ad2b95e74a8489e9e182fed7f309d1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.7

Release history Release notifications | RSS feed

This release

0.3.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page