Skip to main content

TAU Community Detection

PyPI License: MIT Python 3.10+ Downloads Build Status Ruff

tau-community-detection implements TAU, an evolutionary community detection algorithm that couples genetic search with Leiden refinements. It is designed for scalable graph clustering with a simple drop-in run_clustering() API, sensible defaults, and multiprocessing support.


Highlights

  • Evolutionary search: Maintains a population of candidate partitions and applies crossover and mutation tailored for graph clustering.
  • Leiden optimization: Refines every candidate with Leiden to ensure modularity gains each generation.
  • Multiprocessing aware: Utilises parallel worker pools for population optimization with automatic fallback to sequential mode.
  • Fully reproducible: Pass random_seed to seed both TAU's numpy RNG and igraph's Leiden RNG — same seed always produces identical results.
  • Input flexibility: Accepts igraph.Graph, networkx.Graph, or a file path. Edge weights are auto-detected.
  • Simple API: Use run_clustering(graph) for zero-friction usage, or drop down to TauClustering + TauConfig for full control.

Installation

Requires Python 3.10 or newer.

pip install tau-community-detection

To work from a clone:

git clone https://github.com/HillelCharbit/TAU.git
cd TAU
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements-dev.txt
pip install -e .

Quick Start

import igraph as ig
from tau_community_detection import run_clustering

g = ig.Graph.Famous("Zachary")

# Zero-friction default usage
clustering = run_clustering(g)
print(f"Communities: {len(clustering)},  Modularity: {clustering.modularity:.4f}")

# Override only the knobs you care about
clustering = run_clustering(
    g,
    resolution=0.8,
    random_seed=42,
    verbose=True,
    population_size=100,
    max_generations=50,
)

run_clustering() returns an igraph.VertexClustering, so .membership, .modularity, and all standard igraph attributes are available immediately.

NetworkX input

import networkx as nx
from tau_community_detection import run_clustering

g = nx.erdos_renyi_graph(n=500, p=0.02, seed=0)
clustering = run_clustering(g)

Scanpy / AnnData integration

Install TAU with its optional Scanpy dependencies:

pip install "tau-community-detection[scanpy]"

TAU can cluster an existing Scanpy neighbor graph and store the results directly in an AnnData object:

import scanpy as sc
from tau_community_detection.tl import tau

adata = sc.datasets.pbmc3k_processed()

# Compute a neighbor graph if the AnnData object does not already contain one.
if "connectivities" not in adata.obsp:
    sc.pp.neighbors(adata)

tau(
    adata,
    key_added="tau",
    resolution=1.0,
    random_state=42,
    population_size=60,
    max_generations=20,
    worker_count=1,
    verbose=True,
)

print(adata.obs["tau"])
print(adata.uns["tau"]["modularity"])

The cluster assignments are stored as categorical values in adata.obs["tau"]. The parameters and resulting modularity are stored in adata.uns["tau"].

Custom Scanpy graphs can be selected with neighbors_key or obsp:

tau(
    adata,
    neighbors_key="custom_neighbors",
    key_added="tau_custom",
)

The function also supports copy=True, explicit adjacency matrices, and weighted or unweighted graph clustering.

Advanced usage with TauClustering

For full control over the lifecycle — including reusing the worker pool across multiple runs:

from tau_community_detection import TauClustering, TauConfig

config = TauConfig(
    population_size=60,
    max_generations=20,
    resolution=1.0,
    elite_fraction=0.15,
    immigrant_fraction=0.2,
    stopping_generations=10,
    random_seed=42,
    verbose=True,
)

with TauClustering(g, config=config) as tau:
    clustering, stats = tau.run(track_stats=True)

print(f"Ran for {len(stats)} generations")
print(f"Final modularity: {clustering.modularity:.4f}")

track_stats=True returns a list of per-generation dicts with keys generation, top_fitness, average_fitness, time_per_generation, convergence, elite_runtime, crossover_runtime.


Graph Input

Supported sources:

Type Notes
igraph.Graph Passed directly; weights auto-detected from "weight" edge attribute
networkx.Graph Converted internally; weights auto-detected
str (file path) Edgelist/NCOL (.graph, .edgelist, .txt) or adjacency list (.adjlist)

For large graphs or high worker counts, passing a file path is recommended — it avoids serialising the graph object across worker processes.

Edge weights are detected automatically. To override:

from tau_community_detection import TauConfig
config = TauConfig(is_weighted=False)   # force unweighted even if file has weights

Configuration Reference

All hyperparameters live on TauConfig. Every field is validated on construction — invalid values raise ValueError immediately.

Parameter Default Valid range Description
population_size 60 > 0 Number of candidate partitions per generation
max_generations 20 > 0 Hard cap on evolutionary iterations
worker_count None ≥ 1 Parallel workers (default: CPU count, capped by population size)
elite_fraction 0.1 (0, 1] Fraction of best partitions preserved each generation
immigrant_fraction 0.15 (0, 1] Fraction of fresh random partitions injected each generation
selection_power 5 > 0 Sharpness of fitness-proportional parent selection
elite_similarity_threshold 0.9 [0, 1] Jaccard threshold below which two elites are considered diverse
stopping_generations 10 > 0 Generations without improvement before early stopping
stopping_jaccard 0.98 [0, 1] Similarity threshold that counts as "no improvement"
n_iterations 3 > 0 Leiden iterations per fitness evaluation
resolution 1.0 > 0 Leiden resolution — higher values produce more, smaller communities
sample_fraction_range (0.2, 0.9) 0 < low ≤ high ≤ 1 Range for random subgraph sampling during population init
is_weighted None bool or None Override weight auto-detection (None = auto)
sim_sample_size 20 000 int or None Node sample size for Jaccard similarity (None = all nodes)
random_seed None int or None Seeds both numpy and igraph's Leiden RNG for fully deterministic results
verbose False bool Log progress to the standard Python logger

run_clustering() exposes the most common parameters directly. For any other TauConfig field, use TauClustering with a TauConfig directly:

from tau_community_detection import TauClustering, TauConfig

config = TauConfig(elite_fraction=0.2, stopping_generations=5)
with TauClustering(g, config=config) as t:
    clustering = t.run()

Development

pip install -r requirements-dev.txt
pip install -e .
make lint     # ruff checks
make test     # pytest
make coverage # pytest + coverage report
make build    # build sdist + wheel

Continuous Integration

GitHub Actions runs lint, tests (Python 3.10 and 3.11), and a package build on every push and pull request. Set the CODECOV_TOKEN secret to upload coverage reports.

Publishing

  1. Bump version in setup.cfg and commit.
  2. Tag the release: git tag vX.Y.Z && git push --tags.
  3. Run the Publish Package workflow. Use TEST_PYPI_API_TOKEN for a dry run on TestPyPI, or PYPI_API_TOKEN to publish to PyPI.

Reference & Citation

If you use TAU in your research, please cite:

From Leiden to Tel-Aviv University (TAU): exploring clustering solutions via a genetic algorithm Gal Gilad and Roded Sharan. PNAS Nexus, Volume 2, Issue 6, June 2023. DOI: 10.1093/pnasnexus/pgad180

@article{gilad2023tau,
  title={From Leiden to Tel-Aviv University (TAU): exploring clustering solutions via a genetic algorithm},
  author={Gilad, Gal and Sharan, Roded},
  journal={PNAS Nexus},
  volume={2},
  number={6},
  pages={pgad180},
  year={2023},
  publisher={Oxford University Press}
}

License

MIT License © 2023 Hillel Charbit

Release files for tau-community-detection 1.4.7

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tau-community-detection 1.4.7
File Size Uploaded
tau_community_detection-1.4.7.tar.gz 37.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tau-community-detection 1.4.7
File Interpreter ABI Platform
tau_community_detection-1.4.7-py3-none-any.whl Python 3 none any Details

Total release size: 59.8 kB

Release files / tau_community_detection-1.4.7.tar.gz

Download URL tau_community_detection-1.4.7.tar.gz
Size 37.4 kB
Tags Source
SHA-256 checksum
How to use checksums
f4089928b7303f8b1baf0003309531595ea45f6ed01bd00c92d444f238cfc4ac
BLAKE2b-256 checksum
How to use checksums
1021c5f442491f05ed86f32a0765ad499af1bcf5fa9447dbea15ca313615f05e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 14, 2026.

Transparency log

Release files / tau_community_detection-1.4.7-py3-none-any.whl

Download URL tau_community_detection-1.4.7-py3-none-any.whl
Size 22.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
43aefa729e338aa7591cd6259912b82e6e52ce78b4d7aaa2545623cfb8dc0f5e
BLAKE2b-256 checksum
How to use checksums
e6e8b27e19ecdc8430b83af42d97d5c0af9dcb0a6837ac0e89b76d9d26585139
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 14, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.4.7 This release

2 release files

1.4.6

2 release files

1.4.5

2 release files

1.4.4

2 release files

1.4.3

2 release files

1.4.2

2 release files

1.4.1

2 release files

1.4.0

2 release files

1.3.3

2 release files

1.3.2

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.11

2 release files

1.2.9

2 release files

1.2.8

2 release files

1.2.7

2 release files

1.2.6

2 release files

1.2.5

2 release files

1.2.4

2 release files

1.2.3

2 release files

1.2.2

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.9

2 release files

1.1.8

2 release files

1.1.7

2 release files

1.1.6

2 release files

1.1.5

2 release files

1.1.4

2 release files

1.1.3

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.3.23

2 release files

0.3.22

2 release files

0.3.21

2 release files

0.3.20

2 release files

0.3.19

2 release files

0.3.18

2 release files

0.3.17

2 release files

0.3.16

2 release files

0.3.15

2 release files

0.3.14

2 release files

0.3.13

2 release files

0.3.12

2 release files

0.3.11

2 release files

0.3.9

2 release files

0.3.7

2 release files

0.3.6

2 release files

0.3.5

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page