Claim-level narrative clustering with HNSW + Leiden + neighbor-vote assignment
Project description
narrativecluster
Claim-level narrative clustering for large social-media corpora.
Stack: HNSW (FAISS) → Leiden community detection → neighbour-vote cluster assignment
Why this approach?
Clustering millions of social-media posts is an unsupervised problem with no clean ground truth, no fixed number of narratives, and a distribution that shifts every week as new stories emerge. Most standard algorithms break down in one of a few predictable ways.
K-means and its variants require you to specify k up front. For a corpus spanning thousands of distinct claims across a multi-year election cycle, the right k is unknowable in advance — and the algorithm will happily partition noise into fake clusters to fill whatever k you gave it.
Gaussian Mixture Models and DP-Means are theoretically appealing (they can infer k from the data) but brittle at scale. DP-Means is sensitive to its bandwidth parameter in high-dimensional embedding space, where "distance" is dominated by the curse of dimensionality rather than semantic content, and it tends to either over-fragment coherent narratives or lump disparate ones together depending on how that single parameter lands. GMMs have similar issues and become computationally expensive as k grows into the thousands.
DBSCAN / HDBSCAN handle arbitrary cluster shapes well on low-dimensional data but struggle in the intrinsic geometry of dense embedding spaces — nearly every point becomes a "core point" or noise depending on a single eps that is nearly impossible to calibrate in cosine space. They also have no natural mechanism for assigning new documents to existing clusters without re-running the full algorithm.
What narrativecluster does instead
The approach decouples structure discovery (Leiden on a cosine-thresholded KNN graph) from point assignment (neighbour-vote), which buys three things simultaneously: no fixed k, graceful handling of noise, and fast incremental updates.
Structure discovery via graph community detection. Rather than treating each post as a point in metric space, we build a sparse k-NN similarity graph where an edge only exists if two posts exceed a cosine similarity threshold (tau_graph). Leiden community detection then finds densely-connected subgraphs — natural narrative clusters — without any assumption about their number, shape, or size. Leiden is resolution-parameterised, which is semantically interpretable: higher resolution → finer-grained narratives, lower → coarser topical buckets. Critically, isolated posts and borderline cases simply have no edges and fall into small singleton clusters rather than being force-assigned to the nearest centroid.
Neighbour-vote assignment for new data. Once the graph structure is learned, assigning a new post is cheap and explainable: query the HNSW index for its nearest neighbours, filter to neighbours above tau_edge, and take a plurality vote over their cluster memberships. The vote fraction and margin give a per-post confidence score, so you can dial your own precision/recall tradeoff after fitting — without touching the model. This is O(log N) per new post rather than O(N²).
A concrete example. Suppose you have five posts:
A: "Ballots were found in a dumpster in Maricopa County — proof of fraud"
B: "Election officials in Arizona caught destroying ballots"
C: "The mRNA vaccine is causing myocarditis in young athletes"
D: "Pfizer trial data shows spike protein accumulates in organs"
E: "My dog had a great walk this morning :)"
- K-means with k=2 groups {A,B,C,D} vs {E} — conflating two distinct narratives because it must fill both centroids.
- DP-Means may correctly split {A,B} from {C,D}, but will almost certainly pull post E into whichever centroid is geometrically closest rather than leaving it unclaimed.
- narrativecluster: A↔B have high cosine similarity and form one Leiden community; C↔D form another; E has no neighbours above
tau_graphand lands in a singleton — it is simply not assigned when you callpredict(). That is the correct answer for noise.
Speed. HNSW graph construction is O(N log N) and approximate-neighbour search at prediction time is effectively O(log N). On a 12M-post corpus, fit() takes ~20 minutes on a single CPU; predict() on a new 100K-post batch takes under 2 minutes. DP-Means or a full GMM on the same corpus would be intractable.
Install
pip install narrativecluster
# GPU variant (requires faiss-gpu):
pip install "narrativecluster[gpu]"
From source:
git clone https://github.com/patrikgerard/narrativecluster
cd narrativecluster
pip install -e ".[dev]"
Quick start
from narrativecluster import NarrativeClusterer, ClusterConfig
cfg = ClusterConfig(
embed_col="embedding", # column holding np.ndarray / list embeddings
text_col="text",
site_col="platform",
)
model = NarrativeClusterer(config=cfg, checkpoint_dir="./checkpoints")
# Full fit — idempotent: loads from checkpoint if it exists
model.fit(df_train)
# Incremental update — adds new rows, re-runs Leiden, keeps cluster space stable
model.partial_fit(df_new_batch)
# Neighbour-vote assignment for any DataFrame
df_pred = model.predict(df_query)
# df_pred now has: nv_pred_cluster, nv_accepted, nv_vote_frac, …
# Per-site acceptance stats
global_rate, site_tbl = model.site_stats(df_pred)
print(site_tbl)
# Save / load
model.save("./my_model")
model2 = NarrativeClusterer.load("./my_model", config=cfg)
CLI
# Full fit + predict
narrativecluster --raw raw_posts.parquet --out ./checkpoints
# Incremental update
narrativecluster --dedup ./checkpoints/df_unique.parquet \
--partial_fit new_posts.parquet \
--out ./checkpoints
# Force re-predict
narrativecluster --dedup ./checkpoints/df_unique.parquet \
--out ./checkpoints \
--force_repredict
partial_fit semantics
partial_fit is designed for streaming / incremental ingestion:
- New rows are deduplicated against themselves and all previously seen texts.
- They are mean-centred using the original µ (the embedding space is not re-centred — this keeps the geometry stable across batches).
- New vectors are added to the live HNSW index.
- KNN arrays are extended; Leiden is re-run on the combined graph.
Note: Cluster IDs may shift after each
partial_fitbecause Leiden is non-deterministic and the graph changes. If you need stable IDs across incremental updates, pinrandom_seedand store a cluster-label mapping separately.
Configuration
All hyper-parameters live in ClusterConfig:
| Parameter | Default | Description |
|---|---|---|
embed_col |
"embedding_aligned" |
DataFrame column with embeddings |
hnsw_m |
32 | HNSW M parameter |
ef_construction |
200 | HNSW efConstruction |
ef_search |
128 | HNSW efSearch |
k_graph |
64 | KNN neighbours for Leiden graph |
tau_graph |
0.68 | Cosine-sim threshold for graph edges |
do_mutual |
False | Keep only mutual edges |
cap_m |
20 | Max out-degree per node (None = uncapped) |
leiden_resolution |
1.0 | Leiden resolution parameter |
kq |
50 | Neighbours queried at prediction time |
tau_edge |
0.68 | Min sim for a vote to count |
min_cluster_size |
4 | Clusters smaller than this are ineligible |
vote_margin |
2 | Min vote margin (top − second) to accept |
Output columns (predict)
| Column | Type | Description |
|---|---|---|
nv_pred_cluster |
int32 | Predicted Leiden cluster id (-1 = rejected) |
nv_accepted |
bool | Whether the vote passed quality thresholds |
nv_votes_top |
int16 | Votes for winning cluster |
nv_votes_tot |
int16 | Total eligible neighbour votes |
nv_vote_frac |
float32 | Top votes / total votes |
nv_vote_margin |
float32 | top_votes − second_votes |
nv_best_sim |
float32 | Max cosine similarity among winning-cluster neighbours |
License
MIT
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file narrativecluster-0.1.9.tar.gz.
File metadata
- Download URL: narrativecluster-0.1.9.tar.gz
- Upload date:
- Size: 20.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
14f0e01012040e75a3646051b424a89952cf33eff4dc8dff33e4cbc744d30234
|
|
| MD5 |
dd9d99fe5f9ebdd58b5475628aafbc82
|
|
| BLAKE2b-256 |
a3a737bdf918404949fd11fb2dad182b9e20858687501d4ff62605ee0dc8551d
|
Provenance
The following attestation bundles were made for narrativecluster-0.1.9.tar.gz:
Publisher:
publish.yml on patrikgerard/narrativecluster
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
narrativecluster-0.1.9.tar.gz -
Subject digest:
14f0e01012040e75a3646051b424a89952cf33eff4dc8dff33e4cbc744d30234 - Sigstore transparency entry: 1053637029
- Sigstore integration time:
-
Permalink:
patrikgerard/narrativecluster@eae3e44d336af83ee3eae8556fbd10d737f23e2a -
Branch / Tag:
refs/tags/v0.1.9 - Owner: https://github.com/patrikgerard
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@eae3e44d336af83ee3eae8556fbd10d737f23e2a -
Trigger Event:
push
-
Statement type:
File details
Details for the file narrativecluster-0.1.9-py3-none-any.whl.
File metadata
- Download URL: narrativecluster-0.1.9-py3-none-any.whl
- Upload date:
- Size: 17.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
000527c7f597cb3601f09bb171d33b6c2ede04a2c2444294fb724a17871c1704
|
|
| MD5 |
4dc96b3d2d1bb3458402ae24aa128bbe
|
|
| BLAKE2b-256 |
b96c63aa6f78f9e2740538077e105ee8b7697968f821929e9c5a1b6214f4f2e9
|
Provenance
The following attestation bundles were made for narrativecluster-0.1.9-py3-none-any.whl:
Publisher:
publish.yml on patrikgerard/narrativecluster
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
narrativecluster-0.1.9-py3-none-any.whl -
Subject digest:
000527c7f597cb3601f09bb171d33b6c2ede04a2c2444294fb724a17871c1704 - Sigstore transparency entry: 1053637102
- Sigstore integration time:
-
Permalink:
patrikgerard/narrativecluster@eae3e44d336af83ee3eae8556fbd10d737f23e2a -
Branch / Tag:
refs/tags/v0.1.9 - Owner: https://github.com/patrikgerard
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@eae3e44d336af83ee3eae8556fbd10d737f23e2a -
Trigger Event:
push
-
Statement type: