k-LLMmeans clustering algorithm
Project description
k-llmmeans
Scikit-learn compatible implementation of k-LLMmeans for text clustering with summary-based centroids.
This package adapts the original research code into an estimator API you can use with familiar fit, predict, and fit_predict workflows.
- Original implementation: jairoadiazr/k-LLMmeans
- Paper: Summaries as Centroids for Interpretable and Scalable Text Clustering (arXiv:2502.09667)
What This Package Provides
kLLMmeansestimator implementingBaseEstimator+ClusterMixin- scikit-learn style methods:
fit(X)predict(X)fit_predict(X)
- configurable document embedding function (
embedding_fn) - configurable cluster summarization function (
summarizer_fn) or LiteLLM-backed LLM summarization - optional per-cluster sampling before summarization to keep prompts bounded on large clusters
- optional precomputed embedding support for faster iterative experimentation
Installation
pip install k-llmmeans
Or from source:
pip install -e .
Quick Start
from k_llmmeans import kLLMmeans
docs = [
"How to optimize SQL queries for large tables?",
"What is the best way to tune a random forest model?",
"PostgreSQL index strategy for analytics workloads",
"Cross-validation tips for imbalanced classification",
]
model = kLLMmeans(
n_clusters=2,
llm="openai/gpt-4o-mini",
max_llm_iter=5,
random_state=0,
)
labels = model.fit_predict(docs)
print(labels)
print(model.summaries_) # human-readable cluster summaries
LiteLLM Configuration
Pass a model string (any LiteLLM-supported model) or a dict of kwargs forwarded to litellm.completion:
model = kLLMmeans(
n_clusters=2,
llm={
"model": "openai/gpt-4o-mini",
"temperature": 0.2,
},
)
You can also set LITELLM_MODEL or OPENAI_MODEL in the environment instead of passing llm.
Using Custom Embeddings and Summarization
You can fully control both the embedding and summarization steps:
from sentence_transformers import SentenceTransformer
from k_llmmeans import kLLMmeans
encoder = SentenceTransformer("all-MiniLM-L6-v2")
def embedding_fn(texts: list[str]):
return encoder.encode(texts)
def summarizer_fn(cluster_texts: list[str]) -> str:
# Replace with your own deterministic or LLM summarizer
return " | ".join(cluster_texts[:2])
model = kLLMmeans(
n_clusters=3,
embedding_fn=embedding_fn,
summarizer_fn=summarizer_fn,
)
model.fit(["text a", "text b", "text c", "text d"])
Sampling Cluster Documents for Summaries
The paper's few-shot k-LLMmeans variant summarizes at most m representative documents from each cluster instead of sending the full cluster to the LLM. Use cluster_sample_size for this prompt budget:
model = kLLMmeans(
n_clusters=3,
llm="openai/gpt-4o-mini",
cluster_sample_size=10,
cluster_sampling_strategy="kmeans++",
)
Set cluster_sample_size=None to summarize every document assigned to each cluster. Supported sampling strategies are kmeans++, random, centroid, and edge. The selected document indices for each LLM iteration are stored in sampled_indices_evolution_.
Debugging Runs
Use show_progress=True to display progress bars for LLM iterations and cluster summarization. Use verbose=True to print timing, input character counts, sampled document counts, summary lengths, token usage when available, and convergence details:
model = kLLMmeans(
n_clusters=3,
llm="openai/gpt-4o-mini",
cluster_sample_size=10,
show_progress=True,
verbose=True,
)
After fitting, inspect summaries_evolution_ to see how cluster summaries changed, centroids_evolution_ to inspect summary-derived centroid embeddings, and sampled_indices_evolution_ to trace which input texts were used for each cluster summary. summary_workers can parallelize cluster summarization when your LLM provider and rate limits allow concurrent requests.
API Notes
- Input
Xshould belist[str]. - The estimator stores standard fitted attributes such as:
labels_cluster_centers_n_iter_
- Additional clustering interpretability attributes:
summaries_summary_embeddings_summaries_evolution_centroids_evolution_sampled_indices_evolution_
Citation
If you use this package in research or production work, please cite the original paper:
@article{diazrodriguez2025summaries,
title={Summaries as Centroids for Interpretable and Scalable Text Clustering},
author={Diaz-Rodriguez, Jairo},
journal={arXiv preprint arXiv:2502.09667},
year={2025}
}
Paper URL: https://arxiv.org/abs/2502.09667
Acknowledgment
This package is a scikit-learn compatible adaptation of the original project: https://github.com/jairoadiazr/k-LLMmeans
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file k_llmmeans-0.3.0.tar.gz.
File metadata
- Download URL: k_llmmeans-0.3.0.tar.gz
- Upload date:
- Size: 8.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
679fe10c9427a98c80148423a4de2adf0ebe205084515a84d94fde3760c865cb
|
|
| MD5 |
5fa3fb8e9e2ade96a33326b84e8cfda7
|
|
| BLAKE2b-256 |
2afa9a13ace97f941dcb889d71802a777a79957a3f74c1ba478892276e39dcb6
|
File details
Details for the file k_llmmeans-0.3.0-py3-none-any.whl.
File metadata
- Download URL: k_llmmeans-0.3.0-py3-none-any.whl
- Upload date:
- Size: 9.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7658cc7fa99e966e422313177895c6fee4f75d89facfbf1d691f9696b18ea90e
|
|
| MD5 |
ba963332b32e75b257471505d7c36095
|
|
| BLAKE2b-256 |
91f35dc9bdfb475f49fc74524799f6ef7db722320f7aeaf65cfdc917ec062975
|