Skip to main content

k-LLMmeans clustering algorithm

Project description

k-llmmeans

Scikit-learn compatible implementation of k-LLMmeans for text clustering with summary-based centroids.

This package adapts the original research code into an estimator API you can use with familiar fit, predict, and fit_predict workflows.

What This Package Provides

  • kLLMmeans estimator implementing BaseEstimator + ClusterMixin
  • scikit-learn style methods:
    • fit(X)
    • predict(X)
    • fit_predict(X)
  • configurable document embedding function (embedding_fn)
  • configurable cluster summarization function (summarizer_fn) or LiteLLM-backed LLM summarization
  • optional per-cluster sampling before summarization to keep prompts bounded on large clusters
  • optional precomputed embedding support for faster iterative experimentation

Installation

pip install k-llmmeans

Or from source:

pip install -e .

Quick Start

from k_llmmeans import kLLMmeans

docs = [
    "How to optimize SQL queries for large tables?",
    "What is the best way to tune a random forest model?",
    "PostgreSQL index strategy for analytics workloads",
    "Cross-validation tips for imbalanced classification",
]

model = kLLMmeans(
    n_clusters=2,
    llm="openai/gpt-4o-mini",
    max_llm_iter=5,
    random_state=0,
)

labels = model.fit_predict(docs)
print(labels)
print(model.summaries_)  # human-readable cluster summaries

LiteLLM Configuration

Pass a model string (any LiteLLM-supported model) or a dict of kwargs forwarded to litellm.completion:

model = kLLMmeans(
    n_clusters=2,
    llm={
        "model": "openai/gpt-4o-mini",
        "temperature": 0.2,
    },
)

You can also set LITELLM_MODEL or OPENAI_MODEL in the environment instead of passing llm.

Using Custom Embeddings and Summarization

You can fully control both the embedding and summarization steps:

from sentence_transformers import SentenceTransformer
from k_llmmeans import kLLMmeans

encoder = SentenceTransformer("all-MiniLM-L6-v2")

def embedding_fn(texts: list[str]):
    return encoder.encode(texts)

def summarizer_fn(cluster_texts: list[str]) -> str:
    # Replace with your own deterministic or LLM summarizer
    return " | ".join(cluster_texts[:2])

model = kLLMmeans(
    n_clusters=3,
    embedding_fn=embedding_fn,
    summarizer_fn=summarizer_fn,
)

model.fit(["text a", "text b", "text c", "text d"])

Sampling Cluster Documents for Summaries

The paper's few-shot k-LLMmeans variant summarizes at most m representative documents from each cluster instead of sending the full cluster to the LLM. Use cluster_sample_size for this prompt budget:

model = kLLMmeans(
    n_clusters=3,
    llm="openai/gpt-4o-mini",
    cluster_sample_size=10,
    cluster_sampling_strategy="kmeans++",
)

Set cluster_sample_size=None to summarize every document assigned to each cluster. Supported sampling strategies are kmeans++, random, centroid, and edge. The selected document indices for each LLM iteration are stored in sampled_indices_evolution_.

Debugging Runs

Use show_progress=True to display progress bars for LLM iterations and cluster summarization. Use verbose=True to print timing, input character counts, sampled document counts, summary lengths, token usage when available, and convergence details:

model = kLLMmeans(
    n_clusters=3,
    llm="openai/gpt-4o-mini",
    cluster_sample_size=10,
    show_progress=True,
    verbose=True,
)

After fitting, inspect summaries_evolution_ to see how cluster summaries changed, centroids_evolution_ to inspect summary-derived centroid embeddings, and sampled_indices_evolution_ to trace which input texts were used for each cluster summary. summary_workers can parallelize cluster summarization when your LLM provider and rate limits allow concurrent requests.

API Notes

  • Input X should be list[str].
  • The estimator stores standard fitted attributes such as:
    • labels_
    • cluster_centers_
    • n_iter_
  • Additional clustering interpretability attributes:
    • summaries_
    • summary_embeddings_
    • summaries_evolution_
    • centroids_evolution_
    • sampled_indices_evolution_

Citation

If you use this package in research or production work, please cite the original paper:

@article{diazrodriguez2025summaries,
  title={Summaries as Centroids for Interpretable and Scalable Text Clustering},
  author={Diaz-Rodriguez, Jairo},
  journal={arXiv preprint arXiv:2502.09667},
  year={2025}
}

Paper URL: https://arxiv.org/abs/2502.09667

Acknowledgment

This package is a scikit-learn compatible adaptation of the original project: https://github.com/jairoadiazr/k-LLMmeans

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

k_llmmeans-0.3.0.tar.gz (8.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

k_llmmeans-0.3.0-py3-none-any.whl (9.8 kB view details)

Uploaded Python 3

File details

Details for the file k_llmmeans-0.3.0.tar.gz.

File metadata

  • Download URL: k_llmmeans-0.3.0.tar.gz
  • Upload date:
  • Size: 8.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for k_llmmeans-0.3.0.tar.gz
Algorithm Hash digest
SHA256 679fe10c9427a98c80148423a4de2adf0ebe205084515a84d94fde3760c865cb
MD5 5fa3fb8e9e2ade96a33326b84e8cfda7
BLAKE2b-256 2afa9a13ace97f941dcb889d71802a777a79957a3f74c1ba478892276e39dcb6

See more details on using hashes here.

File details

Details for the file k_llmmeans-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: k_llmmeans-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 9.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for k_llmmeans-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 7658cc7fa99e966e422313177895c6fee4f75d89facfbf1d691f9696b18ea90e
MD5 ba963332b32e75b257471505d7c36095
BLAKE2b-256 91f35dc9bdfb475f49fc74524799f6ef7db722320f7aeaf65cfdc917ec062975

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page