Skip to main content

Topica: fast, all-purpose topic modeling for Python — a Rust core for LDA, STM, and more

Project description

Topica: fast, all-purpose topic modeling for Python

PyPI CI Docs License: Apache-2.0

topica is a fast topic-modeling library for Python with more than a dozen models, built for social scientists who want to move from text data to publishable results in a single workflow. It brings together models and tools usually split across JVM software like MALLET and R packages like stm, and runs them on a parallel Rust core competitive with the standard implementations, with reproducible fits: the variational models are identical to the bit, and the samplers reproduce from a fixed seed and thread count. Each model comes with the validation, covariate-effect, and reporting tools to meet the standards reviewers expect.

pip install topica

The core needs only NumPy. Optional extras add features without weighing the core down: topica[viz] (matplotlib plots), topica[formula] (R-style formulas), topica[polars] (Polars frames), and topica[llm] (LLM labels and embeddings, OpenAI or local via ollama). PyTorch is never required.

from topica import LDA

model = LDA(num_topics=2, seed=42)
model.fit([["cat", "dog", "fish"]] * 15 + [["planet", "star", "moon"]] * 15, iters=1000)

for i, words in enumerate(model.top_words(3)):
    print(f"Topic {i}:", " ".join(w for w, _ in words))

See the getting-started guide and the worked examples for end-to-end analyses.

Models

Count-based models learn topics from word counts (collapsed Gibbs, variational EM, or an amortized VAE):

Model What it's for
LDA Classic topics via fast collapsed-Gibbs (SparseLDA); optional multi-threaded and LightLDA alias samplers
STM The Structural Topic Model: correlated topics with prevalence and content covariates
STS Structural Topic and Sentiment-Discourse: covariate-driven topic sentiment/tone on top of STM
CTM Correlated topics (logistic-normal)
DMR Topics conditioned on document metadata (Dirichlet-multinomial regression)
DTM Dynamic topics that evolve across time slices
HDP Nonparametric LDA that infers the number of topics
keyATM / seededlda Guided topics steered by seed words
ProdLDA Sharper, more coherent topics via a product-of-experts word model, fit as an amortized VAE (no PyTorch)
PT / GSDMM Short-text models for tweets, survey answers, headlines
SupervisedLDA Topics shaped to predict a per-document response
LabeledLDA Supervised topics tied to document labels
SAGE Content-covariate topics: the same topic worded differently across groups
PA / HLDA Topic hierarchies (Pachinko, nested-CRP)

Embedding-based models start from document embeddings you supply (no PyTorch, no UMAP/numba in the wheel):

Model What it's for
BERTopic Cluster document embeddings, label topics by class-TF-IDF; topic reduction and a soft per-document distribution
Top2Vec Topics as points in the embedding space; topic words are the nearest word vectors
ETM Generative LDA with the topic-word distribution factored through embeddings (β = softmax(ρ·α)); per-document EM or an amortized VAE (inference="vae")
FASTopic Topics read off two optimal-transport plans between document, topic, and word embeddings

Every model exposes the same shape: fit(docs, …), then topic_word (φ), doc_topic (θ), top_words(n), and save/load. The count-based variational models (CTM/STM/STS/SupervisedLDA/DTM) parallelize across cores while staying bit-for-bit deterministic. The embedding models split into two kinds: BERTopic and Top2Vec run the reduce → cluster → represent pipeline, while ETM and FASTopic are generative and mixed-membership; all of them take vectors from any embedder (sentence-transformers, an API, a local model such as ollama). Full guides: the models and embedding topics.

Diagnostics & analysis

Model-agnostic: they work on any fitted model's topic_word/doc_topic:

  • Quality: coherence (u_mass, c_v, c_uci, c_npmi; computed in the Rust core), exclusivity, topic_diversity, quality_frontier
  • Labeling: label_topics (prob / FREX / lift / score), frex, relevance, find_thoughts, topic_table, summary
  • Validation: word_intrusion, document_intrusion, bootstrap_stability, search_k
  • Reliability: select_model (fit many seeds) and ensemble (combine runs into a consensus more reliable than any single fit — cluster/align/stable methods, the last a gensim EnsembleLda port)
  • Comparison: fighting_words (weighted log-odds) for contrasting corpora
  • Covariate effects: estimate_effect (method of composition, cluster-robust SEs, GLM links), topic_correlation, and the design helpers spline / interaction / one_hot (an stm-style API); posterior_theta_samples draws θ for the logistic-normal models (STM/CTM)
  • Preprocessing: tokenize, learn_phrases / apply_phrases, split_documents, the Corpus class

See diagnostics and covariate effects.

Performance

topica runs on a parallel Rust core. It is several times faster than R stm — the single-threaded field standard — for the structural and other variational models, and it matches the hand-tuned compiled samplers core for core: parity with Java MALLET on plain LDA and with the C++ keyATM on keyword models. On the political-blog corpus (2,000 documents, fit time only, same iterations on both sides):

Model Reference topica speedup
STM R stm 3–6× single-threaded, ~10–22× multithreaded
LDA Java MALLET parity single-threaded; multithread speedup grows with corpus size
keyATM R keyATM parity single-threaded, ~2× multithreaded

For the approximate parallel Gibbs samplers the multithreaded speedup grows with corpus size: the per-sweep count-table merge is fixed overhead, so larger corpora amortize it over more sampling work. LDA's eight-core speedup over MALLET runs about 3× at 2,000 documents and reaches ~4× at 5,000, so the small-corpus figures above understate what large-corpus users see.

Every fit is reproducible from a fixed seed and validated against its reference. See Benchmarks for the full methodology; reproduce the 2,000-document table with python benchmarks/speed_vs_r.py and the size-varying curve with python benchmarks/speed_vs_size.py.

Install from source

pip install maturin
git clone https://github.com/nealcaren/topica && cd topica
python -m venv .venv && source .venv/bin/activate
maturin develop --release --features python

Requires numpy >= 1.21. Use --release (the debug build is much slower).

Acknowledgements

Topica stands on a generation of open topic-modeling research and code. Each entry below lists the reference, its authors and year, and the topica class(es) it underlies; the other models are Rust ports or reimplementations, validated against these reference implementations.

  • MALLET (McCallum, 2002) — LDA, DMR, LabeledLDA: the SparseLDA sampler, Dirichlet-multinomial regression, and hyperparameter optimization. LDA binds David Mimno's RustMallet (Apache-2.0), reproducing its train CLI byte-for-byte; against Java MALLET (a different RNG) it recovers the same topics (cosine 1.000)
  • stm (Roberts, Stewart & Tingley, 2019) — STM, CTM, SAGE: variational EM, estimateEffect, searchK, FREX, spectral initialization, and the method of composition
  • sts (Chen & Mankad, 2024) — STS: the Structural Topic and Sentiment-Discourse model — the joint prevalence/sentiment Laplace E-step and the Poisson topic-word M-step, validated against the package
  • lda-c / ctm-c / dtm and hdp (Blei lab, 2006–2007) — CTM, DTM, HDP: the CTM, Dynamic Topic Model, and HDP samplers
  • gensim (Řehůřek & Sojka, 2010) — DTM, ensemble: coherence measures, the LdaSeqModel DTM reference, and the EnsembleLda (CBDBSCAN stable-topic) method ported for ensemble(method="stable")
  • tomotopy (bab2min, 2020) — API conventions (summary, the short-text models)
  • keyATM (Eshima, Imai & Sasaki, 2024) — KeyATM: the base, covariate, and dynamic models, the information-theory token weighting, and the Chib (1998) change-point HMM, validated against the package
  • seededlda (Watanabe, 2023) — SeededLDA: the seeded-prior scheme
  • LightLDA (Yuan et al., 2015) — LDA: the alias-table Metropolis-Hastings sampler
  • GSDMM (Yin & Wang, 2014) — GSDMM: the movie-group-process mixture for short text
  • ProdLDA / AVITM (Srivastava & Sutton, 2017) — ProdLDA: autoencoding variational inference and the product-of-experts word model
  • BERTopic (Grootendorst, 2022) and Top2Vec (Angelov, 2020) — BERTopic, Top2Vec: the embedding-clustering pipeline, class-based TF-IDF, and the reduce → cluster → represent design
  • ETM (Dieng, Ruiz & Blei, 2020) — ETM: the Embedded Topic Model (per-document variational EM and an amortized VAE)
  • FASTopic (Wu et al., 2024) — FASTopic: the optimal-transport topic model

The embedding-native models build on two pure-Rust crates: petal-clustering for HDBSCAN and umap-rs for the optional UMAP reducer, both BLAS-free.

Full citations for every model and reference implementation, and how to cite topica, are on the Citing page.

License

Apache-2.0 — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

topica-0.16.1.tar.gz (3.0 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

topica-0.16.1-cp39-abi3-win_amd64.whl (2.6 MB view details)

Uploaded CPython 3.9+Windows x86-64

topica-0.16.1-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.6 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ x86-64

topica-0.16.1-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (2.5 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ ARM64

topica-0.16.1-cp39-abi3-macosx_11_0_arm64.whl (2.4 MB view details)

Uploaded CPython 3.9+macOS 11.0+ ARM64

topica-0.16.1-cp39-abi3-macosx_10_12_x86_64.whl (2.5 MB view details)

Uploaded CPython 3.9+macOS 10.12+ x86-64

File details

Details for the file topica-0.16.1.tar.gz.

File metadata

  • Download URL: topica-0.16.1.tar.gz
  • Upload date:
  • Size: 3.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for topica-0.16.1.tar.gz
Algorithm Hash digest
SHA256 df89c82cbcc7065e60ee985e01b51ca66690dcbf5e42960b1ab28ffd543ef9ba
MD5 bb98e81dd472a048b6ce9b403f69bb0c
BLAKE2b-256 b6bd8c99ad41430438b44130fb9142957d39334c0fea5074ab700ba2a1829c72

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.16.1.tar.gz:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.16.1-cp39-abi3-win_amd64.whl.

File metadata

  • Download URL: topica-0.16.1-cp39-abi3-win_amd64.whl
  • Upload date:
  • Size: 2.6 MB
  • Tags: CPython 3.9+, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for topica-0.16.1-cp39-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 90bfc4dd74c3cc5732bbc6bb7c673f2b834d5765390fdfdd310b6ede6204876f
MD5 039946e553ba1d614d8bbf1f8edceece
BLAKE2b-256 00911042bc5138eccb639655044c044eb1601e73885fd688d811653b72b2ffd8

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.16.1-cp39-abi3-win_amd64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.16.1-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for topica-0.16.1-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 fb26a526ce0b3efeeffa1b7ef200368b9d3581381e8d8e64d5bada7389f5a26c
MD5 fda30d991e2a33af995ed1a83e27fd55
BLAKE2b-256 7eadbbb83d526f949b06688c60623fd77a3db4fc8cd1c197d586e19a1ab871ff

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.16.1-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.16.1-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for topica-0.16.1-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 afc79c99ccca345e23c7daf092afc6806eb50194adcd4d17f2a43554f3ef07f6
MD5 f71b3f33a137d13b88956fe03ec0f99e
BLAKE2b-256 2c9aacd8f030c96064c2e2a81c4f872d30e17686b71ed95a520c3fb1c209bf5b

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.16.1-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.16.1-cp39-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for topica-0.16.1-cp39-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 cc79246fa19fdd5590ccb4ec1ab56d8ffcc092718a2b15d817022cac5d4df136
MD5 7a2dca2f9af9c3183b393061ffba2ed0
BLAKE2b-256 be3e9d02ae1a7664b3e2026d84fb72f844b39fe239126d1c2caf3b1c61d65437

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.16.1-cp39-abi3-macosx_11_0_arm64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.16.1-cp39-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for topica-0.16.1-cp39-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 9b8a954d55194e2fa71645a1d549020f09ea0c4c90ecab621cd0d838adcb745f
MD5 a1ec86f024eaf01b02332b1b49a0f34c
BLAKE2b-256 19438427f8643a7c63ca98e4c56d67fe2e088fdb9afa3fc39c3e16af760beb43

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.16.1-cp39-abi3-macosx_10_12_x86_64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page