Skip to main content

Topica: fast, all-in-one topic modeling for Python

PyPI CI Docs Website License: Apache-2.0

topica is a fast, memory-efficient, all-in-one topic-modeling library for Python, built for social scientists. It brings together more than fifty models usually split across JVM tools like MALLET and R packages like stm, including LDA, STM, CTM, keyATM, BERTopic, and neural, dynamic, short-text, and embedding-based models, all under one NumPy-native API. Every model is validated against its reference implementation and reproducible from a fixed seed; all share one set of diagnostics, labeling, validation, and covariate-effect tools, so you learn a single workflow and it applies across the roster. It installs as a single wheel that needs only NumPy and pandas: no JVM, no PyTorch.

pip install topica

Quick start

Point topica at a DataFrame and read the topics. This runs exactly as written, on a bundled example dataset, right after install:

import topica

df = topica.datasets.load_gadarian()          # bundled; loads offline
corpus = topica.from_dataframe(
    df, text_col="open.ended.response", stopwords=topica.data.ENGLISH_STOPWORDS
)

model = topica.LDA(num_topics=5, seed=13)
model.fit(corpus)                             # sensible defaults; no tuning required
print(topica.summary(model))                  # top words per topic

from_dataframe keeps your metadata aligned to the documents that survive pruning, so the same corpus feeds a structural topic model that relates topic prevalence to a covariate and reports the effect with uncertainty:

prevalence = corpus.metadata[["treatment"]]   # a numeric DataFrame goes straight in

stm = topica.STM(num_topics=5, seed=13)
stm.fit(corpus, prevalence, prevalence_names=["treatment"])

draws  = topica.effects.posterior_theta_samples(stm, nsims=30, seed=0)
effect = topica.effects.estimate_effect(draws, prevalence, feature_names=["treatment"])

Your own data is one line away: pass pandas.read_csv("yours.csv") to from_dataframe. See the getting-started guide and the worked examples for analyses end to end.

Fits are reproducible and validated: the variational models are bit-exact against their references, the samplers reproduce from a fixed seed and thread count, and every model is checked against its reference implementation (R stm, MALLET, keyATM, and more).

The core needs only NumPy and pandas. Optional extras add features without weighing it down: topica[viz] (matplotlib plots), topica[formula] (R-style formulas), topica[polars] (Polars frames), and topica[llm] (LLM labels and embeddings, OpenAI or local via ollama).

Models

Find your goal in the left column. Every model is validated against a reference implementation, so the choice is about fit to your research design, not quality. A specialized model is often the right first choice when your data calls for it.

Common openings

If your goal is… Start with Also consider First calls Note
Explore themes with no prior structure LDA NMF search_k(), topic_table() The default first pass. NMF is a fast, deterministic alternative.
Relate topics to metadata (author, date, party) STM DMR estimate_effect(), one_hot(), spline() STM gives covariate effects with uncertainty; DMR is a lighter Gibbs prior.
Measure concepts you can name in advance KeyATM SeededLDA KeyATM(keywords=…), .keyword_rate Anchor named topics with a few seed words each.
Very short documents: tweets, headlines, survey answers GSDMM PT fit() One topic per document; standard LDA over-fragments short text.
Cluster by meaning using embeddings BERTopic ETM fit(docs, doc_embeddings=…) Clustering, not a posterior: topic-proportion uncertainty and effect estimation behave differently than the models above.

Specialized approaches. Start here when your design calls for one.

If your data or goal is… Start with Also consider First calls Note
Topics shift over time slices DTM DETM fit(docs, times=…) Prevalence and content evolve across periods; DETM adds embeddings.
Documents linked in a network (citations, replies) RTM fit(docs, links=…) Models the text and the link graph jointly.
Documents in more than one language PolylingualLDA fit(doc_tuples) Aligned topics across languages from translation-linked tuples.
Place authors or actors on an ideological scale Wordfish TBIP fit(docs) Scaling from word usage; TBIP adds a text-based ideal-point prior.
How tone or sentiment varies with metadata STS estimate_effect() Sentiment-discourse decomposition; reach for it when tone is the question.

topica.list_models(common_start=True) returns the common openings in code; list_models(group=…, brings=…) filters the full roster below.

All models (more than fifty, grouped by what you bring and what you want; click to expand)

Models are organized by what you bring and what you want, not by inference family. The from topica import X namespace is flat; topica.list_models(group=…, brings=…, inference=…, determinism=…) filters this roster in code. Brings is what you supply beyond raw text; Reproducibility is bit-exact (identical regardless of thread count), seed-reproducible (identical from a fixed seed and thread count), or llm-bounded.

Every model below is validated before it enters the roster: against a maintained reference implementation where one exists (MALLET, gensim, R stm, tomotopy, and the like), otherwise by planted recovery on a synthetic corpus with a known answer. See validation for where each model stands. The groupings are about fit to your research design, not quality: a specialized model is the right first choice when your data calls for it.

Common starting points

One per common goal, where most social scientists begin. BERTopic works differently from the others: it clusters document embeddings rather than fitting a posterior, so topic-proportion uncertainty and covariate-effect estimation do not carry over directly.

Model Brings Inference Reproducibility Summary
LDA text gibbs seed-reproducible Classic latent Dirichlet allocation via a fast SparseLDA collapsed-Gibbs sampler.
NMF text matrix-factorization bit-exact Non-negative matrix factorization of the document-term matrix via multiplicative updates.
STM text, metadata variational bit-exact Structural topic model: relate topic prevalence and content to covariates.
KeyATM text, seeds gibbs seed-reproducible Keyword-assisted topics: anchor named topics with a few seed words each.
GSDMM text gibbs seed-reproducible Gibbs-sampling Dirichlet mixture: one topic per short document.
BERTopic text, embeddings clustering seed-reproducible Cluster document embeddings; label topics by class-based TF-IDF.

Specialized approaches

The right first choice when your design calls for one: short text, change over time, document networks, multiple languages, ideological scaling, and more.

General-purpose

Model Brings Inference Reproducibility Summary
OnlineLDA text variational seed-reproducible Online (streaming) variational-Bayes LDA (Hoffman et al. 2010): minibatch stochastic VB with a decaying learning rate and a streaming partial_fit; the gensim LdaModel analogue for very large or streaming corpora.
CTM text variational bit-exact Correlated topic model: a logistic-normal prior that lets topics co-occur.
ProdLDA text vae seed-reproducible Product-of-experts LDA (AVITM) for sharper, more coherent topics; hand-coded VAE.
HDP text gibbs seed-reproducible Hierarchical Dirichlet process: infers the number of topics from the data.
LSA text svd seed-reproducible Latent semantic analysis: a truncated SVD of the weighted document-term matrix.
AnchorLDA text matrix-factorization bit-exact Anchor-words spectral recovery (Arora et al. 2013): deterministic, Gibbs-free topics from the word co-occurrence matrix.
PolylingualLDA text gibbs seed-reproducible Polylingual topic model (Mimno et al. 2009): aligned topics across languages from document tuples that share one topic distribution.
CorEx text information-theoretic seed-reproducible Correlation Explanation: information-theoretic topic model that maximizes total correlation; supports anchor words.
MGLDA text gibbs seed-reproducible Multi-Grain LDA: global (document-level) + local (sliding-window aspect) topics with a per-token grain switch. For reviews / aspect extraction.
TopicalNGrams text gibbs seed-reproducible Topical N-Grams (Wang, McCallum & Wei 2007): an LDA extension that jointly discovers topics and topic-specific multiword phrases. A per-token bigram-status indicator, sampled with the topic, decides whether a token continues a phrase from the previous word given its topic, so phrase structure is learned during fitting rather than fixed beforehand. Exposes top_phrases alongside top_words.

Covariates & structure

Model Brings Inference Reproducibility Summary
STS text, metadata variational bit-exact Structural topic-and-sentiment model over document metadata.
SAGE text, metadata gibbs seed-reproducible Sparse additive generative model: the same topic worded differently across groups.
DMR text, metadata gibbs seed-reproducible Dirichlet-multinomial regression: a document-metadata prior on topic proportions.
GDMR text, metadata gibbs seed-reproducible Generalized DMR with a smooth (Legendre-basis) prior over continuous covariates.
Scholar text, metadata, labels vae seed-reproducible SCHOLAR (Card et al. 2018): a ProdLDA VAE with a covariate-shifted prevalence prior, an optional supervised label head, and optional content (topic-covariate) word deviations — neural STM prevalence + sLDA + SAGE.
RTM text, links variational seed-reproducible Relational topic model (Chang & Blei 2010): jointly models document text and a link graph (citations, hyperlinks, adjacency); predicts links from words and words from links.
FactorialLDA text gibbs seed-reproducible Factorial LDA (Paul & Dredze 2012): each token is a K-tuple of latent factors (e.g. topic x sentiment); structured word priors tie tuples sharing a component and a sparsity prior deactivates unsupported tuples.
AuthorTopic text, metadata gibbs seed-reproducible Author-Topic Model: each author has a topic distribution; documents mix their authors. Answers what an author writes about.
AuthorRecipientTopic text, metadata gibbs seed-reproducible Author-Recipient-Topic (McCallum et al. 2007): topics conditioned on the (sender, recipient) pair, for the language of a directed social network (who talks to whom about what). Realized over the AuthorTopic engine.

Guided & supervised

Model Brings Inference Reproducibility Summary
SeededLDA text, seeds gibbs seed-reproducible Seeded LDA: steer named topics toward supplied seed words.
GuidedNMF text, seeds matrix-factorization seed-reproducible Guided NMF: seed-word-guided semi-supervised NMF; the matrix-factorization analogue of SeededLDA.
LabeledLDA text, labels gibbs seed-reproducible Labeled LDA: each document label is a topic; tokens are restricted to its labels.
SupervisedLDA text, labels variational seed-reproducible Supervised LDA: topics shaped to predict a per-document real-valued response.
DiscLDA text, labels gibbs seed-reproducible Discriminative LDA (Lacoste-Julien et al. 2008): topics split into per-class and shared blocks; reads how classes talk differently.

Short text

Model Brings Inference Reproducibility Summary
PT text gibbs seed-reproducible Pseudo-document topic model: pool short texts into pseudo-documents.
BTM text gibbs seed-reproducible Biterm topic model: learns topics from corpus-level word co-occurrence (biterms).

Dynamic & hierarchical

Model Brings Inference Reproducibility Summary
DTM text, times variational seed-reproducible Dynamic topic model: a fixed topic set whose word distributions drift across time slices.
DETM text, embeddings, times vae seed-reproducible Dynamic embedded topic model: embedding-factored topics that drift across time slices, fit as an amortized VAE.
TopicsOverTime text, times gibbs seed-reproducible Topics over Time: LDA with a per-topic Beta density over continuous timestamps; each topic has a temporal peak. Descriptive continuous-time prevalence (not vocabulary drift).
HLDA text gibbs seed-reproducible Hierarchical LDA (nested CRP): a learned tree of super- and sub-topics.
PA text gibbs seed-reproducible Pachinko allocation: a DAG of super- and sub-topics.

Embedding-based

Model Brings Inference Reproducibility Summary
KeyNMF text, embeddings matrix-factorization bit-exact KeyNMF (Kristensen-McLachlan et al. 2024): NMF over an embedding-derived keyword-importance matrix. For each document it scores its words by the similarity between the document embedding and the word embedding, keeps the top-N positive, and factors that sparse doc-word matrix. The bridge between the count-based NMF family and the embedding backend; sparse, readable topics robust to short/noisy text.
Top2Vec text, embeddings clustering seed-reproducible Topics as dense regions in a joint document-word embedding space.
SemanticSignalSeparation text, embeddings ica seed-reproducible Topics as independent axes of semantic space (S3, Kardos et al. 2025): FastICA over the document embeddings, with each word's importance read off by projecting the vocabulary embeddings onto each axis. Signed poles.
ETM text, embeddings variational seed-reproducible Embedded topic model: topic-word distributions factored through word embeddings.
GaussianLDA text, embeddings gibbs seed-reproducible Gaussian LDA (Das, Zaheer & Dyer 2015): each topic is a Gaussian over the word-embedding space (Normal-Inverse-Wishart prior), so topics generalize over semantically similar words. Collapsed Gibbs with a Student-t posterior predictive and rank-1 Cholesky up/downdates.
FASTopic text, embeddings optimal-transport seed-reproducible Topics from optimal-transport plans between document, topic, and word embeddings.
CombinedTM text, embeddings vae seed-reproducible Contextualized ProdLDA: encoder reads the bag of words plus a document embedding.
ZeroShotTM text, embeddings vae seed-reproducible Contextualized ProdLDA: encoder reads the document embedding alone, enabling cross-lingual transfer.
InfoCTM text, dictionary vae seed-reproducible Cross-lingual: two ProdLDA models aligned by a bilingual dictionary through a mutual-information term.

Ideal point

Model Brings Inference Reproducibility Summary
Wordfish text em bit-exact Poisson scaling (Slapin & Proksch 2008): an unsupervised one-dimensional ideal-point estimate from word frequencies alone, no topics. The word-frequency baseline companion to IdealPointTM.
Wordshoal text, metadata em bit-exact Multi-domain scaling (Lauderdale & Herzog 2016): scales each debate/domain with Wordfish, then combines the within-domain positions into one cross-domain actor scale via a linear factor model. The multi-domain extension of Wordfish, for speeches carrying trusted debate labels.
TBIP text variational seed-reproducible Text-Based Ideal Points (Vafa, Naidu & Blei 2020): a Poisson factorization whose neutral topic-word intensities are rescaled by a per-word ideological factor exp(x_s * eta_kv), with the author position x_s latent. Fit by the paper's mean-field variational inference (reparameterized SVI). Recovers ideological scales from unlabeled text.
PartyEmbeddings text, metadata neural-embedding seed-reproducible Party embeddings (Rheault & Cochrane 2020): a PV-DM paragraph-vector model trained by negative sampling with party-period metadata tags; the leading principal components of the learned party vectors give the ideological scale, and words share the space so a party's language can be read off by proximity. The corpus-trained word-embedding member of the ideal-point family.

LLM-based

Model Brings Inference Reproducibility Summary
TopicGPT text, llm prompting llm-bounded LLM-driven topic discovery: prompt a model to propose, refine, and assign a topic taxonomy with descriptions.

Experimental

Not (yet) on the validated roster, on one of two grounds: the model is unpublished (a topica original with no paper), or it is a published method whose benefit has not held up against a simpler baseline. A published method with faithful inference stays validated even when its basis is planted-recovery, so planted validation alone does not land a model here. Gated: call topica.enable_experimental() (or set TOPICA_EXPERIMENTAL=1) before use. These may change or be removed without a deprecation cycle. For the triple gate a model clears to graduate to the validated roster, and where every model stands on validation evidence, see Validation & graduation.

Model Brings Inference Reproducibility Summary
TensorLDA text svd seed-reproducible Online Tensor LDA (Kangaslahti et al. 2026): deterministic method-of-moments topic modeling via second and third-order cumulants.
NarrativeTM text gibbs seed-reproducible Intra-document narrative trajectory model: captures how topic prevalence shifts across the progress of a text.
CSATM text, links gibbs seed-reproducible Conversational Structure Aware TM (Sun et al. 2020): weights each comment's tokens by a reply-tree 'popularity' score and, after Gibbs, smooths each comment's topics toward its ancestors along the reply path ('transitivity'). For threaded forum data (posts + nested comments). Ported from the paper (no reference implementation); validated by planted recovery + LDA reduction.
ThreadTM text, links variational seed-reproducible ThreadTM (topica-original): a reply-threaded topic model — CTM logistic-normal topics with a reply-tree structured prior, so a reply's topic prior is coupled to the comment it answers (a persistence-smoothing prior, reverting toward its covariate-group baseline). Supports a prevalence covariate and an optional content covariate (topic words shift by depth/level, à la STM's content model). For threaded discussion (posts + nested comments). Validated by planted recovery and a synthetic held-out-beat gate (the parent's topics predict held-out leaf tokens better than the no-tree baseline on persistence-structured data). The real-corpus benefit is genre-dependent; read kappa_ci before claiming persistence.
IdealPointTM text, embeddings variational seed-reproducible Topic model with a latent ideal-point head: each author gets a low-dimensional position that shifts within-topic word choice, with a per-topic discrimination. Consumes word tokens as counts (Wordfish with topics) or, when word embeddings are supplied to fit, factored through them as in ETM. The unsupervised, latent-trait twin of the STM content covariate.
IdealPointSentenceTM text, embeddings em seed-reproducible Continuous ideal-point topic model over sentence/document embeddings: topics are Gaussian clusters whose centroids are displaced by a latent author position. The sentence-embedding sibling of IdealPointTM, fit by EM.
EmbeddingLDA text, embeddings gibbs seed-reproducible LDA anchored by pre-trained embeddings: k-means clusters the vocabulary embeddings, seeds each topic with the words nearest a cluster centroid, and (optionally) biases each document's mixture toward its own embedding. A topica original; validated by planted-recovery only.

Every model exposes the same shape: fit(docs, …), then topic_word (φ), doc_topic (θ), top_words(n), and save/load, so one diagnostic, labeling, and effect-estimation stack applies to all of them and a new model inherits it for free. The embedding-based models take document vectors from any embedder (sentence-transformers, an API, or a local model such as ollama; no PyTorch or UMAP/numba in the wheel). Full guides: the models and embedding topics.

Diagnostics & analysis

Model-agnostic: they work on any fitted model's topic_word/doc_topic:

  • Quality: coherence (u_mass, c_v, c_uci, c_npmi; co-occurrence counting in the Rust core), exclusivity, topic_diversity, quality_frontier
  • Labeling: label_topics (prob / FREX / lift / score), frex, relevance, find_thoughts, topic_table, summary
  • Validation: word_intrusion, document_intrusion, bootstrap_stability, search_k
  • Reliability: select_model (fit many seeds) and ensemble (combine runs into a consensus more reliable than any single fit — cluster/align/stable methods, the last derived from gensim's EnsembleLda)
  • Comparison: fighting_words (weighted log-odds) for contrasting corpora
  • Covariate effects: estimate_effect (method of composition, cluster-robust SEs, GLM links), topic_correlation, and the design helpers one_hot, spline, and interaction (in topica.effects / topica.design; they build covariate bases for any model's design matrix); posterior_theta_samples draws θ for the logistic-normal models (STM/CTM)
  • Preprocessing: tokenize, learn_phrases / apply_phrases, split_documents, the Corpus class

See diagnostics and covariate effects.

Performance

topica runs on a parallel Rust core, so the whole roster is fast and every fit is reproducible from a fixed seed. Core for core it matches the hand-tuned compiled samplers: parity with Java MALLET on plain LDA and with the C++ keyATM on keyword models. For the structural and other variational models it is several times faster than R stm, the single-threaded field standard. Fit to convergence (both at the same emtol, spectral start), on real corpora:

Model Reference topica speedup (to convergence)
STM R stm 1.7–2.7× single-threaded, ~5–7× multicore
LDA Java MALLET parity single-threaded; multithread speedup grows with corpus size
keyATM R keyATM parity single-threaded, ~2× multithreaded

topica also fits in about a quarter of R stm's memory (≈180MB against ≈675MB at 5,000 documents). For the approximate parallel Gibbs samplers the multithreaded speedup grows with corpus size: the per-sweep count-table merge is fixed overhead, so larger corpora amortize it over more sampling work. LDA's eight-core speedup over MALLET runs about 3× at 2,000 documents and reaches ~4× at 5,000.

Every fit is reproducible from a fixed seed and validated against its reference. See Benchmarks for the full methodology; reproduce the structural-model table with python benchmarks/bench_stm_convergence.py and the size-varying LDA curve with python benchmarks/speed_vs_size.py.

Install from source

pip install maturin
git clone https://github.com/nealcaren/topica && cd topica
python -m venv .venv && source .venv/bin/activate
maturin develop --release --features python

Requires numpy >= 1.21. Use --release (the debug build is much slower).

Acknowledgements

Topica was inspired by a post from David Mimno about porting Java MALLET to Rust. As a long-time Python user, I had long been jealous of the topic-modeling tools available in other languages; this seemed like an opportunity to make those capabilities easier to use in Python for me and for others.

As such, Topica stands on a generation of open topic-modeling research and code. Each entry below lists the reference, its authors and year, and the topica class(es) it underlies; the other models are Rust ports or reimplementations, validated against these reference implementations.

  • MALLET (McCallum, 2002) — LDA, DMR, LabeledLDA: the SparseLDA sampler, Dirichlet-multinomial regression, and hyperparameter optimization. LDA began as a port of David Mimno's RustMallet (Apache-2.0) and follows its SparseLDA sampler and fixed-point optimizer closely, but uses its own RNG (PCG), so it is not byte-identical to RustMallet. Against Java MALLET (also a different RNG) it recovers the same topics on a planted corpus (cosine 1.000)
  • stm (Roberts, Stewart & Tingley, 2019) — STM, CTM, SAGE: variational EM, estimateEffect, searchK, FREX, spectral initialization, and the method of composition
  • sts (Chen & Mankad, 2024) — STS: the Structural Topic and Sentiment-Discourse model — the joint prevalence/sentiment Laplace E-step and the Poisson topic-word M-step, validated against the package
  • lda-c / ctm-c / dtm and hdp (Blei lab, 2006–2007) — CTM, DTM, HDP: the CTM, Dynamic Topic Model, and HDP samplers
  • gensim (Řehůřek & Sojka, 2010) — DTM, ensemble, OnlineLDA: the coherence-pipeline conventions (the coherence_type= API and default sliding windows; the measures themselves are Röder et al. 2015 and Mimno et al. 2011), the LdaSeqModel DTM reference, the EnsembleLda (CBDBSCAN stable-topic) method that ensemble(method="stable") derives from (matching it on well-separated inputs, with two documented improvements on edge cases where gensim degenerates), and the LdaModel online-VB reference for OnlineLDA
  • onlineldavb (Hoffman, Blei & Bach, 2010) — OnlineLDA: online (streaming) variational Bayes for LDA — the minibatch stochastic-VB E-step, the decaying Robbins-Monro learning rate, and the streaming partial_fit; written from the paper and validated against onlineldavb.py and gensim's LdaModel as external oracles (both are copyleft, so no code was copied)
  • tomotopy (bab2min, 2020) — API conventions (summary, the short-text models), and GDMR (generalized DMR; Lee & Song, 2020), validated against its GDMRModel
  • scikit-learn (Pedregosa et al., 2011) — NMF: the multiplicative-update solver (Lee & Seung, 2001) and the NNDSVD initialization (Boutsidis & Gallopoulos, 2008), validated against sklearn.decomposition.NMF (BSD-3-Clause); and LSA: latent semantic analysis / indexing (Deerwester et al., 1990), validated against sklearn.decomposition.TruncatedSVD (BSD-3-Clause) including its svd_flip sign convention. The numerics are reimplemented in Rust; the randomized truncated SVD shared by both (it seeds NMF's NNDSVD and is the LSA factorization itself) follows Halko et al. (2011).
  • keyATM (Eshima, Imai & Sasaki, 2024) — KeyATM: the base, covariate, and dynamic models, the information-theory token weighting, and the Chib (1998) change-point HMM, validated against the package
  • seededlda (Watanabe, 2023) — SeededLDA: the corpus-frequency-scaled seed prior (count × weight × 100), validated against the package's seed matrix and seeded topics
  • LightLDA (Yuan et al., 2015) — LDA: the alias-table Metropolis-Hastings sampler
  • GSDMM (Yin & Wang, 2014) — GSDMM: the movie-group-process mixture for short text
  • BTM (Yan, Guo, Lan & Cheng, 2013; R package by Jan Wijffels) — BTM: the biterm co-occurrence topic model for short text
  • Polylingual Topic Models (Mimno, Wallach, Naradowsky, Smith & McCallum, 2009) — PolylingualLDA: LDA over aligned document tuples that share one topic distribution, giving topics aligned across many languages; validated against MALLET's PolylingualTopicModel as a black-box oracle
  • DiscLDA (Lacoste-Julien, Sha & Jordan, 2008) — DiscLDA: discriminative LDA with per-class and shared topic blocks; the fixed block-transform variant, validated against the paper's 20 Newsgroups feature-classification result (no reference implementation exists, so it is paper-derived)
  • Factorial LDA (Paul & Dredze, 2012) — FactorialLDA: sparse multi-dimensional topics, where each token is a K-tuple of latent factors tied by structured log-linear priors; implemented from the paper's mathematics (the reference Java is GPL and non-reproducible, so the port is certified by finite-difference gradient and factor-tying tests plus planted recovery, not seed parity)
  • ProdLDA / AVITM (Srivastava & Sutton, 2017) — ProdLDA: autoencoding variational inference and the product-of-experts word model
  • SCHOLAR (Card, Tan & Smith, 2018; reference dallascard/scholar, Apache-2.0) — Scholar: metadata in a ProdLDA VAE — a covariate-dependent topic-prevalence prior (neural STM prevalence), an optional supervised label head (neural sLDA), and optional content/topic-covariate word deviations (neural SAGE), on topica's ProdLDA backbone, validated against the reference as a numerical oracle
  • BERTopic (Grootendorst, 2022) and Top2Vec (Angelov, 2020) — BERTopic, Top2Vec: the embedding-clustering pipeline, class-based TF-IDF, and the reduce → cluster → represent design
  • CETopic / topicx (Zhang, Fang, Chen & Namazi-Rad, NAACL 2022, MIT) — BERTopic(weighting="tfidf-idf"): the TFIDF×IDF_i topic-word selection scheme (a corpus-level TF-IDF averaged per cluster times a cross-cluster IDF penalty), ported faithfully to the reference's scikit-learn defaults
  • S³ / turftopic (Kardos, Kostkan, Enevoldsen, Vermillet, Nielbo & Rocca, ACL 2025, MIT) — SemanticSignalSeparation: Semantic Signal Separation, FastICA over contextual document embeddings with topic words read off by projecting the vocabulary embeddings onto each independent axis, ported faithfully to the reference's scikit-learn FastICA defaults
  • ETM (Dieng, Ruiz & Blei, 2020) — ETM: the Embedded Topic Model (per-document variational EM and an amortized VAE)
  • DETM (Dieng, Ruiz & Blei, 2019) — DETM: the Dynamic Embedded Topic Model (structured amortized variational inference with a hand-coded LSTM)
  • FASTopic (Wu et al., 2024) — FASTopic: the optimal-transport topic model
  • contextualized-topic-models (Bianchi et al., MIT) — CombinedTM (Bianchi, Terragni & Hovy, 2021) and ZeroShotTM (Bianchi, Nozza & Hovy, 2021): ProdLDA encoders that read a contextual document embedding, alongside or in place of the bag of words
  • quanteda.textmodels (Benoit et al., 2018) — Wordfish: the Slapin & Proksch (2008) Poisson scaling model, validated against its textmodel_wordfish (the recovered scale and the analytic position standard errors both match at correlation 1.00 on a corpus sampled from the model)
  • tbip (Vafa, Naidu & Blei, 2020) — TBIP: Text-Based Ideal Points; the official implementation is TensorFlow 1.x, so topica reimplements the published model and its mean-field variational inference, validated against an independent PyTorch reference
  • TensorLy TLDA (Kangaslahti, Ebanks, Kossaifi, Liu, Alvarez & Anandkumar, 2026) — TensorLDA: online tensor latent Dirichlet allocation via second- and third-order moments. The Rust implementation is experimental; see the TensorLDA validation record for its current evidence and limitations.
  • partyembed (Rheault & Cochrane, 2020) — PartyEmbeddings: party embeddings via a PV-DM paragraph-vector model with party-period metadata tags, placed by PCA of the learned party vectors. The reference builds on gensim's Doc2Vec; topica reimplements the PV-DM negative-sampling training in Rust (from Mikolov et al. 2013 and Le & Mikolov 2014) and is validated against that Doc2Vec scale (correlation 1.00 on a planted ordering)
  • CLNTM (Nguyen & Luu, 2021) — the InfoNCE contrastive regularization on topic vectors offered by the contrastive= flag on the VAE models
  • WHAI / Weibull-Dirichlet VAE (Zhang et al., 2018; Burkhardt & Kramer, 2019) — the Weibull-reparameterized Dirichlet prior offered by prior="dirichlet" on the VAE models
  • Neural variational topic models with alternative priors (Miao, Grefenstette & Blunsom, 2017; Nalisnick & Smyth, 2017) — the Gaussian stick-breaking prior offered by prior="stick_breaking" on the VAE models
  • TopicGPT (Pham et al., NAACL 2024, MIT) — TopicGPT: the generate / refine / assign prompt flow for LLM-driven topic discovery

The embedding-native models build on two pure-Rust crates: petal-clustering for HDBSCAN and umap-rs for the optional UMAP reducer, both BLAS-free.

Full citations for every model and reference implementation, and how to cite topica, are on the Citing page.

Contributing, tests, and support

Contributions are welcome. See CONTRIBUTING for the development setup and workflow, CONTRIBUTING-MODELS for adding a new topic model, and the conventions guide for the cross-model naming and API contract. All participants are expected to follow our Code of Conduct.

To run the test suite after a source build:

cargo test --lib                                   # Rust unit tests
python -m pytest tests/ -q                         # Python tests
mkdocs build --strict                              # docs build clean

Every push runs these on CI (see the badge above). The parity/ checks validate models against their reference implementations (R stm, keyATM, MALLET); they skip cleanly when those toolchains are not installed.

  • Report a bug or request a feature: open an issue.
  • Ask a question or share how you are using topica: start a discussion.

topica is maintained by Neal Caren. Issues and pull requests are triaged on a best-effort basis.

Citation

If you use topica in published work, please cite it. GitHub's Cite this repository button (top right) generates a formatted reference from CITATION.cff. A software paper is in preparation; until it appears, cite the software release:

@software{caren_topica,
  author  = {Caren, Neal},
  title   = {topica: fast, all-purpose topic modeling for Python},
  year    = {2026},
  url     = {https://github.com/nealcaren/topica},
  version = {0.54.0}
}

Replace version with the release you used. For the individual models and their reference implementations, see the Citing page.

License

Apache-2.0 — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

topica-0.58.0.tar.gz (20.3 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

topica-0.58.0-cp39-abi3-win_amd64.whl (5.4 MB view details)

Uploaded CPython 3.9+Windows x86-64

topica-0.58.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (5.2 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ x86-64

topica-0.58.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (4.9 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ ARM64

topica-0.58.0-cp39-abi3-macosx_11_0_arm64.whl (4.7 MB view details)

Uploaded CPython 3.9+macOS 11.0+ ARM64

topica-0.58.0-cp39-abi3-macosx_10_12_x86_64.whl (5.0 MB view details)

Uploaded CPython 3.9+macOS 10.12+ x86-64

File details

Details for the file topica-0.58.0.tar.gz.

File metadata

  • Download URL: topica-0.58.0.tar.gz
  • Upload date:
  • Size: 20.3 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for topica-0.58.0.tar.gz
Algorithm Hash digest
SHA256 73063a797cf81006d37d03448f45c85ea524b8a4d2bd681bfe6bd7a77d2ce9f7
MD5 370f48818f3eca59fa0c155eefd892ef
BLAKE2b-256 25036e67e5bbeba4deb9bcdf2f66e0f2feea5c2ae829c9676336d1ec1ca3823a

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.58.0.tar.gz:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.58.0-cp39-abi3-win_amd64.whl.

File metadata

  • Download URL: topica-0.58.0-cp39-abi3-win_amd64.whl
  • Upload date:
  • Size: 5.4 MB
  • Tags: CPython 3.9+, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for topica-0.58.0-cp39-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 aefbba1af96bf3e4ea58339cd3195b98f768d943e9acc80aedee55282b0bc890
MD5 6935b9ae8892bf99507578e575e00da6
BLAKE2b-256 6a1e64b5af40bd6db73780ae5d6e8efbe02735e243dc9843192b090366b462e4

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.58.0-cp39-abi3-win_amd64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.58.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for topica-0.58.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 3ab32e1c3467d807ed37aaad30f4f601b3474468dc078a31e22e61fb4874aed1
MD5 3a166b5a3c02f84aa0e2647ee1bc5a85
BLAKE2b-256 cb6ad3ffe97288797f6fbb0775dd93f14d1d103396edfb8930b5c3091920f6db

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.58.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.58.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for topica-0.58.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 d69f66bf9e228b3f8e31ca86f0ae42a98c75539e64d066269a92e632ee427d12
MD5 bd00c0cb27f761a4fe066d1719d65131
BLAKE2b-256 e673d91024def77bc4f3e60be57b32ce89f6c816d461a1b70f5cfb3d9f342756

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.58.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.58.0-cp39-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for topica-0.58.0-cp39-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 f9037660fa1a7dc43ca48ea8867fec1e23c5684c40db9ea40aeccd462b3f5936
MD5 1401136529a278f1887123d401867293
BLAKE2b-256 059b39e873ad609e00a7cc1d2ca8b5a0bfaff837fcb35e1a0bf6d9b920dac849

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.58.0-cp39-abi3-macosx_11_0_arm64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.58.0-cp39-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for topica-0.58.0-cp39-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 4058d4e1dec52fd8533df9d7cd01c9b42ead19e4fa160d0d3f76d115827e9963
MD5 c21766f895762ab3221661a5151496e7
BLAKE2b-256 19da3983b7d5ed2e21f038058dbf389e975fba7d135194fc0de8c86904c27a2c

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.58.0-cp39-abi3-macosx_10_12_x86_64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.58.0 This release

6 files

0.57.0

6 files

0.56.0

6 files

0.55.0

6 files

0.54.0

6 files

0.53.0

6 files

0.52.0

6 files

0.51.0

6 files

0.50.0

6 files

0.34.0

6 files

0.32.0

6 files

0.31.0

6 files

0.30.0

6 files

0.29.0

6 files

0.28.0

6 files

0.27.0

6 files

0.26.0

6 files

0.25.0

6 files

0.24.1

6 files

0.24.0

6 files

0.23.1

6 files

0.23.0

6 files

0.22.0

6 files

0.21.0

6 files

0.20.0

6 files

0.19.0

6 files

0.18.0

6 files

0.17.0

6 files

0.16.2

6 files

0.16.1

6 files

0.15.0

6 files

0.14.0

6 files

0.13.0

6 files

0.12.0

6 files

0.11.0

6 files

0.10.0

6 files

0.9.0

6 files

0.8.0

6 files

0.7.1

6 files

0.7.0

6 files

0.1.1

6 files

0.1.0

6 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page