Topica: fast, all-in-one topic modeling for Python
topica is a fast, memory-efficient, all-in-one topic-modeling library for Python, built for social scientists. It brings together more than fifty models usually split across JVM tools like MALLET and R packages like stm, including LDA, STM, CTM, keyATM, BERTopic, and neural, dynamic, short-text, and embedding-based models, all under one NumPy-native API. Every model is validated against its reference implementation and reproducible from a fixed seed; all share one set of diagnostics, labeling, validation, and covariate-effect tools, so you learn a single workflow and it applies across the roster. It installs as a single wheel that needs only NumPy and pandas: no JVM, no PyTorch.
pip install topica
Quick start
Point topica at a DataFrame and read the topics. This runs exactly as written, on a bundled example dataset, right after install:
import topica
df = topica.datasets.load_gadarian() # bundled; loads offline
corpus = topica.from_dataframe(
df, text_col="open.ended.response", stopwords=topica.data.ENGLISH_STOPWORDS
)
model = topica.LDA(num_topics=5, seed=13)
model.fit(corpus) # sensible defaults; no tuning required
print(topica.summary(model)) # top words per topic
from_dataframe keeps your metadata aligned to the documents that survive pruning, so the same corpus feeds a structural topic model that relates topic prevalence to a covariate and reports the effect with uncertainty:
prevalence = corpus.metadata[["treatment"]] # a numeric DataFrame goes straight in
stm = topica.STM(num_topics=5, seed=13)
stm.fit(corpus, prevalence, prevalence_names=["treatment"])
draws = topica.effects.posterior_theta_samples(stm, nsims=30, seed=0)
effect = topica.effects.estimate_effect(draws, prevalence, feature_names=["treatment"])
Your own data is one line away: pass pandas.read_csv("yours.csv") to from_dataframe. See the getting-started guide and the worked examples for analyses end to end.
Fits are reproducible and validated: the variational models are bit-exact against their references, the samplers reproduce from a fixed seed and thread count, and every model is checked against its reference implementation (R stm, MALLET, keyATM, and more).
The core needs only NumPy and pandas. Optional extras add features without weighing it down: topica[viz] (matplotlib plots), topica[formula] (R-style formulas), topica[polars] (Polars frames), and topica[llm] (LLM labels and embeddings, OpenAI or local via ollama).
Models
Find your goal in the left column. Every model is validated against a reference implementation, so the choice is about fit to your research design, not quality. A specialized model is often the right first choice when your data calls for it.
Common openings
| If your goal is… | Start with | Also consider | First calls | Note |
|---|---|---|---|---|
| Explore themes with no prior structure | LDA |
NMF |
search_k(), topic_table() |
The default first pass. NMF is a fast, deterministic alternative. |
| Relate topics to metadata (author, date, party) | STM |
DMR |
estimate_effect(), one_hot(), spline() |
STM gives covariate effects with uncertainty; DMR is a lighter Gibbs prior. |
| Measure concepts you can name in advance | KeyATM |
SeededLDA |
KeyATM(keywords=…), .keyword_rate |
Anchor named topics with a few seed words each. |
| Very short documents: tweets, headlines, survey answers | GSDMM |
PT |
fit() |
One topic per document; standard LDA over-fragments short text. |
| Cluster by meaning using embeddings | BERTopic |
ETM |
fit(docs, doc_embeddings=…) |
Clustering, not a posterior: topic-proportion uncertainty and effect estimation behave differently than the models above. |
Specialized approaches. Start here when your design calls for one.
| If your data or goal is… | Start with | Also consider | First calls | Note |
|---|---|---|---|---|
| Topics shift over time slices | DTM |
DETM |
fit(docs, times=…) |
Prevalence and content evolve across periods; DETM adds embeddings. |
| Documents linked in a network (citations, replies) | RTM |
— | fit(docs, links=…) |
Models the text and the link graph jointly. |
| Documents in more than one language | PolylingualLDA |
— | fit(doc_tuples) |
Aligned topics across languages from translation-linked tuples. |
| Place authors or actors on an ideological scale | Wordfish |
TBIP |
fit(docs) |
Scaling from word usage; TBIP adds a text-based ideal-point prior. |
| How tone or sentiment varies with metadata | STS |
— | estimate_effect() |
Sentiment-discourse decomposition; reach for it when tone is the question. |
topica.list_models(common_start=True) returns the common openings in code;
list_models(group=…, brings=…) filters the full roster below.
All models (more than fifty, grouped by what you bring and what you want; click to expand)
Models are organized by what you bring and what you want, not by inference
family. The from topica import X namespace is flat; topica.list_models(group=…, brings=…, inference=…, determinism=…) filters this roster in code. Brings is
what you supply beyond raw text; Reproducibility is bit-exact (identical
regardless of thread count), seed-reproducible (identical from a fixed seed and
thread count), or llm-bounded.
Every model below is validated before it enters the roster: against a maintained reference implementation where one exists (MALLET, gensim, R stm, tomotopy, and the like), otherwise by planted recovery on a synthetic corpus with a known answer. See validation for where each model stands. The groupings are about fit to your research design, not quality: a specialized model is the right first choice when your data calls for it.
Common starting points
One per common goal, where most social scientists begin. BERTopic works differently from the others: it clusters document embeddings rather than fitting a posterior, so topic-proportion uncertainty and covariate-effect estimation do not carry over directly.
| Model | Brings | Inference | Reproducibility | Summary |
|---|---|---|---|---|
LDA |
text | gibbs | seed-reproducible | Classic latent Dirichlet allocation via a fast SparseLDA collapsed-Gibbs sampler. |
NMF |
text | matrix-factorization | bit-exact | Non-negative matrix factorization of the document-term matrix via multiplicative updates. |
STM |
text, metadata | variational | bit-exact | Structural topic model: relate topic prevalence and content to covariates. |
KeyATM |
text, seeds | gibbs | seed-reproducible | Keyword-assisted topics: anchor named topics with a few seed words each. |
GSDMM |
text | gibbs | seed-reproducible | Gibbs-sampling Dirichlet mixture: one topic per short document. |
BERTopic |
text, embeddings | clustering | seed-reproducible | Cluster document embeddings; label topics by class-based TF-IDF. |
Specialized approaches
The right first choice when your design calls for one: short text, change over time, document networks, multiple languages, ideological scaling, and more.
General-purpose
| Model | Brings | Inference | Reproducibility | Summary |
|---|---|---|---|---|
OnlineLDA |
text | variational | seed-reproducible | Online (streaming) variational-Bayes LDA (Hoffman et al. 2010): minibatch stochastic VB with a decaying learning rate and a streaming partial_fit; the gensim LdaModel analogue for very large or streaming corpora. |
CTM |
text | variational | bit-exact | Correlated topic model: a logistic-normal prior that lets topics co-occur. |
ProdLDA |
text | vae | seed-reproducible | Product-of-experts LDA (AVITM) for sharper, more coherent topics; hand-coded VAE. |
HDP |
text | gibbs | seed-reproducible | Hierarchical Dirichlet process: infers the number of topics from the data. |
LSA |
text | svd | seed-reproducible | Latent semantic analysis: a truncated SVD of the weighted document-term matrix. |
AnchorLDA |
text | matrix-factorization | bit-exact | Anchor-words spectral recovery (Arora et al. 2013): deterministic, Gibbs-free topics from the word co-occurrence matrix. |
PolylingualLDA |
text | gibbs | seed-reproducible | Polylingual topic model (Mimno et al. 2009): aligned topics across languages from document tuples that share one topic distribution. |
CorEx |
text | information-theoretic | seed-reproducible | Correlation Explanation: information-theoretic topic model that maximizes total correlation; supports anchor words. |
MGLDA |
text | gibbs | seed-reproducible | Multi-Grain LDA: global (document-level) + local (sliding-window aspect) topics with a per-token grain switch. For reviews / aspect extraction. |
TopicalNGrams |
text | gibbs | seed-reproducible | Topical N-Grams (Wang, McCallum & Wei 2007): an LDA extension that jointly discovers topics and topic-specific multiword phrases. A per-token bigram-status indicator, sampled with the topic, decides whether a token continues a phrase from the previous word given its topic, so phrase structure is learned during fitting rather than fixed beforehand. Exposes top_phrases alongside top_words. |
Covariates & structure
| Model | Brings | Inference | Reproducibility | Summary |
|---|---|---|---|---|
STS |
text, metadata | variational | bit-exact | Structural topic-and-sentiment model over document metadata. |
SAGE |
text, metadata | gibbs | seed-reproducible | Sparse additive generative model: the same topic worded differently across groups. |
DMR |
text, metadata | gibbs | seed-reproducible | Dirichlet-multinomial regression: a document-metadata prior on topic proportions. |
GDMR |
text, metadata | gibbs | seed-reproducible | Generalized DMR with a smooth (Legendre-basis) prior over continuous covariates. |
Scholar |
text, metadata, labels | vae | seed-reproducible | SCHOLAR (Card et al. 2018): a ProdLDA VAE with a covariate-shifted prevalence prior, an optional supervised label head, and optional content (topic-covariate) word deviations — neural STM prevalence + sLDA + SAGE. |
RTM |
text, links | variational | seed-reproducible | Relational topic model (Chang & Blei 2010): jointly models document text and a link graph (citations, hyperlinks, adjacency); predicts links from words and words from links. |
FactorialLDA |
text | gibbs | seed-reproducible | Factorial LDA (Paul & Dredze 2012): each token is a K-tuple of latent factors (e.g. topic x sentiment); structured word priors tie tuples sharing a component and a sparsity prior deactivates unsupported tuples. |
AuthorTopic |
text, metadata | gibbs | seed-reproducible | Author-Topic Model: each author has a topic distribution; documents mix their authors. Answers what an author writes about. |
AuthorRecipientTopic |
text, metadata | gibbs | seed-reproducible | Author-Recipient-Topic (McCallum et al. 2007): topics conditioned on the (sender, recipient) pair, for the language of a directed social network (who talks to whom about what). Realized over the AuthorTopic engine. |
Guided & supervised
| Model | Brings | Inference | Reproducibility | Summary |
|---|---|---|---|---|
SeededLDA |
text, seeds | gibbs | seed-reproducible | Seeded LDA: steer named topics toward supplied seed words. |
GuidedNMF |
text, seeds | matrix-factorization | seed-reproducible | Guided NMF: seed-word-guided semi-supervised NMF; the matrix-factorization analogue of SeededLDA. |
LabeledLDA |
text, labels | gibbs | seed-reproducible | Labeled LDA: each document label is a topic; tokens are restricted to its labels. |
SupervisedLDA |
text, labels | variational | seed-reproducible | Supervised LDA: topics shaped to predict a per-document real-valued response. |
DiscLDA |
text, labels | gibbs | seed-reproducible | Discriminative LDA (Lacoste-Julien et al. 2008): topics split into per-class and shared blocks; reads how classes talk differently. |
Short text
| Model | Brings | Inference | Reproducibility | Summary |
|---|---|---|---|---|
PT |
text | gibbs | seed-reproducible | Pseudo-document topic model: pool short texts into pseudo-documents. |
BTM |
text | gibbs | seed-reproducible | Biterm topic model: learns topics from corpus-level word co-occurrence (biterms). |
Dynamic & hierarchical
| Model | Brings | Inference | Reproducibility | Summary |
|---|---|---|---|---|
DTM |
text, times | variational | seed-reproducible | Dynamic topic model: a fixed topic set whose word distributions drift across time slices. |
DETM |
text, embeddings, times | vae | seed-reproducible | Dynamic embedded topic model: embedding-factored topics that drift across time slices, fit as an amortized VAE. |
TopicsOverTime |
text, times | gibbs | seed-reproducible | Topics over Time: LDA with a per-topic Beta density over continuous timestamps; each topic has a temporal peak. Descriptive continuous-time prevalence (not vocabulary drift). |
HLDA |
text | gibbs | seed-reproducible | Hierarchical LDA (nested CRP): a learned tree of super- and sub-topics. |
PA |
text | gibbs | seed-reproducible | Pachinko allocation: a DAG of super- and sub-topics. |
Embedding-based
| Model | Brings | Inference | Reproducibility | Summary |
|---|---|---|---|---|
KeyNMF |
text, embeddings | matrix-factorization | bit-exact | KeyNMF (Kristensen-McLachlan et al. 2024): NMF over an embedding-derived keyword-importance matrix. For each document it scores its words by the similarity between the document embedding and the word embedding, keeps the top-N positive, and factors that sparse doc-word matrix. The bridge between the count-based NMF family and the embedding backend; sparse, readable topics robust to short/noisy text. |
Top2Vec |
text, embeddings | clustering | seed-reproducible | Topics as dense regions in a joint document-word embedding space. |
SemanticSignalSeparation |
text, embeddings | ica | seed-reproducible | Topics as independent axes of semantic space (S3, Kardos et al. 2025): FastICA over the document embeddings, with each word's importance read off by projecting the vocabulary embeddings onto each axis. Signed poles. |
ETM |
text, embeddings | variational | seed-reproducible | Embedded topic model: topic-word distributions factored through word embeddings. |
GaussianLDA |
text, embeddings | gibbs | seed-reproducible | Gaussian LDA (Das, Zaheer & Dyer 2015): each topic is a Gaussian over the word-embedding space (Normal-Inverse-Wishart prior), so topics generalize over semantically similar words. Collapsed Gibbs with a Student-t posterior predictive and rank-1 Cholesky up/downdates. |
FASTopic |
text, embeddings | optimal-transport | seed-reproducible | Topics from optimal-transport plans between document, topic, and word embeddings. |
CombinedTM |
text, embeddings | vae | seed-reproducible | Contextualized ProdLDA: encoder reads the bag of words plus a document embedding. |
ZeroShotTM |
text, embeddings | vae | seed-reproducible | Contextualized ProdLDA: encoder reads the document embedding alone, enabling cross-lingual transfer. |
InfoCTM |
text, dictionary | vae | seed-reproducible | Cross-lingual: two ProdLDA models aligned by a bilingual dictionary through a mutual-information term. |
Ideal point
| Model | Brings | Inference | Reproducibility | Summary |
|---|---|---|---|---|
Wordfish |
text | em | bit-exact | Poisson scaling (Slapin & Proksch 2008): an unsupervised one-dimensional ideal-point estimate from word frequencies alone, no topics. The word-frequency baseline companion to IdealPointTM. |
Wordshoal |
text, metadata | em | bit-exact | Multi-domain scaling (Lauderdale & Herzog 2016): scales each debate/domain with Wordfish, then combines the within-domain positions into one cross-domain actor scale via a linear factor model. The multi-domain extension of Wordfish, for speeches carrying trusted debate labels. |
TBIP |
text | variational | seed-reproducible | Text-Based Ideal Points (Vafa, Naidu & Blei 2020): a Poisson factorization whose neutral topic-word intensities are rescaled by a per-word ideological factor exp(x_s * eta_kv), with the author position x_s latent. Fit by the paper's mean-field variational inference (reparameterized SVI). Recovers ideological scales from unlabeled text. |
PartyEmbeddings |
text, metadata | neural-embedding | seed-reproducible | Party embeddings (Rheault & Cochrane 2020): a PV-DM paragraph-vector model trained by negative sampling with party-period metadata tags; the leading principal components of the learned party vectors give the ideological scale, and words share the space so a party's language can be read off by proximity. The corpus-trained word-embedding member of the ideal-point family. |
LLM-based
| Model | Brings | Inference | Reproducibility | Summary |
|---|---|---|---|---|
TopicGPT |
text, llm | prompting | llm-bounded | LLM-driven topic discovery: prompt a model to propose, refine, and assign a topic taxonomy with descriptions. |
Experimental
Not (yet) on the validated roster, on one of two grounds: the model is unpublished (a topica original with no paper), or it is a published method whose benefit has not held up against a simpler baseline. A published method with faithful inference stays validated even when its basis is planted-recovery, so planted validation alone does not land a model here. Gated: call topica.enable_experimental() (or set TOPICA_EXPERIMENTAL=1) before use. These may change or be removed without a deprecation cycle. For the triple gate a model clears to graduate to the validated roster, and where every model stands on validation evidence, see Validation & graduation.
| Model | Brings | Inference | Reproducibility | Summary |
|---|---|---|---|---|
TensorLDA |
text | svd | seed-reproducible | Online Tensor LDA (Kangaslahti et al. 2026): deterministic method-of-moments topic modeling via second and third-order cumulants. |
NarrativeTM |
text | gibbs | seed-reproducible | Intra-document narrative trajectory model: captures how topic prevalence shifts across the progress of a text. |
CSATM |
text, links | gibbs | seed-reproducible | Conversational Structure Aware TM (Sun et al. 2020): weights each comment's tokens by a reply-tree 'popularity' score and, after Gibbs, smooths each comment's topics toward its ancestors along the reply path ('transitivity'). For threaded forum data (posts + nested comments). Ported from the paper (no reference implementation); validated by planted recovery + LDA reduction. |
ThreadTM |
text, links | variational | seed-reproducible | ThreadTM (topica-original): a reply-threaded topic model — CTM logistic-normal topics with a reply-tree structured prior, so a reply's topic prior is coupled to the comment it answers (a persistence-smoothing prior, reverting toward its covariate-group baseline). Supports a prevalence covariate and an optional content covariate (topic words shift by depth/level, à la STM's content model). For threaded discussion (posts + nested comments). Validated by planted recovery and a synthetic held-out-beat gate (the parent's topics predict held-out leaf tokens better than the no-tree baseline on persistence-structured data). The real-corpus benefit is genre-dependent; read kappa_ci before claiming persistence. |
IdealPointTM |
text, embeddings | variational | seed-reproducible | Topic model with a latent ideal-point head: each author gets a low-dimensional position that shifts within-topic word choice, with a per-topic discrimination. Consumes word tokens as counts (Wordfish with topics) or, when word embeddings are supplied to fit, factored through them as in ETM. The unsupervised, latent-trait twin of the STM content covariate. |
IdealPointSentenceTM |
text, embeddings | em | seed-reproducible | Continuous ideal-point topic model over sentence/document embeddings: topics are Gaussian clusters whose centroids are displaced by a latent author position. The sentence-embedding sibling of IdealPointTM, fit by EM. |
EmbeddingLDA |
text, embeddings | gibbs | seed-reproducible | LDA anchored by pre-trained embeddings: k-means clusters the vocabulary embeddings, seeds each topic with the words nearest a cluster centroid, and (optionally) biases each document's mixture toward its own embedding. A topica original; validated by planted-recovery only. |
Every model exposes the same shape: fit(docs, …), then topic_word (φ), doc_topic (θ), top_words(n), and save/load, so one diagnostic, labeling, and effect-estimation stack applies to all of them and a new model inherits it for free. The embedding-based models take document vectors from any embedder (sentence-transformers, an API, or a local model such as ollama; no PyTorch or UMAP/numba in the wheel). Full guides: the models and embedding topics.
Diagnostics & analysis
Model-agnostic: they work on any fitted model's topic_word/doc_topic:
- Quality:
coherence(u_mass,c_v,c_uci,c_npmi; co-occurrence counting in the Rust core),exclusivity,topic_diversity,quality_frontier - Labeling:
label_topics(prob / FREX / lift / score),frex,relevance,find_thoughts,topic_table,summary - Validation:
word_intrusion,document_intrusion,bootstrap_stability,search_k - Reliability:
select_model(fit many seeds) andensemble(combine runs into a consensus more reliable than any single fit — cluster/align/stable methods, the last derived from gensim'sEnsembleLda) - Comparison:
fighting_words(weighted log-odds) for contrasting corpora - Covariate effects:
estimate_effect(method of composition, cluster-robust SEs, GLM links),topic_correlation, and the design helpersone_hot,spline, andinteraction(intopica.effects/topica.design; they build covariate bases for any model's design matrix);posterior_theta_samplesdraws θ for the logistic-normal models (STM/CTM) - Preprocessing:
tokenize,learn_phrases/apply_phrases,split_documents, theCorpusclass
See diagnostics and covariate effects.
Performance
topica runs on a parallel Rust core, so the whole roster is fast and every fit is reproducible from a fixed seed. Core for core it matches the hand-tuned compiled samplers: parity with Java MALLET on plain LDA and with the C++ keyATM on keyword models. For the structural and other variational models it is several times faster than R stm, the single-threaded field standard. Fit to convergence (both at the same emtol, spectral start), on real corpora:
| Model | Reference | topica speedup (to convergence) |
|---|---|---|
| STM | R stm |
1.7–2.7× single-threaded, ~5–7× multicore |
| LDA | Java MALLET | parity single-threaded; multithread speedup grows with corpus size |
| keyATM | R keyATM |
parity single-threaded, ~2× multithreaded |
topica also fits in about a quarter of R stm's memory (≈180MB against ≈675MB at 5,000 documents). For the approximate parallel Gibbs samplers the multithreaded speedup grows with corpus size: the per-sweep count-table merge is fixed overhead, so larger corpora amortize it over more sampling work. LDA's eight-core speedup over MALLET runs about 3× at 2,000 documents and reaches ~4× at 5,000.
Every fit is reproducible from a fixed seed and validated against its reference. See Benchmarks for the full methodology; reproduce the structural-model table with python benchmarks/bench_stm_convergence.py and the size-varying LDA curve with python benchmarks/speed_vs_size.py.
Install from source
pip install maturin
git clone https://github.com/nealcaren/topica && cd topica
python -m venv .venv && source .venv/bin/activate
maturin develop --release --features python
Requires numpy >= 1.21. Use --release (the debug build is much slower).
Acknowledgements
Topica was inspired by a post from David Mimno about porting Java MALLET to Rust. As a long-time Python user, I had long been jealous of the topic-modeling tools available in other languages; this seemed like an opportunity to make those capabilities easier to use in Python for me and for others.
As such, Topica stands on a generation of open topic-modeling research and code. Each entry below lists the reference, its authors and year, and the topica class(es) it underlies; the other models are Rust ports or reimplementations, validated against these reference implementations.
- MALLET (McCallum, 2002) —
LDA,DMR,LabeledLDA: the SparseLDA sampler, Dirichlet-multinomial regression, and hyperparameter optimization.LDAbegan as a port of David Mimno's RustMallet (Apache-2.0) and follows its SparseLDA sampler and fixed-point optimizer closely, but uses its own RNG (PCG), so it is not byte-identical to RustMallet. Against Java MALLET (also a different RNG) it recovers the same topics on a planted corpus (cosine 1.000) - stm (Roberts, Stewart & Tingley, 2019) —
STM,CTM,SAGE: variational EM,estimateEffect,searchK, FREX, spectral initialization, and the method of composition - sts (Chen & Mankad, 2024) —
STS: the Structural Topic and Sentiment-Discourse model — the joint prevalence/sentiment Laplace E-step and the Poisson topic-word M-step, validated against the package - lda-c / ctm-c / dtm and hdp (Blei lab, 2006–2007) —
CTM,DTM,HDP: the CTM, Dynamic Topic Model, and HDP samplers - gensim (Řehůřek & Sojka, 2010) —
DTM,ensemble,OnlineLDA: the coherence-pipeline conventions (thecoherence_type=API and default sliding windows; the measures themselves are Röder et al. 2015 and Mimno et al. 2011), theLdaSeqModelDTM reference, theEnsembleLda(CBDBSCAN stable-topic) method thatensemble(method="stable")derives from (matching it on well-separated inputs, with two documented improvements on edge cases where gensim degenerates), and theLdaModelonline-VB reference forOnlineLDA - onlineldavb (Hoffman, Blei & Bach, 2010) —
OnlineLDA: online (streaming) variational Bayes for LDA — the minibatch stochastic-VB E-step, the decaying Robbins-Monro learning rate, and the streamingpartial_fit; written from the paper and validated againstonlineldavb.pyand gensim'sLdaModelas external oracles (both are copyleft, so no code was copied) - tomotopy (bab2min, 2020) — API conventions (
summary, the short-text models), andGDMR(generalized DMR; Lee & Song, 2020), validated against itsGDMRModel - scikit-learn (Pedregosa et al., 2011) —
NMF: the multiplicative-update solver (Lee & Seung, 2001) and the NNDSVD initialization (Boutsidis & Gallopoulos, 2008), validated againstsklearn.decomposition.NMF(BSD-3-Clause); andLSA: latent semantic analysis / indexing (Deerwester et al., 1990), validated againstsklearn.decomposition.TruncatedSVD(BSD-3-Clause) including itssvd_flipsign convention. The numerics are reimplemented in Rust; the randomized truncated SVD shared by both (it seeds NMF's NNDSVD and is the LSA factorization itself) follows Halko et al. (2011). - keyATM (Eshima, Imai & Sasaki, 2024) —
KeyATM: the base, covariate, and dynamic models, the information-theory token weighting, and the Chib (1998) change-point HMM, validated against the package - seededlda (Watanabe, 2023) —
SeededLDA: the corpus-frequency-scaled seed prior (count × weight × 100), validated against the package's seed matrix and seeded topics - LightLDA (Yuan et al., 2015) —
LDA: the alias-table Metropolis-Hastings sampler - GSDMM (Yin & Wang, 2014) —
GSDMM: the movie-group-process mixture for short text - BTM (Yan, Guo, Lan & Cheng, 2013; R package by Jan Wijffels) —
BTM: the biterm co-occurrence topic model for short text - Polylingual Topic Models (Mimno, Wallach, Naradowsky, Smith & McCallum, 2009) —
PolylingualLDA: LDA over aligned document tuples that share one topic distribution, giving topics aligned across many languages; validated against MALLET'sPolylingualTopicModelas a black-box oracle - DiscLDA (Lacoste-Julien, Sha & Jordan, 2008) —
DiscLDA: discriminative LDA with per-class and shared topic blocks; the fixed block-transform variant, validated against the paper's 20 Newsgroups feature-classification result (no reference implementation exists, so it is paper-derived) - Factorial LDA (Paul & Dredze, 2012) —
FactorialLDA: sparse multi-dimensional topics, where each token is a K-tuple of latent factors tied by structured log-linear priors; implemented from the paper's mathematics (the reference Java is GPL and non-reproducible, so the port is certified by finite-difference gradient and factor-tying tests plus planted recovery, not seed parity) - ProdLDA / AVITM (Srivastava & Sutton, 2017) —
ProdLDA: autoencoding variational inference and the product-of-experts word model - SCHOLAR (Card, Tan & Smith, 2018; reference dallascard/scholar, Apache-2.0) —
Scholar: metadata in a ProdLDA VAE — a covariate-dependent topic-prevalence prior (neural STM prevalence), an optional supervised label head (neural sLDA), and optional content/topic-covariate word deviations (neural SAGE), on topica's ProdLDA backbone, validated against the reference as a numerical oracle - BERTopic (Grootendorst, 2022) and Top2Vec (Angelov, 2020) —
BERTopic,Top2Vec: the embedding-clustering pipeline, class-based TF-IDF, and thereduce → cluster → representdesign - CETopic / topicx (Zhang, Fang, Chen & Namazi-Rad, NAACL 2022, MIT) —
BERTopic(weighting="tfidf-idf"): the TFIDF×IDF_i topic-word selection scheme (a corpus-level TF-IDF averaged per cluster times a cross-cluster IDF penalty), ported faithfully to the reference's scikit-learn defaults - S³ / turftopic (Kardos, Kostkan, Enevoldsen, Vermillet, Nielbo & Rocca, ACL 2025, MIT) —
SemanticSignalSeparation: Semantic Signal Separation, FastICA over contextual document embeddings with topic words read off by projecting the vocabulary embeddings onto each independent axis, ported faithfully to the reference's scikit-learn FastICA defaults - ETM (Dieng, Ruiz & Blei, 2020) —
ETM: the Embedded Topic Model (per-document variational EM and an amortized VAE) - DETM (Dieng, Ruiz & Blei, 2019) —
DETM: the Dynamic Embedded Topic Model (structured amortized variational inference with a hand-coded LSTM) - FASTopic (Wu et al., 2024) —
FASTopic: the optimal-transport topic model - contextualized-topic-models (Bianchi et al., MIT) —
CombinedTM(Bianchi, Terragni & Hovy, 2021) andZeroShotTM(Bianchi, Nozza & Hovy, 2021): ProdLDA encoders that read a contextual document embedding, alongside or in place of the bag of words - quanteda.textmodels (Benoit et al., 2018) —
Wordfish: the Slapin & Proksch (2008) Poisson scaling model, validated against itstextmodel_wordfish(the recovered scale and the analytic position standard errors both match at correlation 1.00 on a corpus sampled from the model) - tbip (Vafa, Naidu & Blei, 2020) —
TBIP: Text-Based Ideal Points; the official implementation is TensorFlow 1.x, so topica reimplements the published model and its mean-field variational inference, validated against an independent PyTorch reference - TensorLy TLDA (Kangaslahti, Ebanks, Kossaifi, Liu, Alvarez & Anandkumar, 2026) —
TensorLDA: online tensor latent Dirichlet allocation via second- and third-order moments. The Rust implementation is experimental; see the TensorLDA validation record for its current evidence and limitations. - partyembed (Rheault & Cochrane, 2020) —
PartyEmbeddings: party embeddings via a PV-DM paragraph-vector model with party-period metadata tags, placed by PCA of the learned party vectors. The reference builds on gensim'sDoc2Vec; topica reimplements the PV-DM negative-sampling training in Rust (from Mikolov et al. 2013 and Le & Mikolov 2014) and is validated against thatDoc2Vecscale (correlation 1.00 on a planted ordering) - CLNTM (Nguyen & Luu, 2021) — the InfoNCE contrastive regularization on topic vectors offered by the
contrastive=flag on the VAE models - WHAI / Weibull-Dirichlet VAE (Zhang et al., 2018; Burkhardt & Kramer, 2019) — the Weibull-reparameterized Dirichlet prior offered by
prior="dirichlet"on the VAE models - Neural variational topic models with alternative priors (Miao, Grefenstette & Blunsom, 2017; Nalisnick & Smyth, 2017) — the Gaussian stick-breaking prior offered by
prior="stick_breaking"on the VAE models - TopicGPT (Pham et al., NAACL 2024, MIT) —
TopicGPT: the generate / refine / assign prompt flow for LLM-driven topic discovery
The embedding-native models build on two pure-Rust crates: petal-clustering for HDBSCAN and umap-rs for the optional UMAP reducer, both BLAS-free.
Full citations for every model and reference implementation, and how to cite topica, are on the Citing page.
Contributing, tests, and support
Contributions are welcome. See CONTRIBUTING for the development setup and workflow, CONTRIBUTING-MODELS for adding a new topic model, and the conventions guide for the cross-model naming and API contract. All participants are expected to follow our Code of Conduct.
To run the test suite after a source build:
cargo test --lib # Rust unit tests
python -m pytest tests/ -q # Python tests
mkdocs build --strict # docs build clean
Every push runs these on CI (see the badge above). The parity/ checks
validate models against their reference implementations (R stm, keyATM,
MALLET); they skip cleanly when those toolchains are not installed.
- Report a bug or request a feature: open an issue.
- Ask a question or share how you are using topica: start a discussion.
topica is maintained by Neal Caren. Issues and pull requests are triaged on a best-effort basis.
Citation
If you use topica in published work, please cite it. GitHub's Cite this
repository button (top right) generates a formatted reference from
CITATION.cff. A software paper is in preparation; until it
appears, cite the software release:
@software{caren_topica,
author = {Caren, Neal},
title = {topica: fast, all-purpose topic modeling for Python},
year = {2026},
url = {https://github.com/nealcaren/topica},
version = {0.54.0}
}
Replace version with the release you used. For the individual models and their
reference implementations, see the Citing
page.
License
Apache-2.0 — see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file topica-0.57.0.tar.gz.
File metadata
- Download URL: topica-0.57.0.tar.gz
- Upload date:
- Size: 20.3 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c3b68be0f61b29e8179ddf81abb38f672413b1c51660b26ea23c1460b7aaedf1
|
|
| MD5 |
7c41165c335d578b9635a5121cf7d5a9
|
|
| BLAKE2b-256 |
46b1b659d476ce81a1f81fdf4d964ec409180ea883338dc2cc58715894708136
|
Provenance
The following attestation bundles were made for topica-0.57.0.tar.gz:
Publisher:
CI.yml on nealcaren/topica
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
topica-0.57.0.tar.gz -
Subject digest:
c3b68be0f61b29e8179ddf81abb38f672413b1c51660b26ea23c1460b7aaedf1 - Sigstore transparency entry: 2701139985
- Sigstore integration time:
-
Permalink:
nealcaren/topica@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Branch / Tag:
refs/tags/v0.57.0 - Owner: https://github.com/nealcaren
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
CI.yml@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Trigger Event:
push
-
Statement type:
File details
Details for the file topica-0.57.0-cp39-abi3-win_amd64.whl.
File metadata
- Download URL: topica-0.57.0-cp39-abi3-win_amd64.whl
- Upload date:
- Size: 5.4 MB
- Tags: CPython 3.9+, Windows x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f4c129b87940e8e22a6089c1d97cecce814842bd07a24cb05256975cc4febc4d
|
|
| MD5 |
8acad18bf67c24e483750ba8e1be8605
|
|
| BLAKE2b-256 |
bacb39eb981c6573e55351a734b15ce870e56f14a3c89ca8df927693be895143
|
Provenance
The following attestation bundles were made for topica-0.57.0-cp39-abi3-win_amd64.whl:
Publisher:
CI.yml on nealcaren/topica
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
topica-0.57.0-cp39-abi3-win_amd64.whl -
Subject digest:
f4c129b87940e8e22a6089c1d97cecce814842bd07a24cb05256975cc4febc4d - Sigstore transparency entry: 2701140085
- Sigstore integration time:
-
Permalink:
nealcaren/topica@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Branch / Tag:
refs/tags/v0.57.0 - Owner: https://github.com/nealcaren
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
CI.yml@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Trigger Event:
push
-
Statement type:
File details
Details for the file topica-0.57.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.
File metadata
- Download URL: topica-0.57.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
- Upload date:
- Size: 5.2 MB
- Tags: CPython 3.9+, manylinux: glibc 2.17+ x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e82c5f8e7f0c95336c0f864f27a9d2276b8fee9e22eba23e9a9d06db49931fed
|
|
| MD5 |
a6444cbbd09eca7267db45526006c60e
|
|
| BLAKE2b-256 |
3b4af5d1c297af98db6028bac496ff367fe576fa3c0b7aa9f504b3f6607ee410
|
Provenance
The following attestation bundles were made for topica-0.57.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:
Publisher:
CI.yml on nealcaren/topica
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
topica-0.57.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl -
Subject digest:
e82c5f8e7f0c95336c0f864f27a9d2276b8fee9e22eba23e9a9d06db49931fed - Sigstore transparency entry: 2701140011
- Sigstore integration time:
-
Permalink:
nealcaren/topica@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Branch / Tag:
refs/tags/v0.57.0 - Owner: https://github.com/nealcaren
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
CI.yml@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Trigger Event:
push
-
Statement type:
File details
Details for the file topica-0.57.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.
File metadata
- Download URL: topica-0.57.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
- Upload date:
- Size: 4.9 MB
- Tags: CPython 3.9+, manylinux: glibc 2.17+ ARM64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
28fe1efeab60df1977c9697930e345b4514c2eec837ee7e11ef49a78d13ab217
|
|
| MD5 |
ba7948aa595471647fe78a5d9496009b
|
|
| BLAKE2b-256 |
f215430085e9c6fa7c302bfc1c72b3ceab1b06732883321fe42c384663ad2162
|
Provenance
The following attestation bundles were made for topica-0.57.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:
Publisher:
CI.yml on nealcaren/topica
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
topica-0.57.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl -
Subject digest:
28fe1efeab60df1977c9697930e345b4514c2eec837ee7e11ef49a78d13ab217 - Sigstore transparency entry: 2701140035
- Sigstore integration time:
-
Permalink:
nealcaren/topica@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Branch / Tag:
refs/tags/v0.57.0 - Owner: https://github.com/nealcaren
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
CI.yml@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Trigger Event:
push
-
Statement type:
File details
Details for the file topica-0.57.0-cp39-abi3-macosx_11_0_arm64.whl.
File metadata
- Download URL: topica-0.57.0-cp39-abi3-macosx_11_0_arm64.whl
- Upload date:
- Size: 4.7 MB
- Tags: CPython 3.9+, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
bde87fde5038f9aa184458891d0218ff743de0f6f7a3842d8a15f983262cc852
|
|
| MD5 |
5d91092190ebaf53b6c357a9b4f200d2
|
|
| BLAKE2b-256 |
4c286f3c02ce25d6496af9813766f587bebe728de0f01ad6c8f5820152345415
|
Provenance
The following attestation bundles were made for topica-0.57.0-cp39-abi3-macosx_11_0_arm64.whl:
Publisher:
CI.yml on nealcaren/topica
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
topica-0.57.0-cp39-abi3-macosx_11_0_arm64.whl -
Subject digest:
bde87fde5038f9aa184458891d0218ff743de0f6f7a3842d8a15f983262cc852 - Sigstore transparency entry: 2701140061
- Sigstore integration time:
-
Permalink:
nealcaren/topica@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Branch / Tag:
refs/tags/v0.57.0 - Owner: https://github.com/nealcaren
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
CI.yml@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Trigger Event:
push
-
Statement type:
File details
Details for the file topica-0.57.0-cp39-abi3-macosx_10_12_x86_64.whl.
File metadata
- Download URL: topica-0.57.0-cp39-abi3-macosx_10_12_x86_64.whl
- Upload date:
- Size: 5.0 MB
- Tags: CPython 3.9+, macOS 10.12+ x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ae8c0a1ad02546d39433279357f79365beb4d5dbce29ee5fa0c42a0ce6a69b30
|
|
| MD5 |
3dc6c945b46a4e32ee28b9d78442bbdc
|
|
| BLAKE2b-256 |
315adc157ed3abc743cb620acb0c4ca43f2e7445c11012077e995e190e702512
|
Provenance
The following attestation bundles were made for topica-0.57.0-cp39-abi3-macosx_10_12_x86_64.whl:
Publisher:
CI.yml on nealcaren/topica
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
topica-0.57.0-cp39-abi3-macosx_10_12_x86_64.whl -
Subject digest:
ae8c0a1ad02546d39433279357f79365beb4d5dbce29ee5fa0c42a0ce6a69b30 - Sigstore transparency entry: 2701140099
- Sigstore integration time:
-
Permalink:
nealcaren/topica@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Branch / Tag:
refs/tags/v0.57.0 - Owner: https://github.com/nealcaren
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
CI.yml@a42fe8cfdbc95a96f0f3a1014b17516003dc9e7c -
Trigger Event:
push
-
Statement type: