Skip to main content

Topica: fast, all-purpose topic modeling for Python — a Rust core for LDA, STM, and more

Project description

Topica: fast, all-purpose topic modeling for Python

PyPI CI Docs License: Apache-2.0

topica is a fast, all-purpose topic-modeling library for Python, built for computational social scientists who want to go from a column of text to publishable results in one workflow. It brings together models usually split across JVM tools like MALLET and R packages like stm, more than two dozen in all (LDA, STM, CTM, plus neural, dynamic, and embedding-based models), each paired with the validation, covariate-effect, and reporting tools reviewers expect. Where general toolkits like Gensim or BERTopic give you topics, topica is built around the question social scientists ask of them next: how topic prevalence and content relate to covariates, with reference-validated models and reproducible fits. It installs as a single wheel that needs only NumPy and pandas: no JVM, no PyTorch.

pip install topica

Quick start

Point topica at a DataFrame and read the topics. This runs exactly as written, on a bundled example dataset, right after install:

import topica

df = topica.datasets.load_gadarian()          # bundled; loads offline
corpus = topica.from_dataframe(
    df, text_col="open.ended.response", stopwords=topica.ENGLISH_STOPWORDS
)

model = topica.LDA(num_topics=5, seed=42)
model.fit(corpus)                             # sensible defaults; no tuning required
print(topica.summary(model))                  # top words per topic

from_dataframe keeps your metadata aligned to the documents that survive pruning, so the same corpus feeds a structural topic model that relates topic prevalence to a covariate, with an honest hypothesis test:

prevalence = corpus.metadata[["treatment"]].to_numpy(float)

stm = topica.STM(num_topics=5, seed=42)
stm.fit(corpus, prevalence, prevalence_names=["treatment"])

draws  = topica.posterior_theta_samples(stm, nsims=30, seed=0)
effect = topica.estimate_effect(draws, prevalence, feature_names=["treatment"])

Your own data is one line away: pass pandas.read_csv("yours.csv") to from_dataframe. See the getting-started guide and the worked examples for analyses end to end.

Fits are reproducible and validated: the variational models are identical to the bit, the samplers reproduce from a fixed seed and thread count, and every model is checked against its reference implementation (R stm, MALLET, keyATM, and more).

The core needs only NumPy and pandas. Optional extras add features without weighing it down: topica[viz] (matplotlib plots), topica[formula] (R-style formulas), topica[polars] (Polars frames), and topica[llm] (LLM labels and embeddings, OpenAI or local via ollama).

Models

Starting out? LDA for general topics, STM to relate topics to covariates, HDP to let the data choose the number of topics, and BERTopic or CombinedTM for embedding-based topics. The full roster follows.

All models (more than two dozen, grouped by what you bring and what you want; click to expand)

Models are organized by what you bring and what you want, not by inference family. The from topica import X namespace is flat; topica.list_models(group=…, brings=…, inference=…, determinism=…) filters this roster in code. Brings is what you supply beyond raw text; Reproducibility is bit-exact (identical regardless of thread count), seed-reproducible (identical from a fixed seed and thread count), or llm-bounded.

General-purpose

Model Brings Inference Reproducibility Summary
LDA text gibbs seed-reproducible Classic latent Dirichlet allocation via a fast SparseLDA collapsed-Gibbs sampler.
CTM text variational bit-exact Correlated topic model: a logistic-normal prior that lets topics co-occur.
ProdLDA text vae seed-reproducible Product-of-experts LDA (AVITM) for sharper, more coherent topics; hand-coded VAE.
HDP text gibbs seed-reproducible Hierarchical Dirichlet process: infers the number of topics from the data.
NMF text matrix-factorization bit-exact Non-negative matrix factorization of the document-term matrix via multiplicative updates.
LSA text svd seed-reproducible Latent semantic analysis: a truncated SVD of the weighted document-term matrix.

Covariates & structure

Model Brings Inference Reproducibility Summary
STM text, metadata variational bit-exact Structural topic model: relate topic prevalence and content to covariates.
STS text, metadata variational bit-exact Structural topic-and-sentiment model over document metadata.
SAGE text, metadata gibbs seed-reproducible Sparse additive generative model: the same topic worded differently across groups.
DMR text, metadata gibbs seed-reproducible Dirichlet-multinomial regression: a document-metadata prior on topic proportions.
GDMR text, metadata gibbs seed-reproducible Generalized DMR with a smooth (Legendre-basis) prior over continuous covariates.

Guided & supervised

Model Brings Inference Reproducibility Summary
KeyATM text, seeds gibbs seed-reproducible Keyword-assisted topics: anchor named topics with a few seed words each.
SeededLDA text, seeds gibbs seed-reproducible Seeded LDA: steer named topics toward supplied seed words.
LabeledLDA text, labels gibbs seed-reproducible Labeled LDA: each document label is a topic; tokens are restricted to its labels.
SupervisedLDA text, labels gibbs seed-reproducible Supervised LDA: topics shaped to predict a per-document real-valued response.

Short text

Model Brings Inference Reproducibility Summary
GSDMM text gibbs seed-reproducible Gibbs-sampling Dirichlet mixture: one topic per short document.
PT text gibbs seed-reproducible Pseudo-document topic model: pool short texts into pseudo-documents.

Dynamic & hierarchical

Model Brings Inference Reproducibility Summary
DTM text, times variational seed-reproducible Dynamic topic model: a fixed topic set whose word distributions drift across time slices.
DETM text, embeddings, times vae seed-reproducible Dynamic embedded topic model: embedding-factored topics that drift across time slices, fit as an amortized VAE.
HLDA text gibbs seed-reproducible Hierarchical LDA (nested CRP): a learned tree of super- and sub-topics.
PA text gibbs seed-reproducible Pachinko allocation: a DAG of super- and sub-topics.

Embedding-based

Model Brings Inference Reproducibility Summary
BERTopic text, embeddings clustering seed-reproducible Cluster document embeddings; label topics by class-based TF-IDF.
Top2Vec text, embeddings clustering seed-reproducible Topics as dense regions in a joint document-word embedding space.
ETM text, embeddings variational seed-reproducible Embedded topic model: topic-word distributions factored through word embeddings.
FASTopic text, embeddings optimal-transport seed-reproducible Topics from optimal-transport plans between document, topic, and word embeddings.
EmbeddingLDA text, embeddings, seeds gibbs seed-reproducible Seeded LDA whose seed sets are expanded with nearest neighbors in an embedding space.
CombinedTM text, embeddings vae seed-reproducible Contextualized ProdLDA: encoder reads the bag of words plus a document embedding.
ZeroShotTM text, embeddings vae seed-reproducible Contextualized ProdLDA: encoder reads the document embedding alone, enabling cross-lingual transfer.
InfoCTM text, dictionary vae seed-reproducible Cross-lingual: two ProdLDA models aligned by a bilingual dictionary through a mutual-information term.

LLM-based

Model Brings Inference Reproducibility Summary
TopicGPT text, llm prompting llm-bounded LLM-driven topic discovery: prompt a model to propose, refine, and assign a topic taxonomy with descriptions.

Experimental

Shipped before a published paper and reference-implementation parity (topica's bar for a validated model). Gated: call topica.enable_experimental() (or set TOPICA_EXPERIMENTAL=1) before use. These may change or be removed without a deprecation cycle.

Model Brings Inference Reproducibility Summary
ECTM text, metadata, times variational bit-exact Evolving content topic model: STM content covariates that vary by group and drift across time periods.

Every model exposes the same shape: fit(docs, …), then topic_word (φ), doc_topic (θ), top_words(n), and save/load, so one diagnostic, labeling, and effect-estimation stack applies to all of them and a new model inherits it for free. The embedding-based models take document vectors from any embedder (sentence-transformers, an API, or a local model such as ollama; no PyTorch or UMAP/numba in the wheel). Full guides: the models and embedding topics.

Diagnostics & analysis

Model-agnostic: they work on any fitted model's topic_word/doc_topic:

  • Quality: coherence (u_mass, c_v, c_uci, c_npmi; co-occurrence counting in the Rust core), exclusivity, topic_diversity, quality_frontier
  • Labeling: label_topics (prob / FREX / lift / score), frex, relevance, find_thoughts, topic_table, summary
  • Validation: word_intrusion, document_intrusion, bootstrap_stability, search_k
  • Reliability: select_model (fit many seeds) and ensemble (combine runs into a consensus more reliable than any single fit — cluster/align/stable methods, the last a gensim EnsembleLda port)
  • Comparison: fighting_words (weighted log-odds) for contrasting corpora
  • Covariate effects: estimate_effect (method of composition, cluster-robust SEs, GLM links), topic_correlation, and the design helpers one_hot, spline, and interaction (all top level; they build covariate bases for any model's design matrix); posterior_theta_samples draws θ for the logistic-normal models (STM/CTM)
  • Preprocessing: tokenize, learn_phrases / apply_phrases, split_documents, the Corpus class

See diagnostics and covariate effects.

Performance

topica runs on a parallel Rust core. It is several times faster than R stm — the single-threaded field standard — for the structural and other variational models, and it matches the hand-tuned compiled samplers core for core: parity with Java MALLET on plain LDA and with the C++ keyATM on keyword models. On the political-blog corpus (2,000 documents, fit time only, same iterations on both sides):

Model Reference topica speedup
STM R stm 4–10× single-threaded, ~11–32× multithreaded
LDA Java MALLET parity single-threaded; multithread speedup grows with corpus size
keyATM R keyATM parity single-threaded, ~2× multithreaded

For the approximate parallel Gibbs samplers the multithreaded speedup grows with corpus size: the per-sweep count-table merge is fixed overhead, so larger corpora amortize it over more sampling work. LDA's eight-core speedup over MALLET runs about 3× at 2,000 documents and reaches ~4× at 5,000, so the small-corpus figures above understate what large-corpus users see.

Every fit is reproducible from a fixed seed and validated against its reference. See Benchmarks for the full methodology; reproduce the 2,000-document table with python benchmarks/speed_vs_r.py and the size-varying curve with python benchmarks/speed_vs_size.py.

Install from source

pip install maturin
git clone https://github.com/nealcaren/topica && cd topica
python -m venv .venv && source .venv/bin/activate
maturin develop --release --features python

Requires numpy >= 1.21. Use --release (the debug build is much slower).

Acknowledgements

Topica stands on a generation of open topic-modeling research and code. Each entry below lists the reference, its authors and year, and the topica class(es) it underlies; the other models are Rust ports or reimplementations, validated against these reference implementations.

  • MALLET (McCallum, 2002) — LDA, DMR, LabeledLDA: the SparseLDA sampler, Dirichlet-multinomial regression, and hyperparameter optimization. LDA began as a port of David Mimno's RustMallet (Apache-2.0) and follows its SparseLDA sampler and fixed-point optimizer closely, but uses its own RNG (PCG), so it is not byte-identical to RustMallet. Against Java MALLET (also a different RNG) it recovers the same topics on a planted corpus (cosine 1.000)
  • stm (Roberts, Stewart & Tingley, 2019) — STM, CTM, SAGE: variational EM, estimateEffect, searchK, FREX, spectral initialization, and the method of composition
  • sts (Chen & Mankad, 2024) — STS: the Structural Topic and Sentiment-Discourse model — the joint prevalence/sentiment Laplace E-step and the Poisson topic-word M-step, validated against the package
  • lda-c / ctm-c / dtm and hdp (Blei lab, 2006–2007) — CTM, DTM, HDP: the CTM, Dynamic Topic Model, and HDP samplers
  • gensim (Řehůřek & Sojka, 2010) — DTM, ensemble: the coherence-pipeline conventions (the coherence_type= API and default sliding windows; the measures themselves are Röder et al. 2015 and Mimno et al. 2011), the LdaSeqModel DTM reference, and the EnsembleLda (CBDBSCAN stable-topic) method ported for ensemble(method="stable")
  • tomotopy (bab2min, 2020) — API conventions (summary, the short-text models), and GDMR (generalized DMR; Lee & Song, 2020), validated against its GDMRModel
  • scikit-learn (Pedregosa et al., 2011) — NMF: the multiplicative-update solver (Lee & Seung, 2001) and the NNDSVD initialization (Boutsidis & Gallopoulos, 2008), validated against sklearn.decomposition.NMF (BSD-3-Clause); and LSA: latent semantic analysis / indexing (Deerwester et al., 1990), validated against sklearn.decomposition.TruncatedSVD (BSD-3-Clause) including its svd_flip sign convention. The numerics are reimplemented in Rust; the randomized truncated SVD shared by both (it seeds NMF's NNDSVD and is the LSA factorization itself) follows Halko et al. (2011).
  • keyATM (Eshima, Imai & Sasaki, 2024) — KeyATM: the base, covariate, and dynamic models, the information-theory token weighting, and the Chib (1998) change-point HMM, validated against the package
  • seededlda (Watanabe, 2023) — SeededLDA: the seeded-prior scheme
  • LightLDA (Yuan et al., 2015) — LDA: the alias-table Metropolis-Hastings sampler
  • GSDMM (Yin & Wang, 2014) — GSDMM: the movie-group-process mixture for short text
  • ProdLDA / AVITM (Srivastava & Sutton, 2017) — ProdLDA: autoencoding variational inference and the product-of-experts word model
  • BERTopic (Grootendorst, 2022) and Top2Vec (Angelov, 2020) — BERTopic, Top2Vec: the embedding-clustering pipeline, class-based TF-IDF, and the reduce → cluster → represent design
  • ETM (Dieng, Ruiz & Blei, 2020) — ETM: the Embedded Topic Model (per-document variational EM and an amortized VAE)
  • DETM (Dieng, Ruiz & Blei, 2019) — DETM: the Dynamic Embedded Topic Model (structured amortized variational inference with a hand-coded LSTM)
  • FASTopic (Wu et al., 2024) — FASTopic: the optimal-transport topic model
  • contextualized-topic-models (Bianchi et al., MIT) — CombinedTM (Bianchi, Terragni & Hovy, 2021) and ZeroShotTM (Bianchi, Nozza & Hovy, 2021): ProdLDA encoders that read a contextual document embedding, alongside or in place of the bag of words
  • CLNTM (Nguyen & Luu, 2021) — the InfoNCE contrastive regularization on topic vectors offered by the contrastive= flag on the VAE models
  • WHAI / Weibull-Dirichlet VAE (Zhang et al., 2018; Burkhardt & Kramer, 2019) — the Weibull-reparameterized Dirichlet prior offered by prior="dirichlet" on the VAE models
  • Neural variational topic models with alternative priors (Miao, Grefenstette & Blunsom, 2017; Nalisnick & Smyth, 2017) — the Gaussian stick-breaking prior offered by prior="stick_breaking" on the VAE models
  • TopicGPT (Pham et al., NAACL 2024, MIT) — TopicGPT: the generate / refine / assign prompt flow for LLM-driven topic discovery

The embedding-native models build on two pure-Rust crates: petal-clustering for HDBSCAN and umap-rs for the optional UMAP reducer, both BLAS-free.

Full citations for every model and reference implementation, and how to cite topica, are on the Citing page.

License

Apache-2.0 — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

topica-0.25.0.tar.gz (5.3 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

topica-0.25.0-cp39-abi3-win_amd64.whl (3.2 MB view details)

Uploaded CPython 3.9+Windows x86-64

topica-0.25.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (3.2 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ x86-64

topica-0.25.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (3.0 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ ARM64

topica-0.25.0-cp39-abi3-macosx_11_0_arm64.whl (2.9 MB view details)

Uploaded CPython 3.9+macOS 11.0+ ARM64

topica-0.25.0-cp39-abi3-macosx_10_12_x86_64.whl (3.1 MB view details)

Uploaded CPython 3.9+macOS 10.12+ x86-64

File details

Details for the file topica-0.25.0.tar.gz.

File metadata

  • Download URL: topica-0.25.0.tar.gz
  • Upload date:
  • Size: 5.3 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for topica-0.25.0.tar.gz
Algorithm Hash digest
SHA256 6fd4814277cb5c4d68f735786794bd6c956712598f03351f0478b4ca6786dd16
MD5 c9baad3122a09d412bd5d0ce1ffaa3a2
BLAKE2b-256 766616ee09d3b215fed86e297b422d70d496c0fad868c01201d4795c316cb4e9

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.25.0.tar.gz:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.25.0-cp39-abi3-win_amd64.whl.

File metadata

  • Download URL: topica-0.25.0-cp39-abi3-win_amd64.whl
  • Upload date:
  • Size: 3.2 MB
  • Tags: CPython 3.9+, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for topica-0.25.0-cp39-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 a2a158a02bc496b94b12195f9cb48d8be1be864768319f397306b31bf5393006
MD5 54fb7ee2681621e124fb3bbb25826303
BLAKE2b-256 435c6fd90d078961516fc60d4604e4c93a8251064ef984c745950d36b9babc9c

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.25.0-cp39-abi3-win_amd64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.25.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for topica-0.25.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 c4ee0f07576a00233e994ba6759402203621089cb3b3792517328cef7f464458
MD5 a4f9c15daf45a064ad5364e0cc12526a
BLAKE2b-256 92b931b7f8af08adf0580637f0ba70de98fc34542e4e90e1f83df8bc33f06b92

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.25.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.25.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for topica-0.25.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 fe566cda6abce39f641e9eb7d331970d652cee8bdcb39c4b26a5f2b83eb6bb2c
MD5 75a5d7c8a01f9b2d1a577e4eca169bc0
BLAKE2b-256 1458db861344071218a0f19907d11df63fdfa76e24bf9ca3d7d1ea8c06e6c594

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.25.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.25.0-cp39-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for topica-0.25.0-cp39-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 dcb46deb9eb1cac0f3c1a4193e8f22ef91c8713e535eef43301622cf3d11d828
MD5 acbb3296c39a6f83398363d533f25daf
BLAKE2b-256 cc27f4ef0f94a604da27fc9494bef4351bae968afb5f28bfd2c8a2711e5717e8

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.25.0-cp39-abi3-macosx_11_0_arm64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.25.0-cp39-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for topica-0.25.0-cp39-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 f35a4a0b66d8fe5fb78bab6ab44e1a5d0efb97e0487d43db1144734c0cacd279
MD5 e4e4845b38eefbc64264abc266cb512d
BLAKE2b-256 68be941996c7c54f06de39e6a5ab1a0f67180b1064726ef6ce4ea0825d0d3e1a

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.25.0-cp39-abi3-macosx_10_12_x86_64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page