Skip to main content

Topica: fast, all-purpose topic modeling for Python — a Rust core for LDA, STM, and more

Project description

Topica: fast, all-purpose topic modeling for Python

PyPI CI Docs License: Apache-2.0

topica is a fast topic-modeling library for Python with more than a dozen models, built for social scientists who want to move from text data to publishable results in a single workflow. It brings together models and tools usually split across JVM software like MALLET and R packages like stm, and runs them on a parallel Rust core competitive with the standard implementations, with reproducible fits: the variational models are identical to the bit, and the samplers reproduce from a fixed seed and thread count. Each model comes with the validation, covariate-effect, and reporting tools to meet the standards reviewers expect.

pip install topica

The core needs only NumPy. Optional extras add features without weighing the core down: topica[viz] (matplotlib plots), topica[formula] (R-style formulas), topica[polars] (Polars frames), and topica[llm] (LLM labels and embeddings, OpenAI or local via ollama). PyTorch is never required.

from topica import LDA

model = LDA(num_topics=2, seed=42)
model.fit([["cat", "dog", "fish"]] * 15 + [["planet", "star", "moon"]] * 15, iters=1000)

for i, words in enumerate(model.top_words(3)):
    print(f"Topic {i}:", " ".join(w for w, _ in words))

See the getting-started guide and the worked examples for end-to-end analyses.

Models

Count-based models learn topics from word counts (collapsed Gibbs, variational EM, or an amortized VAE):

Model What it's for
LDA Classic topics via fast collapsed-Gibbs (SparseLDA); optional multi-threaded and LightLDA alias samplers
STM The Structural Topic Model: correlated topics with prevalence and content covariates
STS Structural Topic and Sentiment-Discourse: covariate-driven topic sentiment/tone on top of STM
CTM Correlated topics (logistic-normal)
DMR Topics conditioned on document metadata (Dirichlet-multinomial regression)
GDMR Generalized DMR over continuous metadata via a Legendre basis, with topic distribution functions across the metadata range
DTM Dynamic topics that evolve across time slices
HDP Nonparametric LDA that learns the number of topics from the data; by default fits with fixed concentrations (topic count steered by gamma), with optional concentration resampling
keyATM / seededlda Guided topics steered by seed words
ProdLDA Sharper, more coherent topics via a product-of-experts word model, fit as an amortized VAE (no PyTorch)
PT / GSDMM Short-text models for tweets, survey answers, headlines
SupervisedLDA Topics shaped to predict a per-document response
LabeledLDA Supervised topics tied to document labels
SAGE Content-covariate topics: the same topic worded differently across groups
PA / HLDA Topic hierarchies (Pachinko, nested-CRP)

Embedding-based models start from document embeddings you supply (no PyTorch, no UMAP/numba in the wheel):

Model What it's for
BERTopic Cluster document embeddings, label topics by class-TF-IDF; topic reduction and a soft per-document distribution
Top2Vec Topics as points in the embedding space; topic words are the nearest word vectors
ETM Embedded Topic Model: a logistic-normal topic model with the topic-word distribution factored through embeddings (β = softmax(ρ·α)); per-document EM or an amortized VAE (inference="vae")
FASTopic Topics read off two optimal-transport plans between document, topic, and word embeddings

Every model exposes the same shape: fit(docs, …), then topic_word (φ), doc_topic (θ), top_words(n), and save/load. The count-based variational models (CTM/STM/STS/SupervisedLDA/DTM) parallelize across cores while staying bit-for-bit deterministic. The embedding models split into two kinds: BERTopic and Top2Vec run the reduce → cluster → represent pipeline, while ETM and FASTopic are generative and mixed-membership; all of them take vectors from any embedder (sentence-transformers, an API, a local model such as ollama). Full guides: the models and embedding topics.

Diagnostics & analysis

Model-agnostic: they work on any fitted model's topic_word/doc_topic:

  • Quality: coherence (u_mass, c_v, c_uci, c_npmi; co-occurrence counting in the Rust core), exclusivity, topic_diversity, quality_frontier
  • Labeling: label_topics (prob / FREX / lift / score), frex, relevance, find_thoughts, topic_table, summary
  • Validation: word_intrusion, document_intrusion, bootstrap_stability, search_k
  • Reliability: select_model (fit many seeds) and ensemble (combine runs into a consensus more reliable than any single fit — cluster/align/stable methods, the last a gensim EnsembleLda port)
  • Comparison: fighting_words (weighted log-odds) for contrasting corpora
  • Covariate effects: estimate_effect (method of composition, cluster-robust SEs, GLM links), topic_correlation, and the design helpers one_hot, spline, and interaction (all top level; they build covariate bases for any model's design matrix); posterior_theta_samples draws θ for the logistic-normal models (STM/CTM)
  • Preprocessing: tokenize, learn_phrases / apply_phrases, split_documents, the Corpus class

See diagnostics and covariate effects.

Performance

topica runs on a parallel Rust core. It is several times faster than R stm — the single-threaded field standard — for the structural and other variational models, and it matches the hand-tuned compiled samplers core for core: parity with Java MALLET on plain LDA and with the C++ keyATM on keyword models. On the political-blog corpus (2,000 documents, fit time only, same iterations on both sides):

Model Reference topica speedup
STM R stm 3–6× single-threaded, ~10–22× multithreaded
LDA Java MALLET parity single-threaded; multithread speedup grows with corpus size
keyATM R keyATM parity single-threaded, ~2× multithreaded

For the approximate parallel Gibbs samplers the multithreaded speedup grows with corpus size: the per-sweep count-table merge is fixed overhead, so larger corpora amortize it over more sampling work. LDA's eight-core speedup over MALLET runs about 3× at 2,000 documents and reaches ~4× at 5,000, so the small-corpus figures above understate what large-corpus users see.

Every fit is reproducible from a fixed seed and validated against its reference. See Benchmarks for the full methodology; reproduce the 2,000-document table with python benchmarks/speed_vs_r.py and the size-varying curve with python benchmarks/speed_vs_size.py.

Install from source

pip install maturin
git clone https://github.com/nealcaren/topica && cd topica
python -m venv .venv && source .venv/bin/activate
maturin develop --release --features python

Requires numpy >= 1.21. Use --release (the debug build is much slower).

Acknowledgements

Topica stands on a generation of open topic-modeling research and code. Each entry below lists the reference, its authors and year, and the topica class(es) it underlies; the other models are Rust ports or reimplementations, validated against these reference implementations.

  • MALLET (McCallum, 2002) — LDA, DMR, LabeledLDA: the SparseLDA sampler, Dirichlet-multinomial regression, and hyperparameter optimization. LDA began as a port of David Mimno's RustMallet (Apache-2.0) and follows its SparseLDA sampler and fixed-point optimizer closely, but uses its own RNG (PCG), so it is not byte-identical to RustMallet. Against Java MALLET (also a different RNG) it recovers the same topics on a planted corpus (cosine 1.000)
  • stm (Roberts, Stewart & Tingley, 2019) — STM, CTM, SAGE: variational EM, estimateEffect, searchK, FREX, spectral initialization, and the method of composition
  • sts (Chen & Mankad, 2024) — STS: the Structural Topic and Sentiment-Discourse model — the joint prevalence/sentiment Laplace E-step and the Poisson topic-word M-step, validated against the package
  • lda-c / ctm-c / dtm and hdp (Blei lab, 2006–2007) — CTM, DTM, HDP: the CTM, Dynamic Topic Model, and HDP samplers
  • gensim (Řehůřek & Sojka, 2010) — DTM, ensemble: the coherence-pipeline conventions (the coherence_type= API and default sliding windows; the measures themselves are Röder et al. 2015 and Mimno et al. 2011), the LdaSeqModel DTM reference, and the EnsembleLda (CBDBSCAN stable-topic) method ported for ensemble(method="stable")
  • tomotopy (bab2min, 2020) — API conventions (summary, the short-text models), and GDMR (generalized DMR; Lee & Song, 2020), validated against its GDMRModel
  • keyATM (Eshima, Imai & Sasaki, 2024) — KeyATM: the base, covariate, and dynamic models, the information-theory token weighting, and the Chib (1998) change-point HMM, validated against the package
  • seededlda (Watanabe, 2023) — SeededLDA: the seeded-prior scheme
  • LightLDA (Yuan et al., 2015) — LDA: the alias-table Metropolis-Hastings sampler
  • GSDMM (Yin & Wang, 2014) — GSDMM: the movie-group-process mixture for short text
  • ProdLDA / AVITM (Srivastava & Sutton, 2017) — ProdLDA: autoencoding variational inference and the product-of-experts word model
  • BERTopic (Grootendorst, 2022) and Top2Vec (Angelov, 2020) — BERTopic, Top2Vec: the embedding-clustering pipeline, class-based TF-IDF, and the reduce → cluster → represent design
  • ETM (Dieng, Ruiz & Blei, 2020) — ETM: the Embedded Topic Model (per-document variational EM and an amortized VAE)
  • FASTopic (Wu et al., 2024) — FASTopic: the optimal-transport topic model

The embedding-native models build on two pure-Rust crates: petal-clustering for HDBSCAN and umap-rs for the optional UMAP reducer, both BLAS-free.

Full citations for every model and reference implementation, and how to cite topica, are on the Citing page.

License

Apache-2.0 — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

topica-0.19.0.tar.gz (3.1 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

topica-0.19.0-cp39-abi3-win_amd64.whl (2.6 MB view details)

Uploaded CPython 3.9+Windows x86-64

topica-0.19.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.7 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ x86-64

topica-0.19.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (2.5 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ ARM64

topica-0.19.0-cp39-abi3-macosx_11_0_arm64.whl (2.4 MB view details)

Uploaded CPython 3.9+macOS 11.0+ ARM64

topica-0.19.0-cp39-abi3-macosx_10_12_x86_64.whl (2.6 MB view details)

Uploaded CPython 3.9+macOS 10.12+ x86-64

File details

Details for the file topica-0.19.0.tar.gz.

File metadata

  • Download URL: topica-0.19.0.tar.gz
  • Upload date:
  • Size: 3.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for topica-0.19.0.tar.gz
Algorithm Hash digest
SHA256 04767574637b791a2819dbb84b684e38359cd137d40f9be93e41957c1323d880
MD5 5855c9b5312e646d2a9c1ba8192cd525
BLAKE2b-256 6523880122c1643c822ca07699ea6a920e3f76346c1a4e7aab5bff143b162db2

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.19.0.tar.gz:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.19.0-cp39-abi3-win_amd64.whl.

File metadata

  • Download URL: topica-0.19.0-cp39-abi3-win_amd64.whl
  • Upload date:
  • Size: 2.6 MB
  • Tags: CPython 3.9+, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for topica-0.19.0-cp39-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 d3fe8e3f0536b4dec6a73b60c807f780d9e68b9cc4911d43bb22f9180c8fa6d4
MD5 15282f0bdb0a508ad1d05fe6e53114a6
BLAKE2b-256 bdb9abbe6945293119d09b0d300505820e4ad21045c470706d5fe7fb7a3081f6

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.19.0-cp39-abi3-win_amd64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.19.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for topica-0.19.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 175f977cc8f34bef314bfd85c1ebec384223a2b5e0bd9bf38659aefdd292fc07
MD5 ea8975ce4786af7c976a286caeb9523a
BLAKE2b-256 3eb9e907187fcc2b8237ebd2cac4f9656054b9631c3404ae1df642e19a7af46f

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.19.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.19.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for topica-0.19.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 e429079130242b29cf3e8652f3f238e819419078f7e8755330290ea83175d9bf
MD5 46022f63f9132595e5e47ba45095ffb3
BLAKE2b-256 598f5de14713ed24182e714cca114fe31d8bd40718373dd1257437f004d8bbcd

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.19.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.19.0-cp39-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for topica-0.19.0-cp39-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 4238d73f5fc94dc1b35586abb0eb6a36e18a4d04166b0af174de25f5a0190a6a
MD5 3daee5c7b7b8916f3bb10ef708a3114e
BLAKE2b-256 34a3dce376c4aeb947efe53ef5508e04bf99397516f6bd015ea19c048c9c3ceb

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.19.0-cp39-abi3-macosx_11_0_arm64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file topica-0.19.0-cp39-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for topica-0.19.0-cp39-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 16daa86077d8618e5eceb68ca7a089bcce63c4a7afd4fe57f47c7918c41fb460
MD5 93c6ad7dce770dcba69240561d3325e3
BLAKE2b-256 7b0a87824a137db2502c6df41488daffd0685a93aca18ca93f3fcd34657e7bbc

See more details on using hashes here.

Provenance

The following attestation bundles were made for topica-0.19.0-cp39-abi3-macosx_10_12_x86_64.whl:

Publisher: CI.yml on nealcaren/topica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page