Skip to main content

Python implementation of the Structural Topic Model (STM), a port of the R stm package with a scikit-learn style API

Project description

pystm — Structural Topic Model in Python

PyPI License: MIT

日本語 README

A Python port of the R stm package (Roberts, Stewart & Tingley) with a scikit-learn style API.

What is STM?

The Structural Topic Model (STM) is a logistic-normal topic model where document metadata ("prevalence covariates") shifts the prior mean of each document's topic proportions. Without covariates it reduces to the Correlated Topic Model (CTM). Estimation uses semi-collapsed variational EM, the same algorithm as the R stm() function.

Installation

pip install structural-topic-model

Quick Start

import numpy as np
from pystm import StructuralTopicModel

# X: (n_docs, n_vocab) word count matrix (dense or scipy.sparse)
# covar: (n_docs, n_covariates) prevalence covariate matrix (intercept added automatically)

model = StructuralTopicModel(n_components=10, init="spectral")
model.fit(X, prevalence=covar)

model.theta_        # topic proportions of training docs (n_docs, K)
model.components_   # topic-word distributions (K, V), each row sums to 1
model.gamma_        # prevalence regression coefficients (1+P, K-1)
model.sigma_        # topic covariance matrix (K-1, K-1)

# Inference on new documents (cf. fitNewDocuments)
theta_new = model.transform(X_new, prevalence=covar_new)

# Top words per topic (cf. labelTopics)
model.top_words(n_words=10)               # by probability
model.top_words(n_words=10, kind="frex")  # FREX: frequency-exclusivity balance

Without covariates, the model is estimated as a CTM:

model = StructuralTopicModel(n_components=10).fit(X)

Content Covariates

Content covariates allow the vocabulary used within topics to vary by document category. Pass one categorical label per document:

model = StructuralTopicModel(n_components=10)
model.fit(X, prevalence=covar, content=party_labels)

model.aspect_components_   # per-level topic-word distributions (n_levels, K, V)
model.kappa_["params"]     # sparse deviations from baseline (lasso-estimated)
model.content_levels_      # sorted unique levels

model.transform(X_new, prevalence=covar_new, content=labels_new)

Estimated via Distributed Poisson regression (equivalent to the R package's default kappa.prior="L1").

Covariate Effects (estimateEffect)

Regress topic proportions on covariates using method of composition, returning coefficients with measurement uncertainty:

from pystm import estimate_effect

eff = estimate_effect(model, covar, uncertainty="Global", nsims=25)
tables = eff.summary()      # {topic: structured array with estimate/std_error/t_value/p_value}
tables[0]["estimate"]       # coefficients for topic 0 (first entry is intercept)

Topic Selection (searchK)

from pystm import search_k

res = search_k(X, K_values=[5, 10, 15], prevalence=covar,
               model_params={"max_iter": 100})
res["heldout"]   # heldout log-likelihood (document completion)
res["residual"]  # residual dispersion — Taddy (2012), closer to 1 is better
res["semcoh"]    # semantic coherence
res["exclus"]    # exclusivity

Diagnostics

from pystm import topic_corr, semantic_coherence, exclusivity, check_residuals

tc = topic_corr(model, cutoff=0.01)    # topic correlation graph (simple method)
tc.posadj                               # positive correlation adjacency matrix
semantic_coherence(model, X, M=10)     # semantic coherence per topic
exclusivity(model, M=10)               # exclusivity per topic
check_residuals(model, X)              # residual dispersion test {dispersion, pvalue, df}

R–Python API Reference

R Python
stm(docs, vocab, K, prevalence=~x, data=meta) StructuralTopicModel(n_components=K).fit(X, prevalence=design)
init.type="Spectral" (recommended, default) init="spectral" (default)
init.type="Random" init="random"
gamma.prior="Pooled" (default) implemented (automatic when covariates are present)
sigma.prior sigma_prior
emtol / max.em.its tol / max_iter
model= (warm restart) warm_start=True
content=~group (kappa.prior="L1", default) fit(X, content=labels)
interactions content_interactions
fitNewDocuments() transform()
labelTopics() top_words()
estimateEffect() / summary() estimate_effect() / .summary()
searchK() search_k()
make.heldout() / eval.heldout() make_heldout() / eval_heldout()
topicCorr(method="simple") topic_corr()
semanticCoherence() / exclusivity() / checkResiduals() semantic_coherence() / exclusivity() / check_residuals()
$theta / $beta / $sigma / $mu$gamma theta_ / components_ / sigma_ / gamma_
$beta$logbeta (content model) aspect_components_ (probability scale)
$beta$kappa kappa_

Differences from scikit-learn LDA

  • components_ contains normalized probability distributions (sklearn LDA stores unnormalized pseudo-counts).
  • Covariates are passed via fit(X, prevalence=...) / transform(X, prevalence=...). R formulas are not supported — encode categorical variables numerically beforehand (e.g. pandas.get_dummies or patsy).
  • perplexity(X) is provided, computed from the variational lower bound (exp(-bound / n_tokens)); lower is better.
  • warm_start=True enables continued learning across repeated fit() calls (cf. R's model= argument).

Not Implemented

  • gamma.prior="L1" (prevalence-side glmnet mode)
  • kappa.prior="Jeffreys" (legacy content estimation)
  • fixedintercept=FALSE
  • LDA (collapsed Gibbs) initialization, ngroups memoized inference, K=0 (Lee & Mimno)
  • estimateEffect() with uncertainty="Local", formula interface (pass pre-expanded basis matrices instead)
  • topicCorr(method="huge"), selectModel(), permutationTest(), plot functions

Development

uv sync
uv run pytest tests/

References

  • Roberts, M., Stewart, B., & Tingley, D. (2019). stm: An R Package for Structural Topic Models. Journal of Statistical Software, 91(2).
  • Arora, S. et al. (2013). A Practical Algorithm for Topic Modeling with Provable Guarantees. ICML.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

structural_topic_model-0.2.2.tar.gz (35.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

structural_topic_model-0.2.2-py3-none-any.whl (28.8 kB view details)

Uploaded Python 3

File details

Details for the file structural_topic_model-0.2.2.tar.gz.

File metadata

  • Download URL: structural_topic_model-0.2.2.tar.gz
  • Upload date:
  • Size: 35.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.10.12 {"installer":{"name":"uv","version":"0.10.12","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for structural_topic_model-0.2.2.tar.gz
Algorithm Hash digest
SHA256 a1dd94625293ef73436b262118d1baf4f3cd8b50e101ac5e5b5e3a7c6a16a8f3
MD5 fdcf57feeb68f24cc9dba18a9019b6fc
BLAKE2b-256 15cdd6793c3cb055e50718c3ec3021dca170ed1e2c1a550cae95abad682f7992

See more details on using hashes here.

File details

Details for the file structural_topic_model-0.2.2-py3-none-any.whl.

File metadata

  • Download URL: structural_topic_model-0.2.2-py3-none-any.whl
  • Upload date:
  • Size: 28.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.10.12 {"installer":{"name":"uv","version":"0.10.12","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for structural_topic_model-0.2.2-py3-none-any.whl
Algorithm Hash digest
SHA256 99c2ca30c002a180d7e6655b3daab5516e3799260c66981a05ae9f15cac21136
MD5 e84ae7f7863bf59fb43c97d3714c3b8e
BLAKE2b-256 a1f60b7d098bdb6416a36f987a0545d0844a0f35f03d61a0c5c538d50ce35381

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page