textmeasures
textmeasures is a Python package for quantitative text measurement in corpus
linguistics, quantitative linguistics, discourse analysis, syntactic
complexity, lexical statistics, collocation analysis, linguistic networks, and
embedding-based semantic analysis.
Version 0.2.0 expands the package from frequency-distribution measures into a
broader toolkit. Many APIs accept ordinary Python sequences; CoNLL-U based
modules use asymgram for dependency
treebank parsing.
Researchers are encouraged to use Codex and other AI agents to read the source code, inspect the formulas and references, and then use the relevant metrics directly in their own analysis workflows.
Installation
pip install textmeasures
For local development:
git clone https://github.com/Yihtsy/textmeasures.git
cd textmeasures
pip install -e .
Six Measure Families
textmeasures currently covers six broad families of indicators.
| Family | What it covers | Main modules / APIs |
|---|---|---|
| 1. Frequency distribution and vocabulary structure | Entropy, repetition, richness, evenness, concentration, inequality, rank-frequency curves, Zipf-family fitting, QUITA-style indicators, and frequency-spectrum measures. | freqdist.py, FreqDist, FreqSpectrum, entropy, repeat_rate, gini, zipf_fitted_parameters, thematic_concentration |
| 2. Length and basic textual quantity | Text length and length-vector helpers for raw text and CoNLL-U style inputs. | length.py, text_length |
| 3. Symbolic, lexical, POS, and numeric sequences | Symbol sequences, word/lemma/POS/dependency-label sequences, lexical richness, MTLD/HD-D/vocd-style measures, dispersion, UPOS frequency profiles, numeric sequence dynamics, and motif sequences. | sequence.py, SymbSeq, WordSeq, LemmaSeq, PosSeq, DeprelSeq, NumSeq, Motif, lexical_richness, symbol_dispersion |
| 4. Syntactic complexity and dependency-tree structure | CoNLL-U dependency-distance measures, UD dependency relation frequencies, MDD/NDD/PMDD, MHD/MHDD, dependency direction, Table-3 style syntactic features, QuanSyn-style measures, linguistic feature counts/rates, dependency and adjacency edges. | syn_complex.py, synstruc.py, mean_dep_dist, dependency_tree_measures, synstruc_features, quansyn_features, linguistic_feature_counts |
| 5. Contingency tables and lexical association | General n x m contingency tables, special 2 x 2 tables, chi-square, likelihood-ratio/G test, Cramer's V, Tschuprow's T, uncertainty coefficient, PMI variants, collocation scores, odds/risk ratios, similarity coefficients, Fisher/binomial/Poisson tests, kappa, and tetrachoric correlation. |
contingency.py, ContingencyTable, TetrachoricTable |
| 6. Linguistic networks and embedding-based semantic measures | NetworkX-backed linguistic networks from edges or CoNLL-U, dependency/adjacency/co-occurrence networks, degree/strength/path/clustering/centrality distributions, vulnerability, semantic cosine similarity, MeanK measures, centroid coherence, temporal dynamics, semantic graph metrics, and perplexity helper. | network.py, embed.py, LingNetwork, network_metrics, semantic_similarity_features |
Several functions return bundles of metrics as dictionaries. This design keeps the public API compact while preserving named, stable metric keys.
Quick Start: Frequency Distributions
from textmeasures import FreqDist, entropy, repeat_rate, gini, normalized_entropy
freqs = FreqDist([10, 5, 3, 1, 1])
print(freqs.to_list())
print(entropy(freqs))
print(repeat_rate(freqs))
print(gini(freqs))
print(normalized_entropy(freqs))
Quick Start: Symbol Sequences
from textmeasures import SymbSeq
seq = SymbSeq("the cat saw the dog and the cat")
print(seq.richness())
print(seq.dispersion("the", part_count=3))
print(seq.mtld())
Quick Start: CoNLL-U and Dependency Structure
from textmeasures import dependency_tree_measures, ud_deprel_frequencies
path = "sample.conllu"
tree = dependency_tree_measures(path)
rels = ud_deprel_frequencies(path, denominator="dependencies")
print(tree["mdd"])
print(tree["ndd"])
print(tree["mhd"])
print(rels["absolute"])
CoNLL-U functions usually skip multiword-token and empty-node rows by requiring
integer token IDs. Dependency counts and dependency-distance denominators
normally exclude root; relation-frequency helpers expose a denominator option
so users can divide by dependency count or word count.
Quick Start: Contingency and Association
from textmeasures import TetrachoricTable
table = TetrachoricTable([[10, 2], [3, 20]])
print(table.pmi())
print(table.odds_ratio(correction=0.5))
print(table.fisher_exact())
print(table.summary())
Quick Start: Linguistic Networks
from textmeasures import LingNetwork
net = LingNetwork.from_conllu(
"sample.conllu",
network_type="dependency",
field="lemma",
directed=True,
weighted=True,
)
print(net.node_count())
print(net.average_degree())
print(net.metrics())
Quick Start: Embedding-Based Semantic Measures
import numpy as np
from textmeasures import semantic_similarity_features
embeddings = np.random.default_rng(42).normal(size=(8, 384))
features = semantic_similarity_features(embeddings)
print(features["mean_k1"])
print(features["global"])
print(features["graph_density"])
textmeasures does not compute embeddings itself. Pass an n x d matrix from
your preferred embedding model, where rows are meaningful units such as words,
clauses, sentences, turns, or utterances.
API Overview
Distribution Objects
FreqDistRelFreqDistCumFreqDistCumRelFreqDistFreqSpectrum
Frequency and Vocabulary Measures
Examples include:
entropy,renyi_entropy,tsallis_entropy,normalized_entropyrepeat_rate,inverse_repeat_rate,inverse_simpson,simpsonrichness,hill_number,hill_evenness,hill_unevennessgini,theil_t,mean_log_deviation,generalized_entropycurve_length,lambda_indicator,b1,b2,b3,b4,b5,b6,b8,b10h_point,k_point,n_point,m_point,r1,r2,r4zipf_fitted_parameters,zipf_mandelbrot_fitted_parameters,zipf_alekseev_fitted_parameters
CoNLL-U Utilities
conllu_to_symbolsconllu_to_freqdistconllu_token_dictconllu_token_is_valid
Supported linguistic units include word, lemma, letter, character,
upos, deprel, and n_gram; availability depends on the selected language.
Syntactic and Dependency Measures
mean_dep_distnormalized_dep_distmean_dep_dist_per_senttotal_dep_distmax_dep_distparameterized_mean_dep_distconditional_mean_dep_distdependency_tree_measuressynstruc_featuresquansyn_featureslinguistic_feature_countslinguistic_feature_ratesud_deprel_frequenciesmean_dep_dist_by_deprel
Sequence Measures
SymbSeq.richness()SymbSeq.dispersion()SymbSeq.dispersion_table()PosSeq.pos_metrics()NumSeq.summary()Motif.motif_metrics()
Network and Semantic Measures
LingNetwork.from_edges()LingNetwork.from_conllu()LingNetwork.metrics()LingNetwork.micro_metrics()LingNetwork.macro_metrics()LingNetwork.distribution_metrics()semantic_similarity_featuressemantic_graph_metricstemporal_dynamics
Inputs and Return Values
- Frequency-distribution functions accept integer counts, probability vectors,
relative-frequency vectors, or
FreqDist-like objects. - CoNLL-U functions accept paths ending in
.conllu,asymgram.TokenList, orasymgram.SentenceListobjects, depending on the function. - Network functions use NetworkX graphs internally but expose a linguistic wrapper class.
- Embedding functions accept two-dimensional numeric arrays.
- Bundle functions return dictionaries with stable snake_case keys.
Citation
If you use textmeasures in academic work, please cite it as research software.
APA
Yih, T. (2026). textmeasures: Quantitative text measurement tools for frequency, sequence, syntax, contingency, network, and semantic measures (Version 0.2.0) [Computer software]. GitHub. https://github.com/Yihtsy/textmeasures
BibTeX
@software{yih_textmeasures_2026,
author = {Yih, Tsy},
title = {textmeasures: Quantitative text measurement tools for frequency, sequence, syntax, contingency, network, and semantic measures},
year = {2026},
version = {0.2.0},
url = {https://github.com/Yihtsy/textmeasures},
note = {Python package}
}
Development Notes
This package is still evolving. New metrics are added conservatively: formulas should be explicit, input requirements should be clear, and public functions should include references in their docstrings.
License
License information will be added here.
Metadata
Release files for textmeasures 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| textmeasures-0.2.0.tar.gz | 74.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| textmeasures-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 148.9 kB
Release files / textmeasures-0.2.0.tar.gz
| Download URL | textmeasures-0.2.0.tar.gz |
|---|---|
| Size | 74.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
58d1be58a2bbe93eb34cc71ec72b6815814d36d8c586a4c874db25c8f2014509
|
|
BLAKE2b-256 checksum How to use checksums |
e0781619b93201f320759e812eb1ac17785097c1c1f1a5ee91e3a8d55887f0d9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency logRelease files / textmeasures-0.2.0-py3-none-any.whl
| Download URL | textmeasures-0.2.0-py3-none-any.whl |
|---|---|
| Size | 74.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d0d71ad0b4efed017bec394aeb9206461f1e4800996e4da420d9e2ff87211726
|
|
BLAKE2b-256 checksum How to use checksums |
5625a6b3d69f31bf245d3d1df2a8f5cf54b214e1f9fea930e8429ae9f95de58e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency log