Skip to main content

saphes logo

saphes

Readability and lexical diversity — two metrics, done carefully, with the parameters other implementations hardcode.

saphes — σαφής, "clear, plain, distinct". Aristotle makes clarity the chief virtue of λέξις (style); the other classical axis is ποικιλία, variety. The two metrics here are exactly those axes: LIX measures clarity, TTR measures variety.

Why this exists

textstat, textdescriptives, lexicalrichness and taaled already cover this ground. Two reasons to still build it:

  1. The LIX long-word threshold is hardcoded at 6 everywhere. That 6 comes from Björnsson's Swedish original, and it is wrong for the languages we work on. Hungarian is agglutinative and Ancient Greek heavily inflected, so at threshold 6 nearly every token counts as "long" and the index saturates into a flat line. Measured over the full Hungarian Webcorpus, 44.5% of running tokens are "long" at threshold 6, against 25.7% in Swedish. On real Hungarian prose that pushes LIX to 60.4 — "very difficult" — where the calibrated threshold gives 43.4. Parameterising the threshold is the whole point.
  2. Implementations disagree. They count words, sentences and long words differently, so they rank the same texts differently. So: expose the counts, document every choice, make results auditable.

Non-goal: becoming another kitchen-sink readability library.

Installation

uv add saphes

The core has no dependencies — plain Python and the standard library. Two optional extras stay out of the core path: saphes[snowball] for Hungarian stemming and saphes[punkt] for the NLTK sentence splitter.

Quickstart

from saphes import lix, lexical_diversity

# LIX takes SURFACE FORMS — word length is the signal.
result = lix("The cat sat on it. Complicated sentences generally frighten us.")
result.score                                       # 45.0
result.words, result.sentences, result.long_words  # (10, 2, 4)  -> A, B, C
result.band                                        # 'standard'

# Diversity takes LEMMAS, and `unit` is required — there is no default.
lexical_diversity(lemmas, unit="lemma")

# Comparing texts of different lengths? Use MATTR, not TTR.
lexical_diversity(lemmas, unit="lemma", window=100).mattr

Hungarian needs three things English does not — a retuned threshold, a letter count, and a stand-in for a lemmatiser:

from saphes import hungarian_letter_count, hungarian_stems, lix, recommended_threshold

# `sz` is one letter, not two characters. The threshold is calibrated per policy.
hu = recommended_threshold("hu-letters")          # 8, not Björnsson's Swedish 6
lix(text, length_policy=hungarian_letter_count, long_word_threshold=int(hu))

# No lemmatiser? Snowball is the fallback, declared as its own unit.
lexical_diversity(hungarian_stems(tokens), unit="stem")

The data contract

The two metrics require opposite token streams.

Metric Wants Because
lexical_diversity lemmas Surface variation is noise — it measures morphology, not vocabulary. Hungarian ház / házak / házban / házakat is four types and one lemma.
lix surface forms Word length is the signal. házakban is 8 characters; its lemma ház is 3.

Feed the same list to both and exactly one is silently wrong — no error, no NaN, just a plausible number. unit is required, the parameter names differ, a raw string is refused where it could only be wrong, and every result records what it measured.

saphes consumes lemmas; it does not produce them. Lemmatisation is language-specific and heavy — CLTK or a treebank for Greek, huspacy for Hungarian. The caller lemmatises; saphes measures.

The one exception is optional and honest about itself: saphes[snowball] gives Hungarian Snowball stemming for callers with no lemmatiser. Its output is declared as unit="stem", a third stream with its own name, because a stemmer both over- and under-merges and a stem-based number is comparable only to another from the same stemmer.

Documentation

saphes.readthedocs.io, organised by Diátaxis:

  • Tutorial — new here? Measure your first text in about ten minutes, with nothing to download.
  • How-to guides — measure Hungarian text, supply a sentence count, count letters rather than characters, stem without a lemmatiser, calibrate a threshold.
  • Reference — every function, its contract and its failure modes, plus the calibration data.
  • Explanation — why the two metrics need opposite input, why the threshold has to move, and why implementations disagree.

Every code block in the docs is executed by CI, so nothing there can drift.

Roadmap

  • LIX with a parameterised long-word threshold
  • TTR and MATTR with a required, recorded token unit
  • RIX (long words per sentence)
  • Empirically calibrated per-language thresholds, from token-weighted word-length distributions — Hungarian ships as recommended_threshold("hu")
  • Phonotactically aware Hungarian letter counting — sz is one letter, ssz is two, and morpheme boundaries are handled by rule plus an attested table
  • Optional Snowball stemming for callers with no lemmatiser, as a declared third unit
  • Mean dependency distance and mean hierarchical distance, from any parser's output
  • Zero-dependency adapters for HuSpaCy/spaCy and CoNLL-U/emtsv
  • Loan-word ratio (idegenszó-arány) against a caller-supplied lexicon
  • A public-domain Hungarian loan-word lexicon to ship with it
  • The same study for Ancient Greek, for the Homer project
  • POS-filtered diversity, once lemmas carry tags
  • MTLD, HD-D, vocd-D, Maas

Research

  • Diachronic analysis: test MDD and loan word density changes on the ParlaMonitor corpus across time
  • Data diversity benchmarking: benchmark MDD and idegenszó-arány across diverse Hungarian registers (Webcorpus, legal texts, literary prose, social media)
  • Multilingual scaling: test the agnostic MDD calculation engine on Universal Dependencies corpora across multiple languages
  • Dependency motifs: chain motif identification for advanced cognitive load mapping, extending the MDD/MHD pair (Jing & Liu 2017)

Maintenance

  • Logo and README banner
  • Mutation-testing baseline

Explicitly out of scope: Flesch, Kincaid, SMOG and relatives. They need syllabification, which is language-specific and a different project. LIX was chosen precisely because it needs only word length and sentence count, so it travels across languages.

Made by

saphes is made by Crow Intelligence.

License

MIT

Release files for saphes 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for saphes 0.2.0
File Size Uploaded
saphes-0.2.0.tar.gz 66.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for saphes 0.2.0
File Interpreter ABI Platform
saphes-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 142.3 kB

Release files / saphes-0.2.0.tar.gz

Download URL saphes-0.2.0.tar.gz
Size 66.0 kB
Tags Source
SHA-256 checksum
How to use checksums
135c1eb3ffd301b00973a76da4a0c8fdd0385b5f992a067ae14202fd94512127
BLAKE2b-256 checksum
How to use checksums
96831fc6a93636050067f55097d936f6952d15772bb33553f6a7b7d99d887bed
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 8, 2026.

Transparency log

Release files / saphes-0.2.0-py3-none-any.whl

Download URL saphes-0.2.0-py3-none-any.whl
Size 76.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ec2607211d84e28c7a2afac786c14f4ebc4b9183f537f0c287153c0c049f9392
BLAKE2b-256 checksum
How to use checksums
02c670203b318272462d2c1ab2c6e3f0f58fd6d3c5bc2ed3ad924836513e5e81
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 8, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.1

2 release files

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page