spokenform
spokenform converts plain written text in one selected language into reviewable
text intended for speech systems. It is a text-to-text frontend: written text in,
spoken form out.
The package provides:
- context-aware abbreviation and source-aligned numeric-unit expansion through
abbr2words; - a locale-aware structured-value stage for quantities, dates, times, currencies, temperatures, labels, and contextual ordinals;
- locale-policy-wrapped number, date, time, currency, decimal, and ordinal verbalization;
- optional provider-neutral spaCy annotations for POS-aware abbreviation rules;
- stage-level provenance through
PreparedText; - composed input-to-output offset maps with left/right boundary bias;
- caller-defined and automatically discovered protected ranges;
- conservative protection for URLs, email addresses, and arbitrary multi-dot semantic-version/ID sequences.
High-confidence URL, e-mail, semantic-version, and contextual Roman rendering
is available explicitly with normalize_literals=True; caller-protected spans
remain absolute. Structured fractions, identifiers, operator-shaped math,
music-context tokens, and controlled biological names use semantic precedence
and source-aligned mappings.
It intentionally does not detect languages, parse or render SSMD, segment mixed
languages, generate phonemes, or depend on kokorog2p.
Installation
python -m pip install spokenform
For optional spaCy integration:
python -m pip install "spokenform[spacy]"
python -m spacy download de_core_news_sm
spaCy and its trained pipelines are separate packages. spokenform never downloads
a model automatically.
Quickstart
from spokenform import prepare
prepared = prepare(
"Prof. Klein bringt am 14.05.2026 um 18:20 Uhr 2 kg mit.",
language="de",
)
print(prepared.spoken_text)
print(prepared.render_changes())
The result contains:
source_text: unchanged caller input;clean_text: plain text used by the normalization pipeline;spoken_text: readable normalized output;language: normalized processing-language code;- ordered stages and mapped edits;
- a composed
offset_map; - structured warnings.
PreparedText.text is an alias for spoken_text.
kokorog2p adapter
Use prepare_for_kokorog2p(text, language=...) for one explicitly selected
language run. The adapter preserves caller-owned run whitespace and protected
overrides, emits exact source-coordinate replacements, and leaves tokenization,
G2P, phonemization, and model punctuation to kokorog2p. German, French, Spanish,
Italian, Portuguese, Czech, and English are parity-gated semantic migration
targets. English is active on the kokorog2p spokenform adapter for reviewed
structured semantics, contextual single-dot release labels, and safe
ordinary-number categories; phoneme-sensitive years, suffix ordinals, Roman
numerals, phone/ID and arbitrary multi-dot sequences, numeric suffixes, and G2P
decisions remain downstream in kokorog2p.
German quantity symbols are recognized by abbr2words.iter_unit_matches().
spokenform owns only the canonical German grammar that realizes those matches,
including gender, invariant Stück, currency decomposition, and lexical decimal
digits. French likewise realizes canonical abbr2words quantity and currency
identities, including French dates, times, ordinals, decimal digits, plural
grammar, temperatures, and major/minor currency units. Spanish realizes
canonical quantities, temperatures, currencies, dates, and ordinary numbers;
Spanish 18:20-style time expressions remain caller-managed. Italian realizes
reviewed dates, quantities, temperatures, currencies, and ordinary numbers;
Italian colon times remain caller-managed. Portuguese realizes reviewed dates,
quantities, temperatures, currencies, and ordinary numbers; Portuguese colon
times remain caller-managed. Czech realizes reviewed dates, ordinary numbers,
quantities, temperatures, currencies, and canonical extended units; Czech colon
times remain caller-managed. No locale copies raw symbol inventories or
downstream tokenizer/phoneme rules.
Language boundary
Each call processes one language. Production callers should always pass
language=...; English remains the API default for compatibility and simple CLI
usage.
Language detection and mixed-language handling belong in the orchestration or G2P
layer. A foreign word may remain unchanged through normalization and be handled
afterward. Existing source spans can be transferred with prepared.offset_map.
Markup must also be parsed outside this package. Pass plain text to prepare() and
use ProtectedSpan for ranges generic normalization must not change.
Configuration
from spokenform import PreparationConfig, prepare
config = PreparationConfig(
language="en",
expand_abbreviations=True,
expand_structured=True,
expand_numbers=True,
normalize_whitespace=True,
context=True,
)
prepared = prepare("The board is 2 in. wide.", config=config)
When a PreparationConfig is supplied, it is authoritative for pipeline options.
spaCy support
spaCy supplies POS annotations for abbreviation rules that opt into POS guards. The public normalization API remains provider-neutral.
abbr2words accepts POS annotations, but its bundled registries do not necessarily
require POS labels. Therefore installing spaCy alone may not change default
normalization output. The integration is usable for custom POS-guarded entries.
Load and inject a pipeline in the application:
import spacy
from spokenform import prepare
nlp = spacy.load("en_core_web_sm")
prepared = prepare(
"The board is 2 in. wide.",
language="en",
nlp=nlp,
)
Or ask spokenform to load an already installed model:
prepared = prepare(
"The board is 2 in. wide.",
language="en",
spacy_model="en_core_web_sm",
strict=True,
)
Model names and paths are passed to spacy.load(). Loaded models are cached by
language/model key. reset_spacy_cache() clears that cache.
The adapter reads the token attributes text, idx, pos_, tag_, lemma_, and lang_. lang_ is carried as provider metadata; it is not used as language detection. Annotation spans are validated against the exact input text and remapped when protected ranges are replaced by internal sentinels.
A trained pipeline with POS or morphological annotations is required for quality
improvement; spacy.blank(...) supplies tokenization but normally no useful POS
tags.
Explicit annotations take precedence over nlp and spacy_model.
Protection
Use ProtectedSpan(start, end) or a (start, end) tuple to protect a source range:
from spokenform import ProtectedSpan, prepare
text = "Keep Dr. literal, but verbalize 12."
start = text.index("Dr.")
prepared = prepare(
text,
language="en",
protected_spans=[ProtectedSpan(start, start + 3)],
)
Invalid or overlapping ranges warn by default and raise ProtectionError with
strict=True. URLs, email addresses, and semantic versions are protected
automatically.
Offset mapping
from spokenform import prepare
source = "Prof. Klein has 2 kg."
prepared = prepare(source, language="de")
start = source.index("Prof.")
end = start + len("Prof.")
spoken_start, spoken_end = prepared.offset_map.map_source_span(start, end)
print(prepared.spoken_text[spoken_start:spoken_end])
Use bias="left" or bias="right" when mapping an individual boundary at an
expansion.
CLI
spokenform --lang de "Prof. Klein hat 2 kg."
spokenform --lang de --changes "Prof. Klein hat 2 kg."
spokenform --lang de --json "Prof. Klein hat 2 kg."
spokenform --lang en --spacy-model en_core_web_sm --strict "The board is 2 in. wide."
echo "The value is 2." | spokenform --lang en
Examples
Executable examples are in examples/:
python examples/basic.py
python examples/german.py
python examples/german.py --spacy-model de_core_news_sm
python examples/protected_text.py
python examples/offset_mapping.py
Documentation
Documentation sources use MyST Markdown. No reStructuredText source files are required.
python -m pip install -e .
python -m pip install -r docs/requirements.txt
sphinx-build -W -b html docs docs/_build/html
Development
python -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"
python -m pytest
python -m ruff check .
python -m ruff format --check .
python -m mypy spokenform examples
python -m build
python -m twine check dist/*
On Windows, activate the environment with .venv\Scripts\activate.
Current limits
- One processing language is supported per call.
- Language detection, mixed-language segmentation, and language marking are external.
- SSMD and other markup must be parsed before calling
spokenform. - Date, time, currency, and ordinal grammar is conservative and not exhaustive.
- Spanish parity ownership covers reviewed dates, quantities, temperatures, currencies, and ordinary numbers; time expressions remain caller-managed.
- Italian parity ownership covers reviewed dates, quantities, temperatures, currencies, and ordinary numbers; colon times remain caller-managed.
- Czech and English own reviewed structured and safe plain-number categories;
Czech colon-time candidates remain caller-managed. English owns contextual
single-dot release labels such as
bot 2.0asbot two point oh, while ordinary decimals retain digit-wise zero wording. English years, suffix ordinals, Roman numerals, phone/ID and arbitrary multi-dot sequences, numeric suffixes, and G2P decisions remain downstream in kokorog2p. abbr2wordscurrently exposes final expanded text rather than semantic replacement objects, so stage edits are reconstructed deterministically from diffs.- Trained spaCy pipelines must be installed and version-compatible with the spaCy runtime.
Dependency direction
abbr2words ─┐
├─ spokenform ── kokorog2p
num2words ──┘
spaCy is an optional quality dependency. spokenform remains independent of
language detection, markup parsing, and phoneme generation.
Release versioning
setuptools-scm derives versions from Git tags and writes
spokenform/_version.py during builds. Use annotated tags such as v0.2.2.
The source-tree fallback when SCM metadata has not been generated is the neutral
version 0+unknown; release builds derive their version from the annotated tag.
Before publishing, ensure the released abbr2words>=0.2.4 prerequisite exists
on the target package index and run the checklist in
docs/release-checklist.md.
License
Apache License 2.0.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file spokenform-0.2.4.tar.gz.
File metadata
- Download URL: spokenform-0.2.4.tar.gz
- Upload date:
- Size: 174.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e2d202a19999d810d1f51c1abba14961cadf9ba146ec875ad3ec7e9f1f434a12
|
|
| MD5 |
cbd53cd9ee7d9e84e1a20629f1b46cc4
|
|
| BLAKE2b-256 |
ba2f2e1f8fcb464a017badeb91a5979847adcac1f332d4b01b96504a2e720660
|
File details
Details for the file spokenform-0.2.4-py3-none-any.whl.
File metadata
- Download URL: spokenform-0.2.4-py3-none-any.whl
- Upload date:
- Size: 105.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3815538cf8921ec40a5ef90babde0d6f0203f08eea3ce92ab0d339da18087cdf
|
|
| MD5 |
15f97a039371a6dc7dff1e9ea3f62a79
|
|
| BLAKE2b-256 |
f89a588e465605e74d2b0cc5645a36b6990209a841ff26396ec094bb34e7a0ac
|