Skip to main content

spokenform

spokenform converts plain written text in one selected language into reviewable text intended for speech systems. It is a text-to-text frontend: written text in, spoken form out.

The package provides:

  • context-aware abbreviation and source-aligned numeric-unit expansion through abbr2words;
  • a locale-aware structured-value stage for quantities, dates, times, currencies, temperatures, labels, and contextual ordinals;
  • locale-policy-wrapped number, date, time, currency, decimal, and ordinal verbalization;
  • optional provider-neutral spaCy annotations for POS-aware abbreviation rules;
  • stage-level provenance through PreparedText;
  • composed input-to-output offset maps with left/right boundary bias;
  • caller-defined and automatically discovered protected ranges;
  • conservative protection for URLs, email addresses, and semantic versions.

It intentionally does not detect languages, parse or render SSMD, segment mixed languages, generate phonemes, or depend on kokorog2p.

Installation

python -m pip install spokenform

For optional spaCy integration:

python -m pip install "spokenform[spacy]"
python -m spacy download de_core_news_sm

spaCy and its trained pipelines are separate packages. spokenform never downloads a model automatically.

Quickstart

from spokenform import prepare

prepared = prepare(
    "Prof. Klein bringt am 14.05.2026 um 18:20 Uhr 2 kg mit.",
    language="de",
)

print(prepared.spoken_text)
print(prepared.render_changes())

The result contains:

  • source_text: unchanged caller input;
  • clean_text: plain text used by the normalization pipeline;
  • spoken_text: readable normalized output;
  • language: normalized processing-language code;
  • ordered stages and mapped edits;
  • a composed offset_map;
  • structured warnings.

PreparedText.text is an alias for spoken_text.

kokorog2p adapter

Use prepare_for_kokorog2p(text, language=...) for one explicitly selected language run. The adapter preserves caller-owned run whitespace and protected overrides, emits exact source-coordinate replacements, and leaves tokenization, G2P, phonemization, and model punctuation to kokorog2p. German and French are parity-gated semantic migration targets; Czech, English, Spanish, Italian, and Portuguese number categories remain caller-managed.

German quantity symbols are recognized by abbr2words.iter_unit_matches(). spokenform owns only the canonical German grammar that realizes those matches, including gender, invariant Stück, currency decomposition, and lexical decimal digits. French likewise realizes canonical abbr2words quantity and currency identities, including French dates, times, ordinals, decimal digits, plural grammar, temperatures, and major/minor currency units. Neither locale copies raw symbol inventories or downstream tokenizer/phoneme rules.

Language boundary

Each call processes one language. Production callers should always pass language=...; English remains the API default for compatibility and simple CLI usage.

Language detection and mixed-language handling belong in the orchestration or G2P layer. A foreign word may remain unchanged through normalization and be handled afterward. Existing source spans can be transferred with prepared.offset_map.

Markup must also be parsed outside this package. Pass plain text to prepare() and use ProtectedSpan for ranges generic normalization must not change.

Configuration

from spokenform import PreparationConfig, prepare

config = PreparationConfig(
    language="en",
    expand_abbreviations=True,
    expand_structured=True,
    expand_numbers=True,
    normalize_whitespace=True,
    context=True,
)

prepared = prepare("The board is 2 in. wide.", config=config)

When a PreparationConfig is supplied, it is authoritative for pipeline options.

spaCy support

spaCy supplies POS annotations for abbreviation rules that opt into POS guards. The public normalization API remains provider-neutral.

abbr2words accepts POS annotations, but its bundled registries do not necessarily require POS labels. Therefore installing spaCy alone may not change default normalization output. The integration is usable for custom POS-guarded entries.

Load and inject a pipeline in the application:

import spacy
from spokenform import prepare

nlp = spacy.load("en_core_web_sm")
prepared = prepare(
    "The board is 2 in. wide.",
    language="en",
    nlp=nlp,
)

Or ask spokenform to load an already installed model:

prepared = prepare(
    "The board is 2 in. wide.",
    language="en",
    spacy_model="en_core_web_sm",
    strict=True,
)

Model names and paths are passed to spacy.load(). Loaded models are cached by language/model key. reset_spacy_cache() clears that cache.

The adapter reads the token attributes text, idx, pos_, tag_, lemma_, and lang_. lang_ is carried as provider metadata; it is not used as language detection. Annotation spans are validated against the exact input text and remapped when protected ranges are replaced by internal sentinels. A trained pipeline with POS or morphological annotations is required for quality improvement; spacy.blank(...) supplies tokenization but normally no useful POS tags.

Explicit annotations take precedence over nlp and spacy_model.

Protection

Use ProtectedSpan(start, end) or a (start, end) tuple to protect a source range:

from spokenform import ProtectedSpan, prepare

text = "Keep Dr. literal, but verbalize 12."
start = text.index("Dr.")
prepared = prepare(
    text,
    language="en",
    protected_spans=[ProtectedSpan(start, start + 3)],
)

Invalid or overlapping ranges warn by default and raise ProtectionError with strict=True. URLs, email addresses, and semantic versions are protected automatically.

Offset mapping

from spokenform import prepare

source = "Prof. Klein has 2 kg."
prepared = prepare(source, language="de")

start = source.index("Prof.")
end = start + len("Prof.")
spoken_start, spoken_end = prepared.offset_map.map_source_span(start, end)

print(prepared.spoken_text[spoken_start:spoken_end])

Use bias="left" or bias="right" when mapping an individual boundary at an expansion.

CLI

spokenform --lang de "Prof. Klein hat 2 kg."
spokenform --lang de --changes "Prof. Klein hat 2 kg."
spokenform --lang de --json "Prof. Klein hat 2 kg."
spokenform --lang en --spacy-model en_core_web_sm --strict "The board is 2 in. wide."
echo "The value is 2." | spokenform --lang en

Examples

Executable examples are in examples/:

python examples/basic.py
python examples/german.py
python examples/german.py --spacy-model de_core_news_sm
python examples/protected_text.py
python examples/offset_mapping.py

Documentation

Documentation sources use MyST Markdown. No reStructuredText source files are required.

python -m pip install -e .
python -m pip install -r docs/requirements.txt
sphinx-build -W -b html docs docs/_build/html

Development

python -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"

python -m pytest
python -m ruff check .
python -m ruff format --check .
python -m mypy spokenform examples
python -m build
python -m twine check dist/*

On Windows, activate the environment with .venv\Scripts\activate.

Current limits

  • One processing language is supported per call.
  • Language detection, mixed-language segmentation, and language marking are external.
  • SSMD and other markup must be parsed before calling spokenform.
  • Date, time, currency, and ordinal grammar is conservative and not exhaustive.
  • abbr2words currently exposes final expanded text rather than semantic replacement objects, so stage edits are reconstructed deterministically from diffs.
  • Trained spaCy pipelines must be installed and version-compatible with the spaCy runtime.

Dependency direction

abbr2words ─┐
            ├─ spokenform ── kokorog2p
num2words ──┘

spaCy is an optional quality dependency. spokenform remains independent of language detection, markup parsing, and phoneme generation.

Release versioning

setuptools-scm derives versions from Git tags and writes spokenform/_version.py during builds. Use annotated tags such as v0.1.0. The source snapshot fallback is 0.1.0.

Before publishing, ensure the required abbr2words version exists on the target package index and run the checklist in docs/release-checklist.md.

License

Apache License 2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

spokenform-0.2.1.tar.gz (78.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

spokenform-0.2.1-py3-none-any.whl (43.8 kB view details)

Uploaded Python 3

File details

Details for the file spokenform-0.2.1.tar.gz.

File metadata

  • Download URL: spokenform-0.2.1.tar.gz
  • Upload date:
  • Size: 78.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for spokenform-0.2.1.tar.gz
Algorithm Hash digest
SHA256 e806cb46581996bddb3e16f077e83fd616cb51decc2b1dafd683102d0f255639
MD5 743a0c95ce8301d8c06e0da8827bcfe3
BLAKE2b-256 968cc243a8e74389028e494c48eb179e338d0e271850758adaed008ab3932439

See more details on using hashes here.

File details

Details for the file spokenform-0.2.1-py3-none-any.whl.

File metadata

  • Download URL: spokenform-0.2.1-py3-none-any.whl
  • Upload date:
  • Size: 43.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for spokenform-0.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 e2c26c3bab9d51272e2b4f02b073abfa48d6c16c367d450b0f326286bd838800
MD5 3b466e42835c897817c674276e829c12
BLAKE2b-256 3d0501a827544dc55ddd0b09bacad86a053867932a762539256ad41d2b6c56a9

See more details on using hashes here.

Release history Release notifications | RSS feed

0.3.1

2 files

0.3.0

2 files

0.2.8

2 files

0.2.7

2 files

0.2.6

2 files

0.2.5

2 files

0.2.4

2 files

0.2.3

2 files

0.2.2

2 files

This release

0.2.1 This release

2 files

0.2.0

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page