Skip to main content

lexhint

Lexhint is a local lexical-evidence engine backed by self-describing SQLite language databases. It provides lexical membership, optional corpus commonness, compact-string segmentation, semantic-domain evidence, and optional rich dictionary entries.

Lexhint does not normalize or speak text. Word boundaries, acronyms, URLs, numbers, versions, and pronunciation policy belong to the consuming application.

Install

python -m pip install lexhint

Lexhint supports Python 3.10 through 3.14. The Python package is code-only and does not include a complete language SQLite artifact. Install a prebuilt artifact from the lexhint-datasets repository, or build one locally as shown below. Keep dataset downloads and their licensing and provenance separate from the Python package.

Quick start

After installing the package and an English artifact, run:

lexhint word compiler
lexhint segment compilerword
lexhint context "The compiler is 8.3.2." --target 16:21
lexhint dictionary word compiler
lexhint dictionary status

For a small local artifact without FrequencyWords enrichment, build from the repository fixture with lexhint dictionary build en --source tests/fixtures/kaikki-mini.jsonl --output /tmp/lexhint-en.sqlite3 --no-frequency and pass --path /tmp/lexhint-en.sqlite3 to the query commands.

1. Common-word lexicon

Open a local artifact with Lexicon:

from lexhint import Lexicon

lexicon = Lexicon.from_path("en.sqlite3")
info = lexicon.word("compiler")
print(info.known, info.frequency_rank)
print(lexicon.segment("compilerword"))

Runtime access is local-only, deterministic, read-only, and never fetches missing entries or mutates the database. segment() and semantic context operations require full authoritative coverage.

The public runtime operations are:

lexicon.word("compiler")
lexicon.contains("compiler")
lexicon.segment("chatgpt")
lexicon.context_domains(text, target=(start, end))
lexicon.supports_domain(text, target=(start, end), domain="computing")
lexicon.entries("compiler")  # rich artifacts only

Dictionary membership is authoritative. Frequency rank and count enrich existing lexemes but never create corpus-only words. Case evidence is retained, so an uppercase-only GPT entry does not validate lowercase gpt.

Semantic results are explainable DomainEvidence values containing score and nearby ContextCue records. The target span is always excluded, and missing evidence is not negative evidence.

Build an artifact

The default build creates a full lexical,semantic,dictionary artifact and automatically acquires the pinned full FrequencyWords source:

lexhint dictionary build en

Use a local or remote dictionary source and explicit build policies when needed:

lexhint dictionary build en --source ./raw-wiktextract-data.jsonl.gz
lexhint dictionary build en --capabilities lexical,semantic --no-frequency
lexhint dictionary build en --profile runtime
lexhint dictionary build en --frequency-source ./en_full.txt
lexhint dictionary build en --refresh-frequency
lexhint --offline dictionary build en --source ./raw-wiktextract-data.jsonl.gz

Capabilities are canonicalized in the order lexical,semantic,dictionary. semantic and dictionary require lexical. Profiles are shortcuts: runtime means lexical,semantic, and rich means lexical,semantic,dictionary.

Frequency enrichment is independent of capabilities. Use --no-frequency for a valid lexical artifact without corpus data. Automatic sources are cached under ~/.cache/lexhint/sources/frequencywords/<revision>/, or an equivalent XDG/LEXHINT_CACHE_DIR location. Builds record source URLs, revisions, hashes, schema, capabilities, and builder metadata. Build configuration and progress are written to stderr, while the final result, including JSON, is written to stdout.

CLI queries

lexhint word compiler
lexhint segment chatgpt
lexhint context "The compiler is 8.3.2." --target 16:21
lexhint dictionary word compiler
lexhint dictionary status

Dictionary word output has three human-readable detail levels. The default standard view shows all senses with compact metadata. Use compact for a deliberately short shell view, or full for every field retained by the local Lexhint dictionary model:

lexhint dictionary word love
lexhint dictionary word love --detail compact
lexhint dictionary word love --detail full
lexhint dictionary word love --detail full --hide examples,tags
lexhint dictionary word love --detail compact --show examples
lexhint dictionary word love --pos noun,verb --exclude-pos proper_noun
lexhint --json dictionary word love --pos noun

The --show and --hide options accept repeatable comma-separated fields. Canonical fields are etymology, pronunciations, forms, tags, topics, examples, synonyms, and antonyms; the all, entry, sense, and relations groups are also supported. --width controls human output from 40 through 240 columns.

Use --json for stable, complete machine-readable output. POS selection applies to JSON entries, while --detail, --show, --hide, and --width are human-only options. dictionary status reports current SQL row counts, capabilities, provenance, size, and build metadata without rebuilding. Use --path as an advanced override when inspecting a specific artifact. Rich dictionary lookup reports a controlled capability error for compact runtime artifacts.

Data and scope

The builder consumes Wiktextract-compatible JSONL, commonly from Kaikki, and FrequencyWords full files for optional corpus enrichment. See DATA_SOURCES.md for source and licensing information.

Lexhint does not implement Spokenform integration, dataset publication, URL parsing, speech rendering, or consumer-specific interpretation rules. The separate buchwandler/lexhint-datasets repository is outside this project.

Development

Contributor setup uses an editable installation with development tools:

git clone https://github.com/buchwandler/lexhint.git
cd lexhint
python -m pip install -e ".[dev]"
pytest -q
ruff check .
mypy lexhint

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

lexhint-0.1.0.tar.gz (56.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

lexhint-0.1.0-py3-none-any.whl (38.7 kB view details)

Uploaded Python 3

File details

Details for the file lexhint-0.1.0.tar.gz.

File metadata

  • Download URL: lexhint-0.1.0.tar.gz
  • Upload date:
  • Size: 56.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for lexhint-0.1.0.tar.gz
Algorithm Hash digest
SHA256 b872ba5ff8cd0359f12601d13e1e6e7e9d203c8949e7e9c97415f40223f497da
MD5 891c86b61a92fca983da6f3cc7a33494
BLAKE2b-256 24dd6489f1ba90b80592c2fb8d13d59e66f482173b53645df5bdb2cc42ebfa56

See more details on using hashes here.

File details

Details for the file lexhint-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: lexhint-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 38.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for lexhint-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e3888a400df0c782703634188eedf030b50947d44e5250334a77b9f8283f52cb
MD5 7ef2b1593e530fcdac10a948aad7afa6
BLAKE2b-256 544ae89bc183a0ed27bfa1de69c29cbcd79af86d003bbeada191c110b18cdb16

See more details on using hashes here.

Release history Release notifications | RSS feed

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

This release

0.1.0 This release

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page