lexhint
Lexhint is a local lexical-evidence engine backed by self-describing SQLite language databases. It provides lexical membership, optional corpus commonness, compact-string segmentation, semantic-domain evidence, and optional rich dictionary entries.
Lexhint does not normalize or speak text. Word boundaries, acronyms, URLs, numbers, versions, and pronunciation policy belong to the consuming application.
Install
python -m pip install lexhint
Lexhint supports Python 3.10 through 3.14. The Python package is code-only and does not include a complete language SQLite artifact. Install a prebuilt artifact from the lexhint-datasets repository, or build one locally as shown below. Keep dataset downloads and their licensing and provenance separate from the Python package.
Quick start
After installing the package and an English artifact, run:
lexhint word compiler
lexhint segment compilerword
lexhint context "The compiler is 8.3.2." --target 16:21
lexhint dictionary word compiler
lexhint dictionary status
For a small local artifact without FrequencyWords enrichment, build from the repository fixture with lexhint dictionary build en --source tests/fixtures/kaikki-mini.jsonl --output /tmp/lexhint-en.sqlite3 --no-frequency and pass --path /tmp/lexhint-en.sqlite3 to the query commands.
1. Common-word lexicon
Open a local artifact with Lexicon:
from lexhint import Lexicon
lexicon = Lexicon.from_path("en.sqlite3")
info = lexicon.word("compiler")
print(info.known, info.frequency_rank)
print(lexicon.segment("compilerword"))
Runtime access is local-only, deterministic, read-only, and never fetches missing entries or mutates the database. segment() and semantic context operations require full authoritative coverage.
The public runtime operations are:
lexicon.word("compiler")
lexicon.contains("compiler")
lexicon.segment("chatgpt")
lexicon.context_domains(text, target=(start, end))
lexicon.supports_domain(text, target=(start, end), domain="computing")
lexicon.entries("compiler") # rich artifacts only
Dictionary membership is authoritative. Frequency rank and count enrich existing lexemes but never create corpus-only words. Case evidence is retained, so an uppercase-only GPT entry does not validate lowercase gpt.
Semantic results are explainable DomainEvidence values containing score and nearby ContextCue records. The target span is always excluded, and missing evidence is not negative evidence.
Build an artifact
The default build creates a full lexical,semantic,dictionary artifact and automatically acquires the pinned full FrequencyWords source:
lexhint dictionary build en
Use a local or remote dictionary source and explicit build policies when needed:
lexhint dictionary build en --source ./raw-wiktextract-data.jsonl.gz
lexhint dictionary build en --capabilities lexical,semantic --no-frequency
lexhint dictionary build en --profile runtime
lexhint dictionary build en --frequency-source ./en_full.txt
lexhint dictionary build en --refresh-frequency
lexhint --offline dictionary build en --source ./raw-wiktextract-data.jsonl.gz
Capabilities are canonicalized in the order lexical,semantic,dictionary. semantic and dictionary require lexical. Profiles are shortcuts: runtime means lexical,semantic, and rich means lexical,semantic,dictionary.
Frequency enrichment is independent of capabilities. Use --no-frequency for a valid lexical artifact without corpus data. Automatic sources are cached under ~/.cache/lexhint/sources/frequencywords/<revision>/, or an equivalent XDG/LEXHINT_CACHE_DIR location. Builds record source URLs, revisions, hashes, schema, capabilities, and builder metadata. Build configuration and progress are written to stderr, while the final result, including JSON, is written to stdout.
CLI queries
lexhint word compiler
lexhint segment chatgpt
lexhint context "The compiler is 8.3.2." --target 16:21
lexhint dictionary word compiler
lexhint dictionary status
Dictionary word output has three human-readable detail levels. The default standard view shows all senses with compact metadata. Use compact for a deliberately short shell view, or full for every field retained by the local Lexhint dictionary model:
lexhint dictionary word love
lexhint dictionary word love --detail compact
lexhint dictionary word love --detail full
lexhint dictionary word love --detail full --hide examples,tags
lexhint dictionary word love --detail compact --show examples
lexhint dictionary word love --pos noun,verb --exclude-pos proper_noun
lexhint --json dictionary word love --pos noun
The --show and --hide options accept repeatable comma-separated fields. Canonical fields are etymology, pronunciations, forms, tags, topics, examples, synonyms, and antonyms; the all, entry, sense, and relations groups are also supported. --width controls human output from 40 through 240 columns.
Use --json for stable, complete machine-readable output. POS selection applies to JSON entries, while --detail, --show, --hide, and --width are human-only options. dictionary status reports current SQL row counts, capabilities, provenance, size, and build metadata without rebuilding. Use --path as an advanced override when inspecting a specific artifact. Rich dictionary lookup reports a controlled capability error for compact runtime artifacts.
Data and scope
The builder consumes Wiktextract-compatible JSONL, commonly from Kaikki, and FrequencyWords full files for optional corpus enrichment. See DATA_SOURCES.md for source and licensing information.
Lexhint does not implement Spokenform integration, dataset publication, URL parsing, speech rendering, or consumer-specific interpretation rules. The separate buchwandler/lexhint-datasets repository is outside this project.
Development
Contributor setup uses an editable installation with development tools:
git clone https://github.com/buchwandler/lexhint.git
cd lexhint
python -m pip install -e ".[dev]"
pytest -q
ruff check .
mypy lexhint
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file lexhint-0.1.0.tar.gz.
File metadata
- Download URL: lexhint-0.1.0.tar.gz
- Upload date:
- Size: 56.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b872ba5ff8cd0359f12601d13e1e6e7e9d203c8949e7e9c97415f40223f497da
|
|
| MD5 |
891c86b61a92fca983da6f3cc7a33494
|
|
| BLAKE2b-256 |
24dd6489f1ba90b80592c2fb8d13d59e66f482173b53645df5bdb2cc42ebfa56
|
File details
Details for the file lexhint-0.1.0-py3-none-any.whl.
File metadata
- Download URL: lexhint-0.1.0-py3-none-any.whl
- Upload date:
- Size: 38.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e3888a400df0c782703634188eedf030b50947d44e5250334a77b9f8283f52cb
|
|
| MD5 |
7ef2b1593e530fcdac10a948aad7afa6
|
|
| BLAKE2b-256 |
544ae89bc183a0ed27bfa1de69c29cbcd79af86d003bbeada191c110b18cdb16
|