tetrak-easyocr-armenian
Armenian language support for EasyOCR — a trained recognition model, installable as a custom network.
Status: alpha, shipping v1 weights.
reader()downloads a trained model and works out of the box. Be clear-eyed about what it is: v1 beats stock EasyOCR by roughly 16x on word recall, but it does not yet beattesseract -l hye. If you want the best available Armenian OCR today and are not tied to EasyOCR, use Tesseract. Use this if you are building on EasyOCR, or want an Armenian base to fine-tune. The numbers, measured on real scans, are on the model card.
Usage
pip install tetrak-easyocr-armenian
import tetrak_hy
reader = tetrak_hy.reader() # an easyocr.Reader, Armenian-ready
results = reader.readtext("scan.png")
reader() accepts everything easyocr.Reader does (gpu=, verbose=,
…) and handles the custom-network plumbing: the network config and weights
are materialised into ~/.tetrak_hy/ on first use, and the quirks of
EasyOCR's custom-model loading path (there are a few) stay our problem
rather than yours.
The weights live in the Hugging Face model repository, which is canonical for them, and each library version pins one immutable Hub revision — so the model you get is decided by the version of this package you installed, never by what happens to be current upstream.
Cached weights are verified against the release's SHA-256 every time they are loaded, not only when they are downloaded. A file that matches is used as it is — so a machine with no outbound access works once the weights are in place — and one that does not, because it was corrupted or because you have upgraded to a release carrying a new model, is replaced by a fresh verified download.
Set TETRAK_HY_HOME to put the cache somewhere other than your home
directory, or choose it per call:
reader = tetrak_hy.reader(cache_dir="/srv/models/tetrak_hy")
A local .pth passed as weights_path= is kept in a local/
subdirectory of the cache, so testing your own weights does not disturb
the released ones.
What it is
EasyOCR does not ship Armenian. This package adds it as a custom recognition network: EasyOCR's own generation2 architecture (VGG + BiLSTM + CTC), trained for the Armenian script — the full alphabet, the և ligature, Armenian punctuation (՝ ՛ ՞ ՜ ։ ֊ « »), digits and basic Latin for mixed material. Detection is untouched: EasyOCR's CRAFT detector already finds Armenian text; reading it is what was missing.
The import name tetrak_hy is also the EasyOCR network name — EasyOCR
imports this package directly as the model architecture. If you prefer to
wire the Reader yourself:
import easyocr
reader = easyocr.Reader(
["en"], # see note below
recog_network="tetrak_hy",
user_network_directory="~/.tetrak_hy",
model_storage_directory="~/.tetrak_hy",
)
(["en"], not ["hy"]: EasyOCR looks up a per-language character file it
does not have for Armenian. The setting is decorative for custom models —
the model's own character list governs decoding — and reader() hides
this entirely.)
Provenance
The model is trained by tetrak-hy-trainer on synthetic line crops: text from human-proofread pages of the Armenian Soviet Encyclopedia on Armenian Wikisource (CC BY-SA 3.0), rendered in Armenian faces at real scan sizes and degraded to look scanned. v1 is trained on that synthetic data alone — fine-tuning on crops cut from actual scans is the next step, and the one expected to close the remaining gap to Tesseract.
Every weights release carries a provenance record — data recipe, fonts,
dataset revision, training config and checksums — published as
provenance.json beside the weights. The training data is published
too, as
tetrak/armenian-ocr-crops.
Built for Tetrak, a local-first transcription
pipeline for archival material, which consumes this package as its
easyocr-hy backend — but nothing here depends on Tetrak.
Licence
Apache License 2.0 — see LICENSE and NOTICE. The architecture re-exported here is EasyOCR's (Apache 2.0).
Development
python3 -m venv .venv && source .venv/bin/activate
pip install -e '.[dev]'
pytest
ruff check src tests && ruff format --check src tests
Commits follow Conventional Commits,
enforced by the hook in .githooks/ (git config core.hooksPath .githooks
after cloning, or lefthook install). Releases and CHANGELOG.md are
generated from those commits automatically on every push to main.
See CONTRIBUTING.md for the full workflow — checks, the dependency lockfile, and what the automation expects — and SECURITY.md for how to report a vulnerability.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file tetrak_easyocr_armenian-0.3.0.tar.gz.
File metadata
- Download URL: tetrak_easyocr_armenian-0.3.0.tar.gz
- Upload date:
- Size: 165.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3fe8d6899ae27403168d1536901cc2cd877d381af9ddaecd4a7e7c023af4222f
|
|
| MD5 |
59a8a60ba6bde6eab1c911b4b30bce86
|
|
| BLAKE2b-256 |
d56fa86e5b7974c7b78fde3cac025e55f7d853cc8e95cfdec1f9d17a724bc607
|
Provenance
The following attestation bundles were made for tetrak_easyocr_armenian-0.3.0.tar.gz:
Publisher:
release.yml on scattercode/tetrak-easyocr-armenian
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
tetrak_easyocr_armenian-0.3.0.tar.gz -
Subject digest:
3fe8d6899ae27403168d1536901cc2cd877d381af9ddaecd4a7e7c023af4222f - Sigstore transparency entry: 2663976139
- Sigstore integration time:
-
Permalink:
scattercode/tetrak-easyocr-armenian@76c35355d045687e68dd52ee82a24c74e31eb485 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/scattercode
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@76c35355d045687e68dd52ee82a24c74e31eb485 -
Trigger Event:
push
-
Statement type:
File details
Details for the file tetrak_easyocr_armenian-0.3.0-py3-none-any.whl.
File metadata
- Download URL: tetrak_easyocr_armenian-0.3.0-py3-none-any.whl
- Upload date:
- Size: 14.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fb218514994a0d51cfb5461712691018d82232223ae02fe7175eaa3dc634fbf1
|
|
| MD5 |
3bdf875d88aca7d6a689bbd431a5cd9f
|
|
| BLAKE2b-256 |
9e1aca430fdc8002fb6e6172eed5133157b4c645bb5b31d622e435492f0c9983
|
Provenance
The following attestation bundles were made for tetrak_easyocr_armenian-0.3.0-py3-none-any.whl:
Publisher:
release.yml on scattercode/tetrak-easyocr-armenian
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
tetrak_easyocr_armenian-0.3.0-py3-none-any.whl -
Subject digest:
fb218514994a0d51cfb5461712691018d82232223ae02fe7175eaa3dc634fbf1 - Sigstore transparency entry: 2663976407
- Sigstore integration time:
-
Permalink:
scattercode/tetrak-easyocr-armenian@76c35355d045687e68dd52ee82a24c74e31eb485 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/scattercode
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@76c35355d045687e68dd52ee82a24c74e31eb485 -
Trigger Event:
push
-
Statement type: