Skip to main content

Egyptian National ID OCR

A local, open-source pipeline that extracts structured data from Egyptian National ID cards — front and back, any scan condition. No cloud APIs, no LLMs: everything runs on your machine, and nothing about the card ever leaves it.

Egyptian IDs only. The card layout, the checksum, the governorate table and the field vocabularies are all specific to this document. It will not read another country's ID.

Read docs/LIMITATIONS.md before relying on this for anything — it documents exactly what does and doesn't work yet, measured against a real card scanned four different ways, not estimated.

Four ways to call it — command line, Python, HTTP API, or MCP server, all below. Every one of them takes the same input and returns the same fields; docs/INTERFACES.md is the full reference for each — exact request/response shapes, every field documented, and real example output (synthetic card, never a real one).

What it extracts

Front: national ID number (checksum-validated), first/full name, address, date of birth, card serial number.

Back: national ID number, issue date, expiry date (with is_expired), profession, gender, religion, marital status.

Every free-text field also gets a machine-transliterated Latin-script version (first_name_english, address_english, …) for display and search. That transliteration is not the official spelling on the holder's documents — see the field's own description in models/id_card.py for why that can't be derived from the Arabic.

Install

pip install egyptian-national-id-ocr

From source:

git clone https://github.com/HassanSalama2001/Egyptian-National-ID-Open-Source.git
cd Egyptian-National-ID-Open-Source
pip install -e .

Command line

egy-nid-ocr extract path/to/card.jpg
egy-nid-ocr extract path/to/card.jpg --output json
egy-nid-ocr doctor    # checks the runtime is set up correctly

Full flag reference and example JSON output: docs/INTERFACES.md § Command line.

Python

import cv2
from egyptian_national_id_ocr.core.pipeline import Pipeline

pipeline = Pipeline()
image = cv2.imread("card.jpg")
result = pipeline.process_image(image)

print(result.status, result.confidence)
if result.front:
    print(result.front.national_id, result.front.first_name)

Full IDCard field reference (every field, every enum value): docs/INTERFACES.md § Python SDK.

HTTP API

uvicorn src.app:app --reload
curl -X POST "http://localhost:8000/ocr?include_images=false" \
     -F "file=@card.jpg"

include_images=false drops the base64 card/crop images from the response — the difference between a payload measured in kilobytes and one measured in megabytes.

Full request/response shapes, error codes, and example output: docs/INTERFACES.md § HTTP API.

MCP server

Exposes the pipeline as tools an AI assistant can call directly.

pip install "egyptian-national-id-ocr[mcp]"
egy-nid-ocr-mcp

Tool signatures and example results: docs/INTERFACES.md § MCP server.

Web demo

A React UI under web_ui/ that shows the extraction alongside the crop it came from, for verifying results by eye.

cd web_ui && npm install && npm run dev

Why it's reliable on bad scans

Real ID photos are rarely clean: black-and-white scans, zoomed-out shots, odd lighting, off-angle cameras. The pipeline is built around that, not around a clean-input assumption:

  • Card detection first, fields second. The card itself is located and perspective-corrected before any field coordinates are applied — fixed boxes are only meaningful once the card is rectified.
  • A binarization ladder for digits (multiple thresholding methods × upscale factors), accepted only when two independent methods agree — not the first plausible-looking read.
  • Checksum validation and repair on the national ID number: Egyptian IDs carry a mod-11 check digit, so a misread digit can often be detected and corrected rather than silently returned wrong.
  • Cross-field validation, not just per-field: the back's issue and expiry dates check each other against the card's fixed 7-year validity period; an unreadable expiry degrades to what the issue date implies (year and month) rather than to nothing — but the day is never invented, and is_expired answers null rather than guess when the day is what it would take to decide.
  • The national ID and the names are the priority fields — everything else (birth date, governorate, gender, century) is derivable from the checksum-validated ID number, so those get the most validation.

None of this is asserted — see docs/LIMITATIONS.md for the actual numbers, field by field, scan condition by scan condition, regenerated by scripts/benchmark_own.py on every change.

Privacy

Extraction happens entirely on your machine. No card image, and no extracted field, is sent anywhere by this library. See docs/LIMITATIONS.md for what "reliable" currently means in practice, and SECURITY.md for how to report a security issue.

Contributing

See CONTRIBUTING.md and docs/RELEASING.md.

License

MIT — see LICENSE.

Release files for egyptian-national-id-ocr 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for egyptian-national-id-ocr 0.4.0
File Size Uploaded
egyptian_national_id_ocr-0.4.0.tar.gz 2.2 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for egyptian-national-id-ocr 0.4.0
File Interpreter ABI Platform
egyptian_national_id_ocr-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 4.4 MB

Release files / egyptian_national_id_ocr-0.4.0.tar.gz

Download URL egyptian_national_id_ocr-0.4.0.tar.gz
Size 2.2 MB
Tags Source
SHA-256 checksum
How to use checksums
a97e9a871655ed8c0ebb1d0869c7b3dc70a8ad507c360323810cfc121dc7ac09
BLAKE2b-256 checksum
How to use checksums
5d3e42079980bdb9b155bfa5de12855e48d3dd3159b432060063f99bac221184
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / egyptian_national_id_ocr-0.4.0-py3-none-any.whl

Download URL egyptian_national_id_ocr-0.4.0-py3-none-any.whl
Size 2.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
8f84123b904953901b343f1c01a36ec70294b135d58e74a4a0cb363044517ab1
BLAKE2b-256 checksum
How to use checksums
428f607ba0243efa743315cbd8d1913b8b40d0cd0e2827a17e7aa2d4bb6cf430
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page