Field Entity Analyzer
field_entity_analyzer is a high-performance, 100% offline, self-contained Python library for identifying, organizing, and classifying entity types from raw text strings and multi-line OCR document scans (such as Email, Phone Number, Postal Code, Credit Card, Government ID, Address, and Person Name).
Key Features
- 100% Offline & Private: Zero network requests, zero external web service dependencies.
- Ultra Lightweight: Embedded model weights packaged inside binary wheel (< 1 MB wheel size).
- Disorganized Multi-Line Document Parsing: Automatically cleans, parses, merges split addresses, and groups disorganized multi-line inputs into structured entity buckets.
- Multi-Engine Cascading Architecture:
- OCR Pre-processing & Noise Repair Engine: Strips boundary scan noise (
===,~~~,|___), normalizes whitespace, and repairs character swaps (O$\leftrightarrow$0,I$\leftrightarrow$1,S$\leftrightarrow$5,Z$\leftrightarrow$2). - Pattern & Algorithmic Rule Engine: Regex matching and checksum validation (e.g. Luhn algorithm for Credit Cards, SSN/PAN/Aadhaar formats).
- Contextual Heuristic & Token Parser: Structural layout analysis, address indicator tokens, key-value anchors (
Name:,Address:,Email:), title prefixes, name patterns. - Lightweight Embedded Classifier: Fast, pre-trained character and token feature classifier asset embedded inside the package.
- OCR Pre-processing & Noise Repair Engine: Strips boundary scan noise (
Supported Entity Types
EMAILPHONEPOSTAL_CODECREDIT_CARDGOVERNMENT_IDADDRESSNAMEUNKNOWN
Installation
pip install field_entity_analyzer
Quick Start & Disorganized Multi-Line Demo
Even if your input text is completely scrambled or disorganized (e.g., Name $\rightarrow$ Email $\rightarrow$ Phone $\rightarrow$ Name $\rightarrow$ Address $\rightarrow$ Government ID), analyze_document() automatically cleans the noise, repairs OCR swaps, merges split address lines, and organizes the entities into clean categorized buckets (grouped_entities).
from field_entity_analyzer import EntityAnalyzer, analyze_entity
# Initialize analyzer
analyzer = EntityAnalyzer()
# Disorganized & Noisy Multi-line Input String (Scrambled Order + OCR Noise)
disorganized_ocr_text = """
=========================================================
|___ CONFIDENTIAL DOCUMENT SCAN ___|
~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~
Name: J0hn D0e
sarah.connor @ sky.net
Phone: +1 (S55) O19-2834
Alexander Hamilton
Flat 4B, Bldg 12
Baker Road, Sector 5
Springfield Pincode 90210
ID No: SSN-123-45-6789
Card No: 4532 O151 1283 S366
=========================================================
"""
# Analyze multi-line text
result = analyzer.analyze_document(disorganized_ocr_text)
# 1. Organized Entity Buckets (Grouped by Category)
print("=== ORGANIZED ENTITIES ===")
for entity_type, items in result["grouped_entities"].items():
if items:
print(f"{entity_type}: {items}")
# 2. Extracted Key-Value Label Anchors
print("\n=== ANCHORED FIELDS ===")
print(result["anchored_fields"])
# 3. Merged Multi-line Address Blocks
print("\n=== MERGED ADDRESSES ===")
print(result["aggregated_addresses"])
Expected Output
=== ORGANIZED ENTITIES ===
EMAIL: ['sarah.connor@sky.net']
PHONE: ['+1 (555) 019-2834']
CREDIT_CARD: ['4532 0151 1283 5366']
GOVERNMENT_ID: ['SSN-123-45-6789']
ADDRESS: ['Flat 4B, Bldg 12, Baker Road, Sector 5, Springfield Pincode 90210']
NAME: ['J0hn D0e', 'Alexander Hamilton']
=== ANCHORED FIELDS ===
{'name': 'J0hn D0e', 'phone': '+1 (S55) O19-2834', 'card_number': '4532 O151 1283 S366', 'id_number': 'SSN-123-45-6789'}
=== MERGED ADDRESSES ===
['Flat 4B, Bldg 12, Baker Road, Sector 5, Springfield Pincode 90210']
Single String Usage
# Analyze a single string value
res = analyze_entity("john.doe@example.com")
print(res.entity_type) # EntityType.EMAIL
print(res.confidence) # 1.0
print(res.engine) # "rules"
License
MIT License.
Metadata
Release files for field-entity-analyzer 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| field_entity_analyzer-0.3.0.tar.gz | 58.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| field_entity_analyzer-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 113.8 kB
Release files / field_entity_analyzer-0.3.0.tar.gz
| Download URL | field_entity_analyzer-0.3.0.tar.gz |
|---|---|
| Size | 58.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
13b5a398d83533e942b0bb8b106cec89542e54f0c16815616c8a5a6d9aecb5a3
|
|
BLAKE2b-256 checksum How to use checksums |
f063bb6b21dd9fdaf5c4141fdc06ed48bad9e8626eda776b1d52e9d665e53082
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.10
|
Release files / field_entity_analyzer-0.3.0-py3-none-any.whl
| Download URL | field_entity_analyzer-0.3.0-py3-none-any.whl |
|---|---|
| Size | 55.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f8023d03dcbb13b04ec46d9de3de61c685ef5549c0c18f330294d9e8dc8ff631
|
|
BLAKE2b-256 checksum How to use checksums |
5fc91a1326e81139934d7213f78697d2646721a3a0a283d9be773e074ecda678
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.10
|