Skip to main content

Field Entity Analyzer

field_entity_analyzer is a high-performance, 100% offline, self-contained Python library for identifying, organizing, and classifying entity types from raw text strings and multi-line OCR document scans (such as Email, Phone Number, Postal Code, Credit Card, Government ID, Address, and Person Name).

Key Features

  • 100% Offline & Private: Zero network requests, zero external web service dependencies.
  • Ultra Lightweight: Embedded model weights packaged inside binary wheel (< 1 MB wheel size).
  • Disorganized Multi-Line Document Parsing: Automatically cleans, parses, merges split addresses, and groups disorganized multi-line inputs into structured entity buckets.
  • Multi-Engine Cascading Architecture:
    1. OCR Pre-processing & Noise Repair Engine: Strips boundary scan noise (===, ~~~, |___), normalizes whitespace, and repairs character swaps (O $\leftrightarrow$ 0, I $\leftrightarrow$ 1, S $\leftrightarrow$ 5, Z $\leftrightarrow$ 2).
    2. Pattern & Algorithmic Rule Engine: Regex matching and checksum validation (e.g. Luhn algorithm for Credit Cards, SSN/PAN/Aadhaar formats).
    3. Contextual Heuristic & Token Parser: Structural layout analysis, address indicator tokens, key-value anchors (Name:, Address:, Email:), title prefixes, name patterns.
    4. Lightweight Embedded Classifier: Fast, pre-trained character and token feature classifier asset embedded inside the package.

Supported Entity Types

  • EMAIL
  • PHONE
  • POSTAL_CODE
  • CREDIT_CARD
  • GOVERNMENT_ID
  • ADDRESS
  • NAME
  • UNKNOWN

Installation

pip install field_entity_analyzer

Quick Start & Disorganized Multi-Line Demo

Even if your input text is completely scrambled or disorganized (e.g., Name $\rightarrow$ Email $\rightarrow$ Phone $\rightarrow$ Name $\rightarrow$ Address $\rightarrow$ Government ID), analyze_document() automatically cleans the noise, repairs OCR swaps, merges split address lines, and organizes the entities into clean categorized buckets (grouped_entities).

from field_entity_analyzer import EntityAnalyzer, analyze_entity

# Initialize analyzer
analyzer = EntityAnalyzer()

# Disorganized & Noisy Multi-line Input String (Scrambled Order + OCR Noise)
disorganized_ocr_text = """
=========================================================
|___ CONFIDENTIAL DOCUMENT SCAN ___|
~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~ ~~~
Name: J0hn D0e
sarah.connor @ sky.net
Phone: +1 (S55) O19-2834
Alexander Hamilton
Flat 4B, Bldg 12
Baker Road, Sector 5
Springfield Pincode 90210
ID No: SSN-123-45-6789
Card No: 4532 O151 1283 S366
=========================================================
"""

# Analyze multi-line text
result = analyzer.analyze_document(disorganized_ocr_text)

# 1. Organized Entity Buckets (Grouped by Category)
print("=== ORGANIZED ENTITIES ===")
for entity_type, items in result["grouped_entities"].items():
    if items:
        print(f"{entity_type}: {items}")

# 2. Extracted Key-Value Label Anchors
print("\n=== ANCHORED FIELDS ===")
print(result["anchored_fields"])

# 3. Merged Multi-line Address Blocks
print("\n=== MERGED ADDRESSES ===")
print(result["aggregated_addresses"])

Expected Output

=== ORGANIZED ENTITIES ===
EMAIL: ['sarah.connor@sky.net']
PHONE: ['+1 (555) 019-2834']
CREDIT_CARD: ['4532 0151 1283 5366']
GOVERNMENT_ID: ['SSN-123-45-6789']
ADDRESS: ['Flat 4B, Bldg 12, Baker Road, Sector 5, Springfield Pincode 90210']
NAME: ['J0hn D0e', 'Alexander Hamilton']

=== ANCHORED FIELDS ===
{'name': 'J0hn D0e', 'phone': '+1 (S55) O19-2834', 'card_number': '4532 O151 1283 S366', 'id_number': 'SSN-123-45-6789'}

=== MERGED ADDRESSES ===
['Flat 4B, Bldg 12, Baker Road, Sector 5, Springfield Pincode 90210']

Single String Usage

# Analyze a single string value
res = analyze_entity("john.doe@example.com")
print(res.entity_type) # EntityType.EMAIL
print(res.confidence)  # 1.0
print(res.engine)      # "rules"

License

MIT License.

Metadata

Release files for field-entity-analyzer 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for field-entity-analyzer 0.3.0
File Size Uploaded
field_entity_analyzer-0.3.0.tar.gz 58.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for field-entity-analyzer 0.3.0
File Interpreter ABI Platform
field_entity_analyzer-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 113.8 kB

Release files / field_entity_analyzer-0.3.0.tar.gz

Download URL field_entity_analyzer-0.3.0.tar.gz
Size 58.7 kB
Tags Source
SHA-256 checksum
How to use checksums
13b5a398d83533e942b0bb8b106cec89542e54f0c16815616c8a5a6d9aecb5a3
BLAKE2b-256 checksum
How to use checksums
f063bb6b21dd9fdaf5c4141fdc06ed48bad9e8626eda776b1d52e9d665e53082
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.10

Release files / field_entity_analyzer-0.3.0-py3-none-any.whl

Download URL field_entity_analyzer-0.3.0-py3-none-any.whl
Size 55.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f8023d03dcbb13b04ec46d9de3de61c685ef5549c0c18f330294d9e8dc8ff631
BLAKE2b-256 checksum
How to use checksums
5fc91a1326e81139934d7213f78697d2646721a3a0a283d9be773e074ecda678
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.10

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page