Skip to main content

Russian KASKO insurance policy field extractor — pulls structured fields (policy terms + policyholder + contacts) from PDF policies via text-layer, tables, and OCR fallback.

Project description

polis-recognizer

A deterministic field extractor for Russian KASKO insurance policy PDFs. Pulls 7 structured fields without LLMs — text-layer extraction (pypdf) plus optional table-aware reading (pdfplumber), with Tesseract OCR fallback for scanned policies.

Status: pre-stable (0.x). API may change before 1.0.

What it extracts

Field Type Example
policy_number str "AC524160804"
policy_period {start, end} (date) {"start": date(2025, 2, 27), "end": date(2026, 2, 26)}
franchise {value, currency, absent} {"value": 30000, "currency": "RUB", "absent": False}
limit {value, currency} {"value": 5525000, "currency": "RUB"}
premium {value, currency} {"value": 220000, "currency": "RUB"}
sum_type "aggregate" / "non_aggregate" "non_aggregate"
repair_mode "dealer" / "service" / "cash" "dealer"
policyholder {type, name, inn, ogrn, kpp, passport, birth_date} {"type": "legal_entity", "name": "ООО \"Альфа\"", "inn": "7707083893", "ogrn": "1027700132195", "kpp": "770701001", "passport": None, "birth_date": None}
policyholder_contacts {phones, emails, address, postal_code} {"phones": ["+74951234567"], "emails": ["contact@alpha.ru"], "address": "101000, г. Москва, ул. Ленина, д. 1", "postal_code": "101000"}

Quick start

from polis_recognizer import PolicyExtractor

extractor = PolicyExtractor()
result = extractor.extract_from_pdf("/path/to/polis.pdf")

print(result.policy_number)
# → "AC524160804"
print(result.policy_period)
# → {"start": date(2025, 2, 27), "end": date(2026, 2, 26)}
print(result.franchise)
# → {"value": 30000.0, "currency": "RUB", "absent": False}

Input methods:

extractor.extract_from_pdf("polis.pdf")
extractor.extract_from_bytes(pdf_bytes, filename="polis.pdf")
extractor.extract_from_text("сырой текст полиса")  # bypass PDF/OCR

Installation

pip install polis-recognizer

System dependencies

polis-recognizer shells out to Tesseract for OCR and to poppler for PDF→image conversion. These are NOT pip-installable; install them through your OS package manager.

Linux (Debian/Ubuntu):

sudo apt-get install -y tesseract-ocr tesseract-ocr-rus poppler-utils libgl1

macOS (Homebrew):

brew install tesseract tesseract-lang poppler

Windows: best-effort. Install Tesseract for Windows and add it to PATH; install Poppler binaries for PDF support. We don't test on Windows in CI.

The Russian language pack (tesseract-ocr-rus / tesseract-lang) is required — without it, OCR silently falls back to English and Cyrillic documents come back as garbage. The library logs a CRITICAL warning at import time if the pack is missing.

How it works

The extractor runs three stages:

  1. PDF ingestionPdfExtractionRouter tries text-layer extraction first (pypdf for text, pdfplumber for tables on the same page). If the result is too short (fewer than 100 chars by default) or detected as glued/EDI-envelope text, it falls back to Tesseract OCR.
  2. Text normalization — Unicode NFKC, NBSP stripping, hyphenated line-break healing, runs of multiple spaces collapsed.
  3. Field extraction — 7 deterministic parsers (one per field) run regex + table-aware patterns and emit Candidates with confidence scores. A ranker picks the winner per field.

There's no LLM and no cloud dependency. Everything runs locally.

PDF extractor choice

The default is "hybrid" — pypdf text plus pdfplumber tables in one pass. Two alternatives:

Option When to use
"hybrid" (default) Best for KASKO. pypdf preserves date/period text quality, pdfplumber's tables fix the column layout for limit/franchise/premium.
"pypdf" Faster, no tables. Use when document quality is uniform and tables aren't needed.
"pdfplumber" Fully layout-aware. Slower; on KASKO it slightly regresses date parsing. Use for table-heavy non-KASKO formats.
extractor = PolicyExtractor(pdf_extractor="pypdf")

Supported insurer formats

The parser ships with patterns for these Russian insurers' KASKO templates: АльфаСтрахование (XLS form-mask), СОГАЗ-АВТО, Чулпан, Ингосстрах, ВСК, АбсолютСтрахование, Росгосстрах, СОГАЗ Diadoc-wrapped PDFs. Recall on real-world KASKO corpora is ~50-65% per field; pulling above that requires per-format parser additions.

If you have a policy from an insurer not on this list — see CONTRIBUTING.md for how to add a parser pattern.

Configuration

Constructor arguments:

extractor = PolicyExtractor(
    ocr_language="rus+eng",       # Tesseract language string
    ocr_timeout_seconds=300,
    ocr_page_limit=50,
    ocr_max_text_size=500_000,
    pdf_extractor="hybrid",       # "pypdf" | "pdfplumber" | "hybrid"
    image_preprocessing="fallback", # "never" | "fallback" | "always"
    psm=None,                     # Tesseract --psm (None = auto)
    oem=None,                     # Tesseract --oem (None = auto)
    max_image_size_bytes=None,    # reject images larger than this
    extract_pii=False,            # opt-in for passport + birth date
)

extract_pii

PolicyExtractor does not extract passport or birth date by default — with extract_pii=False, policyholder.passport and policyholder.birth_date are always None even when the source text contains them. This keeps the default output safe to log / cache / persist without extra redaction work. Set extract_pii=True to opt in. Operational contact data (phone, email, address) is not gated — those handle data the caller has a legal basis to process under 152-ФЗ ст. 6 ч. 1 п. 5 (исполнение договора).

License

MIT © Grigorii Grachev. Free for any use, including commercial.

Roadmap

Next up: ОСАГО support — the underlying field model already accommodates it; only parser patterns need adding.

Past releases — see CHANGELOG.md. The design and implementation plan for the policyholder + contacts work shipped in 0.3.0 is preserved in docs/roadmap-policyholder.md for reference.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

polis_recognizer-0.3.3.tar.gz (108.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

polis_recognizer-0.3.3-py3-none-any.whl (110.3 kB view details)

Uploaded Python 3

File details

Details for the file polis_recognizer-0.3.3.tar.gz.

File metadata

  • Download URL: polis_recognizer-0.3.3.tar.gz
  • Upload date:
  • Size: 108.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for polis_recognizer-0.3.3.tar.gz
Algorithm Hash digest
SHA256 93f2aa10060ce3232acc1e01a242f73a70efa9b9cad5372bf74f634c124a9891
MD5 e4b24f95dbcb2e31dbc6b8c1b69c0abd
BLAKE2b-256 4c402b230f08800ee0c8a0bc7419eb521f8596086a4d3ba89031d50ddc756045

See more details on using hashes here.

Provenance

The following attestation bundles were made for polis_recognizer-0.3.3.tar.gz:

Publisher: publish.yml on grigra27/polis-recognizer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file polis_recognizer-0.3.3-py3-none-any.whl.

File metadata

File hashes

Hashes for polis_recognizer-0.3.3-py3-none-any.whl
Algorithm Hash digest
SHA256 100777a0044a76fc1d5283ab12eb32570746f01e545e4c7c5d3ffe70ec26b051
MD5 e35e05d044fade7c074c7d7b3f4f3033
BLAKE2b-256 e8aa4ca3954917810f41c81302baad128f35a66cf2b2fd7b578df2ea5556e531

See more details on using hashes here.

Provenance

The following attestation bundles were made for polis_recognizer-0.3.3-py3-none-any.whl:

Publisher: publish.yml on grigra27/polis-recognizer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page