Skip to main content

DocArmor (Document Intelligence Gateway)

Python 3.8+ License: MIT PyPI version

DocArmor (Document Intelligence Gateway) is a high-performance document validation, security scanning, quality guardrail, and exact token counting engine. Built in Rust with native Python bindings via PyO3, DocArmor sits between raw document ingestion and downstream LLM/RAG pipelines to prevent system exploitation, database bloat, and unexpected API costs.


Features • Installation • Quick Start • Python API • Telemetry Schema • Supported Formats • Examples • License


Features

  • 🧠 Pre-Ingestion Knowledge Base Engine (v0.2.0) - Converts raw PDFs, images, spreadsheets, and multi-file codebases into hyper-compressed, linked Knowledge Base Markdown (.md) with Table of Contents (TOC), executive summaries, and deep anchor links.
  • 📉 Multi-Model Token Reduction Telemetry - Achieves up to 60-90%+ token reduction before passing content to LLM agents across Claude 3.5/3.7, GPT-4o, Gemini 1.5/2.0, LLaMA 3, and DeepSeek R1/V3.
  • 🌐 "One Brain" Project Repository Ingestion - Recursively aggregates full multi-file codebases or directory trees into a single structured project Knowledge Base with file tree indexes and module breakdowns.
  • ✨ Real GPT Tokenization - Integrates high-performance tiktoken-rs in Rust to calculate exact GPT token budgets (not approximations) for models like GPT-4, GPT-3.5, Claude, or LLaMA.
  • ⚡ Multi-Format Support - Seamlessly extracts text and parses metadata from PDF, TXT, MD, DOCX, PPTX, XLSX, CSV, JSON, XML, HTML, and code files (.py, .rs, .go, .js, .ts, .java, .cpp, .c, .sh, .sql).
  • 🛡️ Ingestion Security - Built-in security scanners inspect compressed documents and file headers to intercept Zip bombs, compression bombs, and oversized resource limits before they reach system memory.
  • 🔍 Text Quality & OCR Necessity Detection - Evaluates page text density, whitespace-to-character ratio, and empty page signals to flag scanned/image-only documents (requires_ocr) before vector database embedding.
  • 🚀 Native Parallel Batch Processing - Utilizes Rust's concurrent work-stealing thread pool (Rayon) to process thousands of files or directory trees in parallel with zero GIL serialization.
  • 💾 Global De-duplication - Computes high-performance SHA-256 content hashes in parallel to identify and skip exact duplicate files inside a batch queue automatically.
  • 💰 Dynamic Cost Estimation - Estimates LLM input cost and vector database embedding cost dynamically before making external API requests.
  • 🎯 Intelligent Agent Routing - Classifies text based on heuristic token frequencies and assigns a target downstream AI Agent (e.g., LegalAgent, ProcurementAgent).
  • 🔒 Rust-Native PII Redaction & Data Masking - Detects and masks Personally Identifiable Information (PII) like emails, phone numbers, SSNs, IP addresses, and credit cards directly in Rust before data leaves your environment.

Installation

From PyPI (Recommended)

Install pre-compiled native binary wheels instantly on Windows, Linux, or macOS:

pip install docarmor

(No Rust compilers, C-libraries, or compilation tools are required on the host system).

From Source

git clone https://github.com/JIVTESH28/docarmor.git
cd docarmor
pip install .

Quick Start

Initialize the Analyzer

import json
import docarmor

# Initialize the gateway analyzer with custom thresholds
analyzer = docarmor.DocumentAnalyzer({
    "target_model": "gpt-4",                   # Target context window check
    "tokenizer_name": "cl100k_base",           # Tiktoken profile
    "embedding_rate_per_million": 0.02,        # Cost per 1M tokens ($)
    "llm_input_rate_per_million": 5.00,        # Cost per 1M tokens ($)
    "max_file_size": 52428800                  # Max file size (50MB)
})

Python API Usage

Single File Ingestion (Local Disk)

report_str = analyzer.analyze_file("contract.pdf")
report = json.loads(report_str)
print(f"Tokens: {report['token_count']} | RAG Ready: {report['rag_ready']}")

In-Memory Bytes Ingestion (API Uploads)

uploaded_bytes = b"Sample document text buffer."
report_str = analyzer.analyze_bytes(uploaded_bytes, "invoice.txt")
report = json.loads(report_str)
print(f"Domain Class: {report['document_class']} | RAG Ready: {report['rag_ready']}")

Natively Parallel Batch Processing

file_list = ["agreement.docx", "data.xlsx", "spec.pdf"]
batch_report_str = analyzer.analyze_batch(file_list)
batch_report = json.loads(batch_report_str)

print(f"Successful files: {batch_report['summary']['successful_files']}")
print(f"Duplicates skipped: {batch_report['summary']['duplicate_files']}")

Directory Ingestion (Recursive Scan)

dir_report_str = analyzer.analyze_directory("./archive", recursive=True)
dir_report = json.loads(dir_report_str)
print(f"Total directory tokens: {dir_report['summary']['total_tokens']}")

🧠 Knowledge Base Pre-Ingestion (.md) Conversion (New in v0.2.0)

Pre-ingests bloated PDFs, documents, images, or full code repositories and converts them into hyper-compressed, linked Knowledge Base Markdown documents with token savings telemetry:

# 1. Top-Level Convenience Helper (File, Directory, or Bytes)
kb_result = docarmor.to_knowledge_base("procurement_agreement.pdf", target_model="claude-3-5-sonnet")

print(kb_result["markdown"])
print(f"Token Reduction : {kb_result['telemetry']['reduction_percentage']}%")
print(f"Cost Savings    : ${kb_result['telemetry']['cost_savings_usd']}")

# 2. Multi-File Project Repository Ingestion ("One Brain")
project_kb = analyzer.convert_directory_to_kb("./my_project", recursive=True, target_model="gemini-1.5-pro")
proj_data = json.loads(project_kb)
print(f"Project Files: {proj_data['telemetry']['total_files']} | Savings: {proj_data['telemetry']['reduction_percentage']}%")

# 3. Hardware-Accelerated OCR to Knowledge Base Markdown
ocr_analyzer = docarmor.OcrDocumentAnalyzer()
ocr_kb_str = ocr_analyzer.convert_file_to_kb("scanned_invoice.png", target_model="gpt-4o")

Ultra-Fast Single-Metric Bypasses

If you only need a single metric and want to bypass the rest of the gateway analysis pipeline (such as security checks, cost estimation, and domain classification), use the sub-millisecond helpers:

# Raw metric count helpers (File-based)
word_count = analyzer.count_words("document.docx")
char_count = analyzer.count_chars("document.docx")
token_count = analyzer.count_tokens("document.docx")

# Raw metric count helpers (Byte-based)
token_count = analyzer.count_tokens_bytes(uploaded_bytes, "invoice.txt")

# Rust-Native PII Redaction & Data Masking
pii_text = "My email is test@example.com and phone is 123-456-7890."
# Redact all supported categories (email, phone, ssn, ip, credit_card)
redacted_all = analyzer.redact_pii(pii_text) # "My email is [EMAIL] and phone is [PHONE]."
# Or redact only specific categories
redacted_email = analyzer.redact_pii(pii_text, ["email"]) # "My email is [EMAIL] and phone is 123-456-7890."

Telemetry Output Schema

DocArmor generates a comprehensive, metadata-rich telemetry report for every analyzed file:

{
  "file_name": "contract_agreement.pdf",
  "file_type": "pdf",
  "sha256": "07c270b274dae324f906e0aa3a8d606471931e9c1afc241ddbc8f9ae52baffe7",
  "token_count": 2424,
  "word_count": 1612,
  "character_count": 11448,
  "page_count": 4,
  "requires_ocr": false,
  "quality_score": 0.8,
  "duplicate": false,
  "security_risk": "low",
  "fits_context": true,
  "rag_ready": true,
  "requires_summarization": false,
  "recommended_chunking": "semantic chunking",
  "document_class": "Legal",
  "recommended_agent": "LegalAgent",
  "contains_pii": true,
  "pii_categories_found": ["email", "phone"],
  "estimated_embedding_cost": 0.0,
  "estimated_llm_cost": 0.0121,
  "processing_time_ms": 12.34
}

Telemetry Field Descriptions

Field Type Description
file_name String Base name of the analyzed file.
file_type String Lowercase file extension (e.g. pdf, docx, txt).
sha256 String Cryptographic SHA-256 hash representing the exact content payload.
token_count Integer Exact token count matching the selected model tokenizer profile.
word_count Integer Number of words counted based on unicode whitespace dividers.
character_count Integer UTF-8 character length of the extracted document text.
page_count Integer Page count (e.g. PDF pages, PowerPoint slides, Excel sheets, estimated text lines).
requires_ocr Boolean Flags true if document has page structures but low text density (image-only scanned).
quality_score Float Cleanliness index (0.0 - 1.0) graded by density, metadata, ratio, and OCR markers.
duplicate Boolean Flags true if identical SHA-256 has already been processed in the concurrent batch queue.
security_risk String Security score (low, medium, high) validating Zip bombs and size thresholds.
fits_context Boolean Checks if token_count fits inside the target model's context window.
rag_ready Boolean Evaluates suitability for search databases (true if secure, non-scanned, and clean).
requires_summarization Boolean Recommends pre-summarizing if the token count or page density is excessively large.
recommended_chunking String Suggested chunking strategy (no chunking, fixed, semantic, hierarchical, agentic).
document_class String Classified topical domain (Finance, Procurement, Legal, HR, Tech Doc, Research, etc.).
recommended_agent String Recommended target downstream AI Agent target (e.g. LegalAgent).
contains_pii Boolean Flags true if document text contains common PII entities (email, phone, SSN, IP, credit card).
pii_categories_found List Names of PII categories found in the document (e.g., ["email", "phone"]).
estimated_embedding_cost Float Predicted vector database indexing cost.
estimated_llm_cost Float Predicted input processing cost.
processing_time_ms Float Internal Gateway execution latency in milliseconds.

Supported Formats

Format Extension Extraction Method Key Features
PDF .pdf Native lopdf Parser Structural reading, scanned detection, page extraction
Word .docx Native docx XML Parser Direct paragraph and table text extraction
PowerPoint .pptx Native pptx XML Parser Shape text, slide processing, bullet analysis
Excel .xlsx Calamine Engine Spreadsheet parsing, cell extraction, rows estimation
CSV .csv CSV Parser Direct row, column parsing, delimiter validation
Plain Text .txt, .md Unicode Parser Streaming flat extraction, lossy fallback encoding
JSON .json Serde JSON Recursive nested key-value string extraction
XML .xml Quick XML Parser Tag-stripped text, element-wise traversal
HTML .html Quick XML Parser Element parsing, script/style extraction filtering

Configuration Limits

Setting Default Value Purpose
target_model "gpt-4" Target context size limit check
tokenizer_name "cl100k_base" Tokenizer profile (cl100k_base, r50k_base, p50k_base)
max_file_size 52,428,800 bytes (50MB) Intercept oversized documents
embedding_rate_per_million $0.02 Custom embedding cost rate

Ingestion Pipeline Flow

graph TD
    File[Document Uploaded] --> Security[Security Scanner: check sizes, corruption, zip bombs]
    Security -->|High Risk| Block[Abort: return error / flag security_risk]
    Security -->|Safe| Parser[Select Parser based on Extension: PDF, Docx, Xlsx, etc.]
    Parser --> Quality[Quality Evaluator: calculate density, pages, readability]
    
    Quality -->|Text Empty / Scanned| OCR[Lazy-load OCR: GPU Accelerated MPS/CUDA]
    Quality -->|Readable Text| Metrics[Metrics Evaluator: TikToken Token Counting, PII detection]
    
    OCR --> Metrics
    Metrics --> Router[Heuristic Domain Classifier & Cost Estimator]
    Router --> JSON[Generate Telemetry JSON Report]

How the OCR Integration Works

DocArmor implements a high-performance hybrid OCR gateway under the OcrDocumentAnalyzer class:

  1. Rust-Native Gatekeeping: When a file is submitted, DocArmor first uses its sub-millisecond Rust parsers to check the file type and structure.
    • If the document is a clean digital file (e.g., text PDF, Word doc, or markdown), the text is extracted instantly, and the heavy OCR engine is completely bypassed.
    • If the file is an image (.png, .jpg, .jpeg, etc.) or is flagged by the Rust quality scanner as a scanned/text-empty PDF (requires_ocr: True), the OCR engine is initialized.
  2. Lazy Loading: To keep package imports sub-millisecond, PyTorch and EasyOCR model weights are loaded lazily on-demand only when the first scanned document or raw image is encountered.
  3. Hardware Auto-Detection: The engine dynamically autodetects your host hardware to run deep learning models at maximum speed:
    • macOS (Apple Silicon): Natively offloads tensor computations to the GPU via Metal Performance Shaders (MPS).
    • Windows/Linux with GPU: Automatically targets your Nvidia GPU via CUDA.
    • Fallback: Runs on optimized multi-threaded CPU.
  4. Rust Telemetry Reconciliation: Once text is extracted via OCR, the raw text bytes are passed back into DocArmor's Rust core using a virtual text buffer. The Rust engine then computes exact GPT token budgets (tiktoken-rs), counts words/characters, runs domain classification, and generates cost estimations—reconciling all statistics back into a single unified JSON schema.

Examples

Example 1: RAG Ingestion Security & Quality Gatekeeper

Ensure that only secure, high-quality, digital documents enter your vector database:

import json
import docarmor

analyzer = docarmor.DocumentAnalyzer()
report = json.loads(analyzer.analyze_file("user_upload.pdf"))

# Intercept risks at the gateway
if report["security_risk"] == "high":
    raise ValueError(f"CRITICAL: Security exception triggered for {report['file_name']}")

if report["requires_ocr"]:
    print(f"Routing {report['file_name']} to hardware-accelerated OCR pipeline.")
elif not report["rag_ready"]:
    print(f"Skipping {report['file_name']} due to low text quality score: {report['quality_score']}")
else:
    print(f"Ingesting clean document text. Context Size: {report['token_count']} tokens.")

Example 2: API Cost Budgeting & Model Window Check

Calculate API transaction costs and verify if a document fits within a model's context window:

import json
import docarmor

analyzer = docarmor.DocumentAnalyzer({
    "target_model": "gpt-3.5-turbo",
    "llm_input_rate_per_million": 1.50
})

report = json.loads(analyzer.analyze_file("long_transcript.txt"))

if not report["fits_context"]:
    print(f"Document exceeds target context window. Recommended chunking strategy: {report['recommended_chunking']}")
else:
    print(f"Document fits. Estimated processing cost: ${report['estimated_llm_cost']:.4f}")

Example 3: Hardware-Accelerated OCR Integration (Metal/CUDA)

Incorporate unified OCR for scanned files directly from the installed package:

import json
from docarmor import OcrDocumentAnalyzer

# Initialize unified OcrDocumentAnalyzer (auto-routes to Apple Metal MPS or CUDA)
gateway = OcrDocumentAnalyzer()

report_json = gateway.analyze_file("scanned_receipt.jpg")
report = json.loads(report_json)

print(f"OCR Text: {report['text']}")
print(f"OCR Tokens: {report['token_count']} | RAG Ready: {report['rag_ready']}")

License

This project is licensed under the MIT License.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

docarmor-0.2.0.tar.gz (71.3 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

docarmor-0.2.0-cp38-abi3-win_amd64.whl (3.0 MB view details)

Uploaded CPython 3.8+Windows x86-64

docarmor-0.2.0-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (3.3 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.17+ x86-64

docarmor-0.2.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (3.3 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.17+ ARM64

docarmor-0.2.0-cp38-abi3-macosx_11_0_arm64.whl (3.1 MB view details)

Uploaded CPython 3.8+macOS 11.0+ ARM64

docarmor-0.2.0-cp38-abi3-macosx_10_12_x86_64.whl (3.2 MB view details)

Uploaded CPython 3.8+macOS 10.12+ x86-64

File details

Details for the file docarmor-0.2.0.tar.gz.

File metadata

  • Download URL: docarmor-0.2.0.tar.gz
  • Upload date:
  • Size: 71.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for docarmor-0.2.0.tar.gz
Algorithm Hash digest
SHA256 96489637f4facff73e937e6099764895b7655a1b78b794fa58dee774b77946e1
MD5 3aa68bb144f63d22e9540b956cc07620
BLAKE2b-256 f37d596fe8c43508da081002eac4d0233784f70c9bfadfab4ecce9e41fdfd7f5

See more details on using hashes here.

Provenance

The following attestation bundles were made for docarmor-0.2.0.tar.gz:

Publisher: pypi.yml on JIVTESH28/docarmor

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file docarmor-0.2.0-cp38-abi3-win_amd64.whl.

File metadata

  • Download URL: docarmor-0.2.0-cp38-abi3-win_amd64.whl
  • Upload date:
  • Size: 3.0 MB
  • Tags: CPython 3.8+, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for docarmor-0.2.0-cp38-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 e98e21b40b97e1b996f6b1eb339fcbfc53c11f5e8b79a210d8ba1ee5c915b8c2
MD5 7321e5fd3673286671321c8bc5adac31
BLAKE2b-256 3cd86c44023e5157d6ad558d132c982bc7b6dc762fe05519aea1cf0c6a0859e3

See more details on using hashes here.

Provenance

The following attestation bundles were made for docarmor-0.2.0-cp38-abi3-win_amd64.whl:

Publisher: pypi.yml on JIVTESH28/docarmor

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file docarmor-0.2.0-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for docarmor-0.2.0-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 074da59fc7dc36370fa2a64b3be0e8d3ebe5d01d8fa0f2107af261fb2bdcce74
MD5 3154488a8bd98e5f06578ba3e55ee663
BLAKE2b-256 c81fefc2e74be317dcfafec7fba4124b9c3b4e0d95767dfbd49f6e068e6f2809

See more details on using hashes here.

Provenance

The following attestation bundles were made for docarmor-0.2.0-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: pypi.yml on JIVTESH28/docarmor

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file docarmor-0.2.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for docarmor-0.2.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 1950640526b97c349b54d977f21e4baa48ec8bdce1ab8e31782fa7667ff6bdaf
MD5 78b0d8d05ee19bee22eda8400df971f3
BLAKE2b-256 6205ecd5de9a5df6e1f2c5c4e91c4fff19db413bb90c39a1c9bfe280a2c9cbc2

See more details on using hashes here.

Provenance

The following attestation bundles were made for docarmor-0.2.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: pypi.yml on JIVTESH28/docarmor

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file docarmor-0.2.0-cp38-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for docarmor-0.2.0-cp38-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 3d2fa7a5fdf3205b3bd99534ff31650663abb8830fd04eaef13ceee47580b90a
MD5 55bc962a10c06aa28ee0ff1dbfb83d66
BLAKE2b-256 848579314925e7d10ca6382357ab492c68ee317e21ebde10db2e9c88fc612b34

See more details on using hashes here.

Provenance

The following attestation bundles were made for docarmor-0.2.0-cp38-abi3-macosx_11_0_arm64.whl:

Publisher: pypi.yml on JIVTESH28/docarmor

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file docarmor-0.2.0-cp38-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for docarmor-0.2.0-cp38-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 291be1d7875e1627eeb9ab51899f47dd7af4b6052012a471b75c466757089780
MD5 7d60a7a932687dd8aeedaaf1c2779410
BLAKE2b-256 4890be6b46d1d790e2b200255e3e1602c7e56b8eafc13b8d49d7a6b574f37b94

See more details on using hashes here.

Provenance

The following attestation bundles were made for docarmor-0.2.0-cp38-abi3-macosx_10_12_x86_64.whl:

Publisher: pypi.yml on JIVTESH28/docarmor

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.2.1

6 files

This release

0.2.0 This release

6 files

0.1.14

6 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page