Skip to main content

Alt text PyPI version

This is the package implementation of the code for the kinship ties and relational information from large text corpora, such as genealogies, biographies, and historical dictionaries.

Important information:

  • For example package usage, see test_workflow.ipynb
  • For information about individual functions before the package implementation and release, see examples folder.

Expected input files

Each document should normally have these three files together in the same folder under documents_root:

0_GenelogiesAndBiographies/
├── Example genealogy/
│   ├── Example genealogy_mod.htm       # source HTML used for the inventory
│   ├── Example genealogy_mod.pdf       # PDF with a readable OCR text layer
│   └── Example genealogy_Original.pdf  # original PDF without OCR
└── Biographies/
    └── Example biography/
        ├── Example biography_mod.htm
        ├── Example biography_mod.pdf
        └── Example biography_Original.pdf

The filenames must share the same base name. The package uses the HTML for inventory creation, prefers *_mod.pdf for page matching, and falls back to *_Original.pdf, which can be OCRed with Tesseract when necessary.

Current release:

Current release includes inventory creation, setting text bounds, assigning font usage and selecting appropriate text chunks from the files (including restriction to "MainText" fonts and JumPJumP insertion).

Files structure:

  1. cleaner.py

Contains the HTMLCleaner class and thus the main logic.

  1. utils.py

Contains helper functions (like JumPJumP insertion)

  1. io.py

Contains the file processing logic, such as loading and saving the JSON inventory.

  1. page_matching.py

Adds one-based PDF start_page and end_page values to chunk metadata using exact and fuzzy text matching. Missing PDFs are skipped with null page values and explicit filename guidance.

Add PDF page numbers

If you already created an HTMLCleaner, use cleaner.add_pdf_pages(). If you only have an inventory JSON file, use add_pages_to_inventory_file() directly.

summary = cleaner.add_pdf_pages(
    documents_root="../0_GenelogiesAndBiographies",
    output_path="exact_html_inventory_new_ids_cleaned_pages.json",
)

The default fuzzy threshold is 0.80. The matcher prefers *_mod.pdf, falls back to _Original.pdf when needed, and can use Tesseract for image-only originals when Tesseract is installed.

Human Labeler Allocation

When text chunks are allocated to labelers, each generated TXT filename indicates its task type:

  • Task_0 contains chunks shared across all labelers.
  • Task_1 through Task_9 contain labeler-specific, unshared chunks.

New section 4 assessment workflow

The reusable assessment follows sections 4.2.2.1 and 4.2.2.2 of 240426 Information extraction system 260801.ipynb. Install the optional plot dependencies, load the processed human and AI tables, and then create the assessment tables before drawing anything:

pip install -e ".[plots]"
from text2rel import (
    ReliabilityAssessment,
    build_sweep_report,
    plot_sweep_report,
    plot_threshold_curves,
)

assessment = ReliabilityAssessment("inventory.json")
assessment.load_relations_data(relations_csv="human_relations.csv")
assessment.load_chatgpt_relations("machine_relations.csv")

# one_to_one is the package's stricter default for headline metrics.
relation_results = assessment.assess_llm_against_humans(data="relations")
plot_threshold_curves(relation_results["threshold_metrics"], direction="one_to_one")

# directional_best reproduces the notebook's two reusable-candidate views.
diagnostics = assessment.assess_llm_against_humans(
    data="relations", matching_mode="directional_best"
)
fn_report = build_sweep_report(
    diagnostics["human_to_llm_best_pairs"], data="relations"
)
plot_sweep_report(fn_report)

assess_llm_against_humans() returns the selected pairs, threshold metrics, false-negative and false-positive review tables, structural exclusions, and configuration metadata. Plot functions only render these returned tables; they do not change matching or scores. See the separate New section 4 assessment plots and tables section at the end of test_workflow.ipynb for relations and events, notebook-style titles, confidence/length diagnostics, and review summaries.

Labelbox event construction modes

labelbox_events_to_dataframe(..., construction_mode="graph") is the default. It follows section 2.4.3.2 of the 260801 notebook: connected annotations are rebuilt around one Verb, safe reversed arrows are corrected, and incomplete or ambiguous structures remain auditable. Use construction_mode="legacy_simple" only when explicitly reproducing the package's earlier permissive conversion. Relation conversion uses its audited endpoint joins separately.

Release files for text2rel 1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for text2rel 1.0
File Size Uploaded
text2rel-1.0.tar.gz 150.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for text2rel 1.0
File Interpreter ABI Platform
text2rel-1.0-py3-none-any.whl Python 3 none any Details

Total release size: 308.4 kB

Release files / text2rel-1.0.tar.gz

Download URL text2rel-1.0.tar.gz
Size 150.3 kB
Tags Source
SHA-256 checksum
How to use checksums
89174bda2402a77dc25c1787e753e705dd81348a60ba8e20348ce692dc108e80
BLAKE2b-256 checksum
How to use checksums
49b0657295e620df10dbdd4b89dbda48e5e3eb1cb1ceb48d8ca6d78ed7f81def
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.15

Release files / text2rel-1.0-py3-none-any.whl

Download URL text2rel-1.0-py3-none-any.whl
Size 158.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
df257f8444b8c4efc803ca823fc36094af142d523b015854c3036c517fae095e
BLAKE2b-256 checksum
How to use checksums
4820413102ca59cef6a88bf7465713d25246cb75f2065d09e9940cfd54205abf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.15

Release history Release notifications | RSS feed

1.0.1

2 release files

This release

1.0 This release

2 release files

0.5

2 release files

0.2

2 release files

0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page