Skip to main content

Extracts HTML tables and converts them to Markdown, preserving structure via row/colspan handling.

Project description

html-table-rescuer

A robust Python tool to extract complex HTML tables and convert them into clean Markdown, JSON, or CSV formats.

Unlike simple formatters, this library features a Grid Logic Solver that correctly interprets rowspan and colspan attributes, normalizing complex HTML grids into perfectly aligned data structures.

How is this different from other table2md packages?

There is an existing package called table2md on PyPI.

  • The existing package is a formatter. You feed it Python lists/dicts, and it draws a Markdown table.

  • This package is an extractor and parser. It takes raw HTML source code, uses BeautifulSoup to parse the tags, mathematically resolves complex cell spans (rowspans/colspans), and builds an internal representation before exporting to Markdown.

The Problem:

Most parsers turn this <td rowspan="2"> into a misaligned mess:

| Header | Value |
|---|---|
| Spanned | Row 1 |
| Row 2 | |   <-- Everything shifts!

The Solution:

html-table-rescuer uses a grid solver to correctly normalize the matrix:

| Header | Value |
|---|---|
| Spanned | Row 1 |
| dito (Spanned) | Row 2 |

Architecture

Our pipeline ensures that complex HTML structures are safely converted without data loss or misalignment:

graph TD
    n1["HTML Input"] --> n2["BeautifulSoup Parser"]
    n2 --> n3["Grid Logic<br>(Rowspan/Colspan Solver)"]
    n3 --> n4["ParsedTable Data Object"]
    n4 --> n5["Markdown Export"]
    n4 --> n6["JSON/CSV Export"]
    n4 --> n7["LangChain/LlamaIndex/Haystack Wrappers"]

Installation

pip install html-table-rescuer

# With framework integrations
pip install "html-table-rescuer[langchain]"
pip install "html-table-rescuer[llamaindex]"
pip install "html-table-rescuer[haystack]"

Quick Start

from html_table_rescuer import TableParser

html_content = """
<table border="1">
  <tr>
    <th colspan="2">Header</th>
  </tr>
  <tr>
    <td>Data 1</td>
    <td>Data 2</td>
  </tr>
</table>
"""

# Initialize parser with your HTML
parser = TableParser(html_content)

# Parse all tables in the HTML
tables = parser.parse()

if tables:
    table = tables[0]
    
    # Export to Markdown (perfect for LLM context windows)
    print(table.to_markdown())
    
    # Export to JSON
    # print(table.to_json())
    
    # Export to CSV
    # print(table.to_csv())

Command Line

The package installs an html-table-rescuer command that reads from a file, a URL, or stdin:

# From a file
html-table-rescuer page.html

# From a URL
html-table-rescuer https://example.com/page.html

# From stdin (pipe or '-')
curl -s https://example.com/page.html | html-table-rescuer
cat page.html | html-table-rescuer - --format json

Options:

Option Description
--format, -f Output format: markdown (default), json, csv
--strategy, -s Rowspan fill strategy: fill_dito (default), repeat, empty
--dito-prefix Prefix used by the fill_dito strategy
--table, -t Extract only the table with this index (0-based)
--output, -o Write to a file instead of stdout; with CSV and multiple tables, writes name_1.csv, name_2.csv, …
--no-links / --no-bold / --no-italic Strip the respective inline formatting
--parser BeautifulSoup backend (lxml default, or html.parser)

AI Framework Integrations

Every table becomes its own document, so retrieval never splits a table in half.

LangChain

from html_table_rescuer.integrations.langchain import HTMLTableRescuerLoader

docs = HTMLTableRescuerLoader("page.html").load()
print(docs[0].page_content)   # Markdown table
print(docs[0].metadata)       # {'source': 'page.html', 'table_index': 0, 'parser': 'html_table_rescuer'}

LlamaIndex

from html_table_rescuer.integrations.llamaindex import HTMLTableRescuerReader

docs = HTMLTableRescuerReader().load_data("page.html")

Works as a file_extractor in SimpleDirectoryReader, so entire folders are handled for you:

from llama_index.core import SimpleDirectoryReader

reader = SimpleDirectoryReader(
    "./docs",
    file_extractor={".html": HTMLTableRescuerReader()},
)
docs = reader.load_data()

Haystack

from html_table_rescuer.integrations.haystack import HTMLTableRescuerConverter

converter = HTMLTableRescuerConverter()
docs = converter.run(sources=["page.html"])["documents"]

Drops straight into a pipeline — unreadable sources are skipped with a warning rather than failing the run, and ParseConfig survives pipeline serialization:

from haystack import Pipeline

pipe = Pipeline()
pipe.add_component("converter", HTMLTableRescuerConverter())
result = pipe.run({"converter": {"sources": ["page.html"]}})

All three accept a ParseConfig to control rowspan handling and inline formatting:

from html_table_rescuer import ParseConfig, RowspanStrategy

config = ParseConfig(rowspan_strategy=RowspanStrategy.REPEAT_VALUE)
docs = HTMLTableRescuerReader(config=config).load_data("page.html")

Features

  • HTML parsing via BeautifulSoup
  • Recursive inline-tag formatting (keeps links, bold, and italic tags alive even if nested in divs)
  • Complex rowspan and colspan grid resolution (using flexible strategies like filling cells with "dito" to preserve context for LLMs)
  • Clean Markdown export
  • Data Exports: JSON and CSV serialization from the ParsedTable object
  • CLI: html-table-rescuer command with file/URL/stdin input and Markdown/JSON/CSV output
  • AI Integrations: Ready-to-use LangChain Document Loader, LlamaIndex Reader (works with SimpleDirectoryReader), and Haystack Converter
  • Robust against broken real-world HTML: invalid colspan/rowspan values, HTML comments, and oversized spans are handled gracefully

🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

This project is licensed under the MIT License.

Blog-Posts:

** Read the 1. blog:** How to Stop LLMs from Hallucinating on Complex HTML Tables ** Read the 2. blog:** How to Stop LLMs from Hallucinating on Complex HTML Tables

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

html_table_rescuer-0.3.0.tar.gz (22.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

html_table_rescuer-0.3.0-py3-none-any.whl (16.3 kB view details)

Uploaded Python 3

File details

Details for the file html_table_rescuer-0.3.0.tar.gz.

File metadata

  • Download URL: html_table_rescuer-0.3.0.tar.gz
  • Upload date:
  • Size: 22.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for html_table_rescuer-0.3.0.tar.gz
Algorithm Hash digest
SHA256 275ace483c902e6d1ce3160a3521fdf2a00f94576504735bfa62e33e02e730d4
MD5 522c991df868f399dd75c44b964644f9
BLAKE2b-256 981977f8560226f7cb8eb5df4c97ed3f65bac7a4bb5d01c0a3b52ae57836eb6d

See more details on using hashes here.

Provenance

The following attestation bundles were made for html_table_rescuer-0.3.0.tar.gz:

Publisher: publish.yml on Encephos/html-table-rescuer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file html_table_rescuer-0.3.0-py3-none-any.whl.

File metadata

File hashes

Hashes for html_table_rescuer-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 18d53686f9b5f028e6cee29dfbad717547521ba14b3be2573b6235472adb2205
MD5 421f55441ced37ead8254bbec844c403
BLAKE2b-256 665adc831d7e303ce5d1816194defc2a6cbf1b29602837e434bc4ae9ba5f2e11

See more details on using hashes here.

Provenance

The following attestation bundles were made for html_table_rescuer-0.3.0-py3-none-any.whl:

Publisher: publish.yml on Encephos/html-table-rescuer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page