Skip to main content

xml2table

A small, predictable Python SDK for converting XML and PDF documents to CSV, Excel, XML, and plain text.

  • XML → CSV / Excel. One core recursive flattening function turns nested XML into flat rows; CSV and Excel are thin writers on top of it. No data loss by default: repeated elements can be joined into one cell, exploded into extra rows, or spread across indexed columns — you choose.
  • PDF → XML. A structure-preserving extractor (pdfplumber-backed) groups a PDF's words into paragraphs and detects tables, keeping both in their original top-to-bottom reading order — no data lost, no tables flattened into loose words.
  • PDF → text. A fast, lightweight raw-text extractor (pypdf-backed) that doesn't need table detection at all — just the PDF's text, in reading order, with its visual whitespace layout preserved by default.
  • Small dependency footprint. Only openpyxl is required for the XML-to-table side. pandas (DataFrames), pypdf (PDF text), and pdfplumber (PDF XML/tables) are all optional extras — importing xml2table never requires any of them.
  • Three ways in, for each conversion: one-line functions, a reusable converter object, or the xml2table CLI.

Install

pip install -e .             # from a checkout: XML -> CSV/Excel only
pip install -e ".[pandas]"   # + optional DataFrame support
pip install -e ".[pdf-text]" # + PDF -> text (pypdf only, lightweight)
pip install -e ".[pdf]"      # + PDF -> XML and text (pdfplumber + pypdf)

Quick start

from xml2table import xml_to_csv, xml_to_excel

xml_to_csv("orders.xml", "orders.csv")
xml_to_excel("orders.xml", "orders.xlsx")

Given:

<orders>
  <order id="1">
    <customer><name>Jane Doe</name></customer>
    <total>99.99</total>
  </order>
</orders>

you get one row per <order>:

@id customer.name total
1 Jane Doe 99.99

Nested elements are flattened with .-separated keys, attributes get an @ prefix, and the record element (order here) is auto-detected as "the repeated child of the root". When detection is ambiguous, pass record_path explicitly.

Reusable converter

Parse once, write many times:

from xml2table import XMLConverter, FlattenOptions

converter = XMLConverter("orders.xml", options=FlattenOptions(record_path="Order"))
converter.to_csv("orders.csv")
converter.to_excel("orders.xlsx", sheet_name="Orders")
rows = converter.to_records()      # list[dict]
df = converter.to_dataframe()      # requires pandas

Handling repeated elements (array_mode)

Given an order with two line items, FlattenOptions.array_mode controls the shape of the output:

mode Result Use when
"join" (default) One row per order; items collapsed into one joined cell You just want a quick, human-readable table
"explode" One row per item (order fields repeat) You want a normalized, analysis-ready table, like a SQL join
"index" One row per order; items spread into items.item.0.*, items.item.1.*, ... You need every field as its own column with no row duplication
from xml2table import FlattenOptions, xml_to_records

xml_to_records("orders.xml", options=FlattenOptions(array_mode="explode"))

Multi-sheet Excel from one document

Turn different parts of the same XML document into separate, related sheets (e.g. an "orders" table and an "items" table):

from xml2table import xml_to_excel

xml_to_excel(
    "orders.xml", "orders.xlsx",
    sheets={"orders": "Order", "items": ".//Order/Items/Item"},
)

Record paths

record_path uses the same syntax as Element.findall (a subset of XPath): "Order", "Orders/Order", ".//Item", "Order[@status='shipped']", etc.

CLI (XML)

xml2table csv orders.xml orders.csv --record-path Order --array-mode explode
xml2table excel orders.xml orders.xlsx --sheet-name Orders
xml2table excel orders.xml orders.xlsx --sheet orders=Order --sheet items=".//Item"

Run xml2table csv --help or xml2table excel --help for all flags (--separator, --attribute-prefix, --no-attributes, --keep-namespaces, --delimiter, --encoding, ...).

PDF → XML / text

pdf_to_xml and pdf_to_text are two independent, purpose-built backends behind one options object and one CLI command:

  • pdf_to_xml (pdfplumber) groups the PDF's words into paragraphs (by line, then by vertical gap) and detects tables separately via ruling lines or text alignment, then places both back in the order they appear on the page. Nothing is dropped, and table text never bleeds into surrounding paragraphs.
  • pdf_to_text (pypdf) is a much lighter path: it just extracts each page's text in reading order, preserving the PDF's visual whitespace layout by default (so simple tables and columns still read naturally) without doing any table detection.
from xml2table import pdf_to_xml, pdf_to_text

pdf_to_xml("invoice.pdf", "invoice.xml")    # paragraphs + tables (needs xml2table[pdf])
pdf_to_text("invoice.pdf", "invoice.txt")   # fast raw text (needs xml2table[pdf-text])

invoice.xml looks like:

<document source="invoice.pdf" pages="1">
  <page number="1" width="595.28" height="841.89">
    <paragraph bbox="42.83,43.13,159.90,61.13">Invoice #1024</paragraph>
    <table bbox="40.00,220.00,540.00,316.00" rows="4" cols="3">
      <row><cell>Item</cell><cell>Qty</cell><cell>Price</cell></row>
      <row><cell>Widget</cell><cell>2</cell><cell>$10.00</cell></row>
      ...
    </table>
  </page>
</document>

invoice.txt is pypdf's layout-preserving text, e.g.:

Invoice #1024

Bill To: Jane Doe
123 Example Street
Springfield, USA

Thank you for your business. Payment is due within thirty days...

Item                                                          Qty        Price
Widget                                                        2          $10.00
...

Reuse one PDFConverter for both (it lazily parses with pdfplumber only if you call to_xml()/pages, and always uses pypdf for to_text()):

from xml2table import PDFConverter

converter = PDFConverter("invoice.pdf")
converter.to_xml("invoice.xml")
converter.to_text("invoice.txt")
for page in converter.pages:
    print(page.number, len(page.paragraphs), len(page.tables))

PDFOptions reference

Option Default Used by Description
line_tolerance 3.0 XML Max vertical gap (points) for words to count as the same line
paragraph_gap 6.0 XML Min vertical gap (points) between lines that starts a new paragraph
table_settings None XML Passed through to pdfplumber's find_tables() for unusual tables (e.g. borderless). Caution: this applies to the whole page, not just the table — see Validation below before using it on documents with narrative text.
cell_na "" XML String used for empty/missing table cells
page_separator "\n----- Page {page} -----\n" text Inserted between pages
keep_layout True text Preserve the PDF's whitespace layout (pypdf "layout" mode) vs. plain, whitespace-normalized text

CLI (PDF)

xml2table pdf invoice.pdf invoice.xml --to xml
xml2table pdf invoice.pdf invoice.txt --to text

Run xml2table pdf --help for all flags (--line-tolerance, --paragraph-gap, --cell-na, --page-separator, --no-layout).

Validation: tested on 1,200 financial-report PDFs

pdf_to_xml and pdf_to_text were benchmarked against 1,200 generated financial-report PDFs (balance sheets, income statements, cash-flow statements, MD&A-style narrative text, footnotes — real financial formatting: $1,234,567, (123,456) negatives, N/A blanks, multi-page, ruled and borderless tables) with exact ground truth for every paragraph and table cell, so the numbers below are measured, not estimated. Full methodology, the generator, and raw per-document results are in benchmarks/RESULTS.md.

Metric Result
Documents converted without error 1,200 / 1,200 (100%)
Paragraph text fidelity (XML) 100.000%
Ruled-table shape + cell fidelity 100.000%
Borderless-table shape detection 0.000% (documented limitation — see below)
Raw-text content recall (pdf_to_text) 100.000%, including for borderless tables
Throughput 20.4 PDFs/sec (pdf_to_xml), 184.6 PDFs/sec (pdf_to_text)

The one real limitation, quantified: pdfplumber's default table finder needs ruling lines, so it doesn't detect borderless (text-only-aligned) tables — but nothing is lost when it doesn't: the un-detected table's text still comes through as ordinary paragraph text (pdf_to_text recall stays at 100%). We also tested the obvious "fix" — pdfplumber's table_settings override for borderless tables — across all 1,200 documents, and it made things worse: because the override applies to the whole page, it started misreading ordinary paragraph sentences as table cells, corrupting paragraph output in 97.9% of documents (paragraph fidelity dropped from 100% to 13.9%). We did not ship that as a recommended workaround; see PDFOptions.table_settings's docstring and benchmarks/RESULTS.md for the full numbers and why.

We could not download real financial filings for this test — this sandboxed session's network policy blocks direct access to sites like sec.gov — so the corpus is synthetic but built to real financial-statement conventions specifically so every value has a known-correct answer to grade against. benchmarks/RESULTS.md explains this in more detail and gives the exact commands to reproduce or extend the benchmark (e.g. against real filings, on a machine with broader network access).

FlattenOptions reference

Option Default Description
record_path None (auto-detect) Path to the repeated record element
attribute_prefix "@" Prefix for attribute-derived columns
text_key "#text" Key for an element's own text when it also has attributes/children
separator "." Separator for nested key paths
array_mode "join" "join", "explode", or "index"
join_separator "; " Separator used by "join" mode
include_attributes True Include XML attributes as columns
strip_namespaces True Strip {namespace} from tag/attribute names
encoding "utf-8" Text encoding for reads/writes

Errors

All exceptions inherit from xml2table.XMLConversionError:

  • XMLParseError — malformed XML input
  • RecordPathNotFoundError — record_path matched nothing, or automatic record detection was ambiguous (the error message tells you what to pass)
  • PDFExtractionError — a PDF file could not be opened or parsed
  • MissingOptionalDependencyError — e.g. calling to_dataframe() without pandas, pdf_to_xml() without pdfplumber, or pdf_to_text() without pypdf, installed

Development

pip install -e ".[dev]"   # includes pandas, pdfplumber, pypdf, and fpdf2 (for regenerating PDF fixtures)
pytest
python examples/quickstart.py

Project layout:

src/xml2table/
  parser.py          # XML -> list[dict] flattening engine
  options.py         # FlattenOptions
  converter.py       # XMLConverter + module-level convenience functions
  writers/           # CSV and Excel output
  pdf_extract.py     # pdfplumber: PDF -> Document(pages of Paragraph/Table), in reading order
  pdf_xml_writer.py  # Document -> XML
  pdf_text_extract.py # pypdf: PDF -> plain text, independent of pdf_extract.py
  pdf_options.py     # PDFOptions
  pdf_converter.py   # PDFConverter + pdf_to_xml/pdf_to_text
  cli.py             # `xml2table` command-line tool
tests/
  fixtures/          # sample XML documents and PDF fixtures (fixtures/pdf/)
  test_*.py
examples/
  quickstart.py

Metadata

Release files for xml2table 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for xml2table 0.1.0
File Size Uploaded
xml2table-0.1.0.tar.gz 32.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for xml2table 0.1.0
File Interpreter ABI Platform
xml2table-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 60.2 kB

Release files / xml2table-0.1.0.tar.gz

Download URL xml2table-0.1.0.tar.gz
Size 32.2 kB
Tags Source
SHA-256 checksum
How to use checksums
2163d1feb9459a3932e8685ffe9e65d73e02d6c62a3a42d87a81e2d9e9cd9df8
BLAKE2b-256 checksum
How to use checksums
c427e8c7e39d6e3af12a8abd7fee4017e95c7cbcb8e9a56dead83ce0555744a1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release files / xml2table-0.1.0-py3-none-any.whl

Download URL xml2table-0.1.0-py3-none-any.whl
Size 28.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3637c6b0d4b5869a9d22c7e5ebbe6be311c16c2a18d507caf1266c4477e9e7a9
BLAKE2b-256 checksum
How to use checksums
7ef005301aedb67f2e4bc0a3f2be1549bd260e9c5e02a381ad9d7b9b01835e0b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page