xml2table
A small, predictable Python SDK for converting XML and PDF documents to CSV, Excel, XML, and plain text.
- XML → CSV / Excel. One core recursive flattening function turns nested XML into flat rows; CSV and Excel are thin writers on top of it. No data loss by default: repeated elements can be joined into one cell, exploded into extra rows, or spread across indexed columns — you choose.
- PDF → XML. A structure-preserving extractor (
pdfplumber-backed) groups a PDF's words into paragraphs and detects tables, keeping both in their original top-to-bottom reading order — no data lost, no tables flattened into loose words. - PDF → text. A fast, lightweight raw-text extractor
(
pypdf-backed) that doesn't need table detection at all — just the PDF's text, in reading order, with its visual whitespace layout preserved by default. - Small dependency footprint. Only
openpyxlis required for the XML-to-table side.pandas(DataFrames),pypdf(PDF text), andpdfplumber(PDF XML/tables) are all optional extras — importingxml2tablenever requires any of them. - Three ways in, for each conversion: one-line functions, a reusable
converter object, or the
xml2tableCLI.
Install
pip install -e . # from a checkout: XML -> CSV/Excel only
pip install -e ".[pandas]" # + optional DataFrame support
pip install -e ".[pdf-text]" # + PDF -> text (pypdf only, lightweight)
pip install -e ".[pdf]" # + PDF -> XML and text (pdfplumber + pypdf)
Quick start
from xml2table import xml_to_csv, xml_to_excel
xml_to_csv("orders.xml", "orders.csv")
xml_to_excel("orders.xml", "orders.xlsx")
Given:
<orders>
<order id="1">
<customer><name>Jane Doe</name></customer>
<total>99.99</total>
</order>
</orders>
you get one row per <order>:
| @id | customer.name | total |
|---|---|---|
| 1 | Jane Doe | 99.99 |
Nested elements are flattened with .-separated keys, attributes get an
@ prefix, and the record element (order here) is auto-detected as "the
repeated child of the root". When detection is ambiguous, pass
record_path explicitly.
Reusable converter
Parse once, write many times:
from xml2table import XMLConverter, FlattenOptions
converter = XMLConverter("orders.xml", options=FlattenOptions(record_path="Order"))
converter.to_csv("orders.csv")
converter.to_excel("orders.xlsx", sheet_name="Orders")
rows = converter.to_records() # list[dict]
df = converter.to_dataframe() # requires pandas
Handling repeated elements (array_mode)
Given an order with two line items, FlattenOptions.array_mode controls the
shape of the output:
| mode | Result | Use when |
|---|---|---|
"join" (default) |
One row per order; items collapsed into one joined cell | You just want a quick, human-readable table |
"explode" |
One row per item (order fields repeat) | You want a normalized, analysis-ready table, like a SQL join |
"index" |
One row per order; items spread into items.item.0.*, items.item.1.*, ... |
You need every field as its own column with no row duplication |
from xml2table import FlattenOptions, xml_to_records
xml_to_records("orders.xml", options=FlattenOptions(array_mode="explode"))
Multi-sheet Excel from one document
Turn different parts of the same XML document into separate, related sheets (e.g. an "orders" table and an "items" table):
from xml2table import xml_to_excel
xml_to_excel(
"orders.xml", "orders.xlsx",
sheets={"orders": "Order", "items": ".//Order/Items/Item"},
)
Record paths
record_path uses the same syntax as
Element.findall
(a subset of XPath): "Order", "Orders/Order", ".//Item",
"Order[@status='shipped']", etc.
CLI (XML)
xml2table csv orders.xml orders.csv --record-path Order --array-mode explode
xml2table excel orders.xml orders.xlsx --sheet-name Orders
xml2table excel orders.xml orders.xlsx --sheet orders=Order --sheet items=".//Item"
Run xml2table csv --help or xml2table excel --help for all flags
(--separator, --attribute-prefix, --no-attributes, --keep-namespaces,
--delimiter, --encoding, ...).
PDF → XML / text
pdf_to_xml and pdf_to_text are two independent, purpose-built backends
behind one options object and one CLI command:
pdf_to_xml(pdfplumber) groups the PDF's words into paragraphs (by line, then by vertical gap) and detects tables separately via ruling lines or text alignment, then places both back in the order they appear on the page. Nothing is dropped, and table text never bleeds into surrounding paragraphs.pdf_to_text(pypdf) is a much lighter path: it just extracts each page's text in reading order, preserving the PDF's visual whitespace layout by default (so simple tables and columns still read naturally) without doing any table detection.
from xml2table import pdf_to_xml, pdf_to_text
pdf_to_xml("invoice.pdf", "invoice.xml") # paragraphs + tables (needs xml2table[pdf])
pdf_to_text("invoice.pdf", "invoice.txt") # fast raw text (needs xml2table[pdf-text])
invoice.xml looks like:
<document source="invoice.pdf" pages="1">
<page number="1" width="595.28" height="841.89">
<paragraph bbox="42.83,43.13,159.90,61.13">Invoice #1024</paragraph>
<table bbox="40.00,220.00,540.00,316.00" rows="4" cols="3">
<row><cell>Item</cell><cell>Qty</cell><cell>Price</cell></row>
<row><cell>Widget</cell><cell>2</cell><cell>$10.00</cell></row>
...
</table>
</page>
</document>
invoice.txt is pypdf's layout-preserving text, e.g.:
Invoice #1024
Bill To: Jane Doe
123 Example Street
Springfield, USA
Thank you for your business. Payment is due within thirty days...
Item Qty Price
Widget 2 $10.00
...
Reuse one PDFConverter for both (it lazily parses with pdfplumber only if
you call to_xml()/pages, and always uses pypdf for to_text()):
from xml2table import PDFConverter
converter = PDFConverter("invoice.pdf")
converter.to_xml("invoice.xml")
converter.to_text("invoice.txt")
for page in converter.pages:
print(page.number, len(page.paragraphs), len(page.tables))
PDFOptions reference
| Option | Default | Used by | Description |
|---|---|---|---|
line_tolerance |
3.0 |
XML | Max vertical gap (points) for words to count as the same line |
paragraph_gap |
6.0 |
XML | Min vertical gap (points) between lines that starts a new paragraph |
table_settings |
None |
XML | Passed through to pdfplumber's find_tables() for unusual tables (e.g. borderless). Caution: this applies to the whole page, not just the table — see Validation below before using it on documents with narrative text. |
cell_na |
"" |
XML | String used for empty/missing table cells |
page_separator |
"\n----- Page {page} -----\n" |
text | Inserted between pages |
keep_layout |
True |
text | Preserve the PDF's whitespace layout (pypdf "layout" mode) vs. plain, whitespace-normalized text |
CLI (PDF)
xml2table pdf invoice.pdf invoice.xml --to xml
xml2table pdf invoice.pdf invoice.txt --to text
Run xml2table pdf --help for all flags (--line-tolerance,
--paragraph-gap, --cell-na, --page-separator, --no-layout).
Validation: tested on 1,200 financial-report PDFs
pdf_to_xml and pdf_to_text were benchmarked against 1,200 generated
financial-report PDFs (balance sheets, income statements, cash-flow
statements, MD&A-style narrative text, footnotes — real financial
formatting: $1,234,567, (123,456) negatives, N/A blanks, multi-page,
ruled and borderless tables) with exact ground truth for every paragraph
and table cell, so the numbers below are measured, not estimated. Full
methodology, the generator, and raw per-document results are in
benchmarks/RESULTS.md.
| Metric | Result |
|---|---|
| Documents converted without error | 1,200 / 1,200 (100%) |
| Paragraph text fidelity (XML) | 100.000% |
| Ruled-table shape + cell fidelity | 100.000% |
| Borderless-table shape detection | 0.000% (documented limitation — see below) |
Raw-text content recall (pdf_to_text) |
100.000%, including for borderless tables |
| Throughput | 20.4 PDFs/sec (pdf_to_xml), 184.6 PDFs/sec (pdf_to_text) |
The one real limitation, quantified: pdfplumber's default table finder
needs ruling lines, so it doesn't detect borderless (text-only-aligned)
tables — but nothing is lost when it doesn't: the un-detected table's text
still comes through as ordinary paragraph text (pdf_to_text recall stays
at 100%). We also tested the obvious "fix" — pdfplumber's table_settings
override for borderless tables — across all 1,200 documents, and it made
things worse: because the override applies to the whole page, it started
misreading ordinary paragraph sentences as table cells, corrupting
paragraph output in 97.9% of documents (paragraph fidelity dropped from
100% to 13.9%). We did not ship that as a recommended workaround; see
PDFOptions.table_settings's docstring and benchmarks/RESULTS.md for the
full numbers and why.
We could not download real financial filings for this test — this sandboxed
session's network policy blocks direct access to sites like sec.gov — so
the corpus is synthetic but built to real financial-statement conventions
specifically so every value has a known-correct answer to grade against.
benchmarks/RESULTS.md explains this in more detail and gives the exact
commands to reproduce or extend the benchmark (e.g. against real filings, on
a machine with broader network access).
FlattenOptions reference
| Option | Default | Description |
|---|---|---|
record_path |
None (auto-detect) |
Path to the repeated record element |
attribute_prefix |
"@" |
Prefix for attribute-derived columns |
text_key |
"#text" |
Key for an element's own text when it also has attributes/children |
separator |
"." |
Separator for nested key paths |
array_mode |
"join" |
"join", "explode", or "index" |
join_separator |
"; " |
Separator used by "join" mode |
include_attributes |
True |
Include XML attributes as columns |
strip_namespaces |
True |
Strip {namespace} from tag/attribute names |
encoding |
"utf-8" |
Text encoding for reads/writes |
Errors
All exceptions inherit from xml2table.XMLConversionError:
XMLParseError— malformed XML inputRecordPathNotFoundError—record_pathmatched nothing, or automatic record detection was ambiguous (the error message tells you what to pass)PDFExtractionError— a PDF file could not be opened or parsedMissingOptionalDependencyError— e.g. callingto_dataframe()withoutpandas,pdf_to_xml()withoutpdfplumber, orpdf_to_text()withoutpypdf, installed
Development
pip install -e ".[dev]" # includes pandas, pdfplumber, pypdf, and fpdf2 (for regenerating PDF fixtures)
pytest
python examples/quickstart.py
Project layout:
src/xml2table/
parser.py # XML -> list[dict] flattening engine
options.py # FlattenOptions
converter.py # XMLConverter + module-level convenience functions
writers/ # CSV and Excel output
pdf_extract.py # pdfplumber: PDF -> Document(pages of Paragraph/Table), in reading order
pdf_xml_writer.py # Document -> XML
pdf_text_extract.py # pypdf: PDF -> plain text, independent of pdf_extract.py
pdf_options.py # PDFOptions
pdf_converter.py # PDFConverter + pdf_to_xml/pdf_to_text
cli.py # `xml2table` command-line tool
tests/
fixtures/ # sample XML documents and PDF fixtures (fixtures/pdf/)
test_*.py
examples/
quickstart.py
Metadata
Release files for xml2table 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| xml2table-0.1.0.tar.gz | 32.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| xml2table-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 60.2 kB
Release files / xml2table-0.1.0.tar.gz
| Download URL | xml2table-0.1.0.tar.gz |
|---|---|
| Size | 32.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2163d1feb9459a3932e8685ffe9e65d73e02d6c62a3a42d87a81e2d9e9cd9df8
|
|
BLAKE2b-256 checksum How to use checksums |
c427e8c7e39d6e3af12a8abd7fee4017e95c7cbcb8e9a56dead83ce0555744a1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency logRelease files / xml2table-0.1.0-py3-none-any.whl
| Download URL | xml2table-0.1.0-py3-none-any.whl |
|---|---|
| Size | 28.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3637c6b0d4b5869a9d22c7e5ebbe6be311c16c2a18d507caf1266c4477e9e7a9
|
|
BLAKE2b-256 checksum How to use checksums |
7ef005301aedb67f2e4bc0a3f2be1549bd260e9c5e02a381ad9d7b9b01835e0b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency log