Skip to main content

pdfschema

Pull just the columns you need out of a PDF, so your LLM prompt carries data instead of a document.

pip install pdfschema
import pdfschema

rows = pdfschema.extract_rows(
    "statement.pdf",
    schema=["Date", "Particulars", "Debit"],
)
[
  {"Date": "01-04-2025", "Particulars": "PAYMENT MADE VIA UPI TO VPA MERCHANT7X2K@EXAMPLEPAY AGAINST REF 100000000001", "Debit": "390.00"}
]

Every example here uses fictional transaction data. The measurements below are real, taken from a 125-page statement that is not published.

The problem this solves

Feeding a whole PDF to a model is the expensive way to answer a narrow question. A 125-page bank statement is ~44,000 tokens of prompt for every single call — and most of it is page furniture, letterhead, disclaimers and columns the question never touches. At scale that is the bill.

pdfschema lets you name the columns you actually need and get structured rows back. You extract once, then filter, aggregate or sample in Python before anything reaches the model.

Measured on a real 125-page statement

All six columns, typed schema:

What you send Tokens (approx.) vs. raw page text
Raw extracted page text 43,951 —
to_csv() 28,540 −35%
to_tson() 29,974 −32%
to_toon() 30,005 −32%
to_tson_columnar() 30,246 −31%
to_json() 50,577 +15% ⚠️

Narrow the columns and it compounds: three columns as CSV is 21,565 tokens (−51%), two columns 4,687 (−89%).

Don't put JSON in a prompt. An array of objects repeats every key on every row, which on this document costs more than the raw text it replaced. Keep to_json() for the code that consumes the rows afterwards — that is what it's good at — and send one of the formats that names the fields once:

table = pdfschema.extract("statement.pdf", schema={
    "Date": "date", "Particulars": "string", "Debit": "number",
})[0]

prompt = f"Transactions:\n{table.to_toon(name='transactions')}\n\nWhich merchant took the most?"
transactions[970]{Date,Particulars,Debit}:
  2025-04-01,PAYMENT MADE VIA UPI TO VPA MERCHANT7X2K@EXAMPLEPAY AGAINST REF …,390
  2025-04-02,PAYMENT RECEIVED VIA UPI FROM VPA J.DOE-1@EXAMPLEBANK FROM …,null

(… marks text elided for width.)

The real win is narrowing. Most of the saving comes from dropping columns and boilerplate the question never needed — and from filtering rows in Python, which no amount of prompt engineering over raw PDF text does reliably:

april = [r for r in rows if r["Date"].month == 4]   # with schema={"Date": "date"}

Why a schema, and not a grid detector

Most extractors look for the lines a page draws and infer columns from them. Bank statements, invoices and most generated business PDFs draw row bands but no vertical rules — so there is nothing to find, and those tools hand back one column with every row collapsed into a single string.

pdfschema starts from the header. Your labels are located on the page and their x positions become the column boundaries, so an unruled table reads exactly like a ruled one. The same labels identify the table's continuation on the next page, which is how a 125-page statement comes back as one table rather than 125 fragments.

Modes

Six independent switches that compose freely. Pick what you care about, leave the rest at their defaults.

Mode Values Default Set with Controls
Extraction schema / discovery discovery schema= Which tables come back, and what the keys are called
Value string, number, date string mapping schema= The Python type of each cell
Output json, csv, toon, tson, tson-columnar, lists, DataFrame — to_*() method The text or object you hand on
Cell join smart, space, newline smart join= How a cell wrapped over several lines is rebuilt
Merge on / off on merge= Whether a table split across pages returns as one
Date reading day-first / month-first day-first dayfirst= How 01/04/2025 is read
Failure lenient / strict lenient strict= Whether a cell that won't convert is None or an error

table.strategy additionally reports how rows were found — "ruled" (the page drew them), "clustered" (inferred from word positions) or "mixed".

Extraction mode 1 — schema

Name the columns; get exactly those keys back, in your order and your spelling.

rows = pdfschema.extract_rows("statement.pdf", schema=["Date", "Balance"])

Matching is normalised subset: a table matches when its header supplies every label you asked for, compared ignoring case and whitespace. Columns you didn't ask for are dropped. "transaction id", "Transaction ID", "TRANSACTION ID" and "TransactionID" all find the same column.

Headers that wrap across lines — common in narrow ruled columns, where Application No prints as Application / No — are read cell by cell from the drawn grid, so you ask for the label as it reads: schema=["Application No", "Hypothecation Type"].

If no table carries your columns you get SchemaNotFoundError — naming the unmatched labels and listing the headers that do exist — not an empty list your pipeline silently passes along.

Extraction mode 2 — discovery

Omit the schema and every table found comes back, keyed by the labels printed on the page:

for table in pdfschema.extract("statement.pdf"):
    print(table.header, table.n_rows, table.page_span)
# ['Date', 'Transaction ID', 'Particulars', 'Debit', 'Credit', 'Balance'] 970 1-125

Those labels are exactly what schema= accepts, so this is the first step when onboarding an unfamiliar document.

Value modes

Value mode Python type None when
string str never — an empty cell is ""
number float the cell is empty or a lone dash
date datetime.date the cell is empty or a lone dash
rows = pdfschema.extract_rows("statement.pdf", schema={
    "Date": "date", "Debit": "number", "Particulars": "string",
})
rows[0]["Date"]   # datetime.date(2025, 4, 1)
rows[0]["Debit"]  # 390.0

number handles currency symbols, thousands separators, percentages and accounting negatives like (250.50). date tries a fixed list of formats.

Other modes

pdfschema.extract(path, schema, pages="1,3,5-7")   # page selection
pdfschema.extract(path, schema, merge=False)       # one table per page
pdfschema.extract(path, schema, join="newline")    # keep the PDF's line breaks
pdfschema.extract(path, schema, dayfirst=False)    # 01/04/2025 is 4 January
pdfschema.extract(path, schema, strict=True)       # raise instead of storing None

join decides how a cell wrapped over several lines is rebuilt. PDFs often chop text every N characters regardless of word boundaries, so a naive join gives you either @EXAMPLEPAYAGAINST or VP A MERCHANT; "smart" (the default) works the wrap style out per column and rebuilds the original string.

Output modes and types

Every extraction returns a list[Table]. The output mode is whichever method you then call — nothing about the extraction changes, only the shape you hand on.

Method Returns Use for
.rows / to_dicts() list[dict] Python code
to_json(indent=) str APIs, files, another program. Not prompts.
to_lists() list[list] Spreadsheet writers, the csv module
to_dataframe() pandas.DataFrame Analysis, aggregation, joins
to_csv() str Prompts (cheapest), spreadsheets, DB load
to_toon(name=) str Prompts (most reliable) — TOON, adds a row count and field list for ~5% over CSV
to_tson() str zenoaihq TSON — single line, delimiter-based
to_tson_columnar(name=) str tsonformat.com TSON — indentation-based, transposed
to_prompt(fmt, **opts) str The format comes from config or a flag

to_dataframe() needs pip install pdfschema[pandas]; the rest are always available. The two TSONs are unrelated formats that share a name — one delimiter-based, one indentation-based — and pdfschema implements both rather than picking for you.

Picking a combination:

You want Extraction Value Output
Rows in an LLM prompt, cheapest schema typed to_csv()
Rows in a prompt, model must not miscount schema typed to_toon(name=…)
Rows for your own code schema typed to_json() / .rows
Analysis and aggregation schema typed to_dataframe()
Onboarding an unfamiliar document discovery string .header
Faithful transcription schema string to_csv()

One thing worth knowing: every prompt format must quote a string that would otherwise read as a number, so an untyped Debit column emits "390.00" where schema={"Debit": "number"} emits 390. Typing the numeric columns makes the prompt smaller and less ambiguous.

These formats are newer than CSV and how reliably a given model reads one varies by model — benchmark against yours before switching a pipeline over.

Command line

$ pdfschema discover statement.pdf
[1] pages 1-125  rows 970  (ruled)
    schema: ["Date", "Transaction ID", "Particulars", "Debit", "Credit", "Balance"]

$ pdfschema extract statement.pdf --schema "Date:date,Debit:number" -o rows.json
$ pdfschema extract statement.pdf --schema "Date,Balance" --pages 1-5 --format csv
$ pdfschema extract statement.pdf --schema "Date:date" --format toon --name transactions

--format takes json (default), csv, toon, tson, tson-columnar. Also --pages, --name, --join, --no-merge, --monthfirst, --strict, -o/--output, --indent. extract exits 2 if the schema is not found.

What you get

  • Deterministic and auditable. No model in the loop, no temperature, no hallucinated cells. The same PDF gives the same rows every time — which is what you want in the layer feeding an LLM.
  • Wrapped cells rejoined correctly, multi-page tables merged, multi-table pages handled.
  • Errors you can branch on. SchemaNotFoundError carries .unmatched and .headers_found; everything deliberate derives from PdfSchemaError.

Typical uses

  • RAG ingestion — turn statements, invoices and reports into rows worth embedding, instead of embedding page images or raw text dumps.
  • Agent tools — back a get_transactions(start, end) tool with a real extraction rather than a whole-document prompt.
  • Batch pipelines — thousands of documents a day, where a 50% prompt reduction is the difference in the invoice.
  • Pre-flight validation — check a document has the columns you expect before spending a model call on it.
  • Plain ETL — no LLM anywhere; PDF to CSV, database or DataFrame.

Requirements

Python 3.11+ and pdfplumber. Optional: pdfschema[pandas] for to_dataframe().

Documentation

Full documentation, limitations and examples: https://github.com/nullbite-coder/pdfschema

MIT licensed.

Metadata

Release files for pdfschema 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pdfschema 0.2.0
File Size Uploaded
pdfschema-0.2.0.tar.gz 63.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pdfschema 0.2.0
File Interpreter ABI Platform
pdfschema-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 110.5 kB

Release files / pdfschema-0.2.0.tar.gz

Download URL pdfschema-0.2.0.tar.gz
Size 63.5 kB
Tags Source
SHA-256 checksum
How to use checksums
db55f05702522abc8a2b159ef85d1ad2159fc92e69a1c274aeecabff36bdf619
BLAKE2b-256 checksum
How to use checksums
f6b2cf1de6f9fce9ed473c0bd7c2796ab5e0b335316da94807d3df8309261851
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release files / pdfschema-0.2.0-py3-none-any.whl

Download URL pdfschema-0.2.0-py3-none-any.whl
Size 47.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5c0efed1031499404ab4701477d6c6d2f7e75c9fc94f4c8a8089d1feb7b935d9
BLAKE2b-256 checksum
How to use checksums
9c8ec70e5e1668d75abf582c422e43a3ba57d71032c883bf4c6a87f2e31f80d8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page