pdfschema
Pull just the columns you need out of a PDF, so your LLM prompt carries data instead of a document.
pip install pdfschema
import pdfschema
rows = pdfschema.extract_rows(
"statement.pdf",
schema=["Date", "Particulars", "Debit"],
)
[
{"Date": "01-04-2025", "Particulars": "PAYMENT MADE VIA UPI TO VPA MERCHANT7X2K@EXAMPLEPAY AGAINST REF 100000000001", "Debit": "390.00"}
]
Every example here uses fictional transaction data. The measurements below are real, taken from a 125-page statement that is not published.
The problem this solves
Feeding a whole PDF to a model is the expensive way to answer a narrow question. A 125-page bank statement is ~44,000 tokens of prompt for every single call — and most of it is page furniture, letterhead, disclaimers and columns the question never touches. At scale that is the bill.
pdfschema lets you name the columns you actually need and get structured rows back. You extract once, then filter, aggregate or sample in Python before anything reaches the model.
Measured on a real 125-page statement
All six columns, typed schema:
| What you send | Tokens (approx.) | vs. raw page text |
|---|---|---|
| Raw extracted page text | 43,951 | — |
to_csv() |
28,540 | −35% |
to_tson() |
29,974 | −32% |
to_toon() |
30,005 | −32% |
to_tson_columnar() |
30,246 | −31% |
to_json() |
50,577 | +15% ⚠️ |
Narrow the columns and it compounds: three columns as CSV is 21,565 tokens (−51%), two columns 4,687 (−89%).
Don't put JSON in a prompt. An array of objects repeats every key on every
row, which on this document costs more than the raw text it replaced. Keep
to_json() for the code that consumes the rows afterwards — that is what it's
good at — and send one of the formats that names the fields once:
table = pdfschema.extract("statement.pdf", schema={
"Date": "date", "Particulars": "string", "Debit": "number",
})[0]
prompt = f"Transactions:\n{table.to_toon(name='transactions')}\n\nWhich merchant took the most?"
transactions[970]{Date,Particulars,Debit}:
2025-04-01,PAYMENT MADE VIA UPI TO VPA MERCHANT7X2K@EXAMPLEPAY AGAINST REF …,390
2025-04-02,PAYMENT RECEIVED VIA UPI FROM VPA J.DOE-1@EXAMPLEBANK FROM …,null
(… marks text elided for width.)
The real win is narrowing. Most of the saving comes from dropping columns and boilerplate the question never needed — and from filtering rows in Python, which no amount of prompt engineering over raw PDF text does reliably:
april = [r for r in rows if r["Date"].month == 4] # with schema={"Date": "date"}
Why a schema, and not a grid detector
Most extractors look for the lines a page draws and infer columns from them. Bank statements, invoices and most generated business PDFs draw row bands but no vertical rules — so there is nothing to find, and those tools hand back one column with every row collapsed into a single string.
pdfschema starts from the header. Your labels are located on the page and their x positions become the column boundaries, so an unruled table reads exactly like a ruled one. The same labels identify the table's continuation on the next page, which is how a 125-page statement comes back as one table rather than 125 fragments.
Modes
Six independent switches that compose freely. Pick what you care about, leave the rest at their defaults.
| Mode | Values | Default | Set with | Controls |
|---|---|---|---|---|
| Extraction | schema / discovery | discovery | schema= |
Which tables come back, and what the keys are called |
| Value | string, number, date |
string |
mapping schema= |
The Python type of each cell |
| Output | json, csv, toon, tson, tson-columnar, lists, DataFrame |
— | to_*() method |
The text or object you hand on |
| Cell join | smart, space, newline |
smart |
join= |
How a cell wrapped over several lines is rebuilt |
| Merge | on / off | on | merge= |
Whether a table split across pages returns as one |
| Date reading | day-first / month-first | day-first | dayfirst= |
How 01/04/2025 is read |
| Failure | lenient / strict | lenient | strict= |
Whether a cell that won't convert is None or an error |
table.strategy additionally reports how rows were found — "ruled" (the page
drew them), "clustered" (inferred from word positions) or "mixed".
Extraction mode 1 — schema
Name the columns; get exactly those keys back, in your order and your spelling.
rows = pdfschema.extract_rows("statement.pdf", schema=["Date", "Balance"])
Matching is normalised subset: a table matches when its header supplies every
label you asked for, compared case-insensitively with whitespace collapsed.
Columns you didn't ask for are dropped. "transaction id", "Transaction ID"
and "TRANSACTION ID" all find the same column.
If no table carries your columns you get SchemaNotFoundError — naming the
unmatched labels and listing the headers that do exist — not an empty list your
pipeline silently passes along.
Extraction mode 2 — discovery
Omit the schema and every table found comes back, keyed by the labels printed on the page:
for table in pdfschema.extract("statement.pdf"):
print(table.header, table.n_rows, table.page_span)
# ['Date', 'Transaction ID', 'Particulars', 'Debit', 'Credit', 'Balance'] 970 1-125
Those labels are exactly what schema= accepts, so this is the first step when
onboarding an unfamiliar document.
Value modes
| Value mode | Python type | None when |
|---|---|---|
string |
str |
never — an empty cell is "" |
number |
float |
the cell is empty or a lone dash |
date |
datetime.date |
the cell is empty or a lone dash |
rows = pdfschema.extract_rows("statement.pdf", schema={
"Date": "date", "Debit": "number", "Particulars": "string",
})
rows[0]["Date"] # datetime.date(2025, 4, 1)
rows[0]["Debit"] # 390.0
number handles currency symbols, thousands separators, percentages and
accounting negatives like (250.50). date tries a fixed list of formats.
Other modes
pdfschema.extract(path, schema, pages="1,3,5-7") # page selection
pdfschema.extract(path, schema, merge=False) # one table per page
pdfschema.extract(path, schema, join="newline") # keep the PDF's line breaks
pdfschema.extract(path, schema, dayfirst=False) # 01/04/2025 is 4 January
pdfschema.extract(path, schema, strict=True) # raise instead of storing None
join decides how a cell wrapped over several lines is rebuilt. PDFs often chop
text every N characters regardless of word boundaries, so a naive join gives you
either @EXAMPLEPAYAGAINST or VP A MERCHANT; "smart" (the default) works the wrap
style out per column and rebuilds the original string.
Output modes and types
Every extraction returns a list[Table]. The output mode is whichever method you
then call — nothing about the extraction changes, only the shape you hand on.
| Method | Returns | Use for |
|---|---|---|
.rows / to_dicts() |
list[dict] |
Python code |
to_json(indent=) |
str |
APIs, files, another program. Not prompts. |
to_lists() |
list[list] |
Spreadsheet writers, the csv module |
to_dataframe() |
pandas.DataFrame |
Analysis, aggregation, joins |
to_csv() |
str |
Prompts (cheapest), spreadsheets, DB load |
to_toon(name=) |
str |
Prompts (most reliable) — TOON, adds a row count and field list for ~5% over CSV |
to_tson() |
str |
zenoaihq TSON — single line, delimiter-based |
to_tson_columnar(name=) |
str |
tsonformat.com TSON — indentation-based, transposed |
to_prompt(fmt, **opts) |
str |
The format comes from config or a flag |
to_dataframe() needs pip install pdfschema[pandas]; the rest are always
available. The two TSONs are unrelated formats that share a name — one
delimiter-based, one indentation-based — and pdfschema implements both rather
than picking for you.
Picking a combination:
| You want | Extraction | Value | Output |
|---|---|---|---|
| Rows in an LLM prompt, cheapest | schema | typed | to_csv() |
| Rows in a prompt, model must not miscount | schema | typed | to_toon(name=…) |
| Rows for your own code | schema | typed | to_json() / .rows |
| Analysis and aggregation | schema | typed | to_dataframe() |
| Onboarding an unfamiliar document | discovery | string | .header |
| Faithful transcription | schema | string | to_csv() |
One thing worth knowing: every prompt format must quote a string that would
otherwise read as a number, so an untyped Debit column emits "390.00" where
schema={"Debit": "number"} emits 390. Typing the numeric columns makes the
prompt smaller and less ambiguous.
These formats are newer than CSV and how reliably a given model reads one varies by model — benchmark against yours before switching a pipeline over.
Command line
$ pdfschema discover statement.pdf
[1] pages 1-125 rows 970 (ruled)
schema: ["Date", "Transaction ID", "Particulars", "Debit", "Credit", "Balance"]
$ pdfschema extract statement.pdf --schema "Date:date,Debit:number" -o rows.json
$ pdfschema extract statement.pdf --schema "Date,Balance" --pages 1-5 --format csv
$ pdfschema extract statement.pdf --schema "Date:date" --format toon --name transactions
--format takes json (default), csv, toon, tson, tson-columnar. Also
--pages, --name, --join, --no-merge, --monthfirst, --strict,
-o/--output, --indent. extract exits 2 if the schema is not found.
What you get
- Deterministic and auditable. No model in the loop, no temperature, no hallucinated cells. The same PDF gives the same rows every time — which is what you want in the layer feeding an LLM.
- Wrapped cells rejoined correctly, multi-page tables merged, multi-table pages handled.
- Errors you can branch on.
SchemaNotFoundErrorcarries.unmatchedand.headers_found; everything deliberate derives fromPdfSchemaError.
Typical uses
- RAG ingestion — turn statements, invoices and reports into rows worth embedding, instead of embedding page images or raw text dumps.
- Agent tools — back a
get_transactions(start, end)tool with a real extraction rather than a whole-document prompt. - Batch pipelines — thousands of documents a day, where a 50% prompt reduction is the difference in the invoice.
- Pre-flight validation — check a document has the columns you expect before spending a model call on it.
- Plain ETL — no LLM anywhere; PDF to CSV, database or DataFrame.
Requirements
Python 3.11+ and pdfplumber. Optional:
pdfschema[pandas] for to_dataframe().
Documentation
Full documentation, limitations and examples: https://github.com/nullbite-coder/pdfschema
MIT licensed.
Metadata
Release files for pdfschema 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pdfschema-0.1.0.tar.gz | 53.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pdfschema-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 95.2 kB
Release files / pdfschema-0.1.0.tar.gz
| Download URL | pdfschema-0.1.0.tar.gz |
|---|---|
| Size | 53.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6bab247a798c27953ac8dfa50a7779c79bddaa153024ed967f462e7b31772170
|
|
BLAKE2b-256 checksum How to use checksums |
5803b5b0a326ff34788c6ab53dbb38d4780bd4de6e47d1be0a646b0484743109
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.
Transparency logRelease files / pdfschema-0.1.0-py3-none-any.whl
| Download URL | pdfschema-0.1.0-py3-none-any.whl |
|---|---|
| Size | 41.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8d066d833e32f774907393251ca5c8a1de26ee2ca1b947f89a26a0d6ce184f4b
|
|
BLAKE2b-256 checksum How to use checksums |
f97bed6ebc336320079263a9f6d531292104092b5d4efe64dfa14b27b8661f26
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.
Transparency log