FactFrame
PDF (and docx, xlsx) to markdown, where every block of text, every table row and every figure has a stable id — and any id resolves back to a rectangle on a page of the original PDF. An answer that cites an id is an answer a reader can check in one click.
pdf ──▶ provider ──▶ layout.json ──▶ annotate ──▶ document.md (ids inline)
└──▶ elements.json (id → page + box)
└──▶ manifest.json (pages, counts)
{{para:a1b2c3 bbox=1.00,2.35,7.50,3.10}}
The device operates from a single supply.
{{table:9f8e7d bbox=1.00,3.40,7.50,5.90}}
| {{row:0a2f13 table=9f8e7d idx=0}} Symbol | Min | Max |
| --- | --- | --- |
| {{row:44c1a0 table=9f8e7d idx=1}} VIH | 2.0 | 5.5 |
Use it from your AI assistant (MCP)
FactFrame runs as a local MCP server: point it at folders of PDF, docx and xlsx files and your assistant can convert, search, read and cite them — with a link from every citation to the exact place in the original. Everything happens on your machine; no document leaves it. Free for use on your own computer.
One command. It installs uv if you do not have it (uv is what downloads and runs the server), asks which folder to use, and registers the server with the desktop clients it finds — Claude Desktop, Cursor, Windsurf:
# macOS / Linux
curl -LsSf https://factframe.tech/install.sh | sh
# Windows (PowerShell): install uv, open a new window, then set up FactFrame
winget install --id=astral-sh.uv -e
uvx factframe-mcp install
(irm https://factframe.tech/install.ps1 | iex does both in one go, but
ad blockers and security tools often stop that kind of line from being copied.)
Already have uv? Skip the script:
uvx factframe-mcp install --root ~/Documents/contracts
Repeat --root for more folders; add --client cursor (or claude-desktop,
windsurf) to set up just one.
Claude Code:
claude mcp add factframe -- uvx factframe-mcp --root ~/Documents/contracts
Then ask your assistant to convert a folder and answer from it. Details and options are under MCP server below.
Not read: scanned PDFs with no text layer (there is no OCR in the local version). Need your team's Google Drive or SharePoint connected, or a secured deployment? Write to dany@factframe.tech.
Try it
pip install -e ".[local,dev]"
factframe convert your.pdf -o out/ # markdown + provenance index
factframe convert your.xlsx -o out/ # also .docx; needs the [xlsx] extra for workbooks
factframe show out/ 44c1a0 # where did that element come from?
factframe ask out/ "What is the max supply voltage?" # sourced answers (needs [openai])
# --llm anthropic for the Anthropic API instead of the default Azure OpenAI model:
factframe ask out/ "What is the max supply voltage?" --llm anthropic
Web demo
Side-by-side highlight demo — rendered markdown on the left, the source PDF on the right, click either to light up the other.
Install dependencies once:
cd web
npm install
The demo reads converted documents from web/demo/public/<slug>/ (or, for a
folder of PDFs converted together, web/demo/public/<folder>/<slug>/) —
document.md, elements.json, manifest.json, plus the source PDF itself
as source.pdf. You can drop one in by hand:
factframe convert your.pdf -o web/demo/public/sample
cp your.pdf web/demo/public/sample/source.pdf
Then run it:
cd web
npm run demo
Vite serves it on http://localhost:5173/ and opens it in a browser. Pick a
different port if that one is taken (npm run demo -- --port 5174). /
lists every converted document; each opens at /doc/<slug> (or
/doc/<folder>/<slug>), optionally with ?highlight=<id> to land scrolled
and flashed on one element — a bookmarkable, shareable deep link.
You don't have to convert up front, either — the doc list has a Convert &
view button for a single PDF or .xlsx and a Choose folder picker to convert
a whole folder of them in one go. Both post to a dev-server-only /api/convert
endpoint (see demo/vite.config.ts), which shells out to factframe convert
(the same local/pymupdf provider, titled from the filename) into
web/demo/public/<slugified-filename>/ (or, with a folder, one level
deeper). Needs factframe importable from the project's .venv (or on
$PATH) in whatever shell runs npm run demo — it only exists in the dev
server, not in npm run build.
MCP server
A local, stdio-transport MCP server. Six tools:
| Tool | What it does |
|---|---|
convert_folder |
convert every PDF, docx and xlsx in a folder and its subfolders; unchanged files are skipped on the next run |
list_documents |
what is already converted |
search_documents |
keyword search across documents (or one folder / one document), returning element ids |
read_document |
a document's markdown, ids inline, a range of pages at a time |
get_element |
one element's text by id, to check a citation |
make_link |
a URL that opens the viewer scrolled and highlighted to one element |
Options, on the command line the client runs:
| Option | Meaning |
|---|---|
--root FOLDER |
a folder the server may read from; repeatable. Required: without one, convert_folder refuses to run |
--store DIR |
where conversions are written. Default ~/.factframe/store |
--mirror |
also write a markdown copy next to each source file (report.pdf.md) |
FACTFRAME_ALLOWED_ROOTS (path-separator-separated), FACTFRAME_PUBLIC_DIR
and FACTFRAME_MIRROR=1 are the same three as environment variables.
uvx factframe-mcp install writes the entry for you. To add it by hand, every
client reads the same shape, from different files (claude_desktop_config.json,
.cursor/mcp.json, .vscode/mcp.json, ...):
{
"mcpServers": {
"factframe": {
"command": "/absolute/path/to/uvx",
"args": ["factframe-mcp", "--root", "/absolute/path/to/your/folder"]
}
}
}
Use absolute paths: a desktop app does not inherit your shell's PATH, and
some clients do not expand ~. VS Code wants "servers" instead of
"mcpServers".
The viewer. make_link's URLs open a read-only viewer that the server
itself runs on 127.0.0.1 (port 47200, or a free one if that is taken),
started the first time a link is asked for. It lives as long as the server
does, so a link opens on this computer while the client is running. Set
FACTFRAME_WEB_BASE_URL to point links at a viewer hosted elsewhere instead.
From a checkout. An editable install (pip install -e ".[mcp,local,xlsx]")
writes into web/demo/public/, so anything converted through the server shows
up in a running npm run demo; run npm run viewer:build in web/ to give
the built-in viewer something to serve.
Licensing. FactFrame is Apache-2.0. The local PDF reader is PyMuPDF, which is AGPL-3.0 (or commercially licensed by Artifex); that matters if you redistribute or host this as a service, not for running it on your own machine.
Configuration
Nothing is required to start. The default provider reads the PDF's own text layer locally — no account, no key, no per-page cost. Credentials only unlock the optional paths:
| Variable | Needed for | Notes |
|---|---|---|
AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT |
--provider azure |
e.g. https://<resource>.cognitiveservices.azure.com/ |
AZURE_DOCUMENT_INTELLIGENCE_KEY |
--provider azure |
one key, or several separated by commas |
AZURE_DOCUMENT_INTELLIGENCE_PAGES_PER_CHUNK |
free-tier Azure keys | set to 2 — see below |
AZURE_OPENAI_ENDPOINT |
factframe ask (default --llm openai) |
e.g. https://<resource>.openai.azure.com/ |
AZURE_OPENAI_API_KEY |
factframe ask (default --llm openai) |
|
AZURE_OPENAI_API_VERSION |
factframe ask (default --llm openai) |
defaults to 2025-04-01-preview (needed for the Responses API) |
ANTHROPIC_API_KEY |
factframe ask --llm anthropic |
conversion and highlighting work without it |
FACTFRAME_AUTH_USER |
Web demo & API | Username for HTTP Basic Auth |
FACTFRAME_AUTH_PASS |
Web demo & API | Password for HTTP Basic Auth |
FACTFRAME_HOST |
npm run demo (Vite dev server) |
Shell env var, not a .env entry — defaults to 127.0.0.1; set to 0.0.0.0 (or a specific address) to opt into binding all interfaces, e.g. for a remote dev session over a forwarded port. Same idea as HOST for npm run demo:serve (prod-server.mjs), which already defaults to 127.0.0.1. |
factframe ask defaults to --llm openai, a model deployed on Azure OpenAI
(gpt-5.4-mini by default) — pip install factframe[openai]. Pass --llm anthropic to use the Anthropic API instead (pip install factframe[anthropic]), and --model to override either adapter's default
model (for Azure OpenAI, this is the deployment name).
Working with free-tier Azure keys
The free (F0) tier reads only the first two pages of any document you submit. There is no error: pages three onward are simply absent from the result, so a 40-page datasheet silently converts to a two-page one. If you are on a free key, this is the single most important thing to know about the provider.
The fix is to split the PDF into two-page pieces, analyse each one, and stitch the results back into a single document:
export AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT="https://<resource>.cognitiveservices.azure.com/"
export AZURE_DOCUMENT_INTELLIGENCE_KEY="key1,key2,key3"
export AZURE_DOCUMENT_INTELLIGENCE_PAGES_PER_CHUNK=2
factframe convert your.pdf -o out/ --provider azure
AZURE_DOCUMENT_INTELLIGENCE_KEY takes several comma-separated keys. Chunks
are spread across them round-robin, and a retry always moves to the next key —
so a key that has hit its rate or monthly page limit doesn't fail the same chunk
twice. Three free keys give you three times the throughput and three times the
monthly page allowance.
Everything lives in providers/azure.py: split_pdf cuts the PDF, shift_pages
moves each chunk's page numbers to where they really belong (a chunk always
comes back numbered from 1), and _combine reassembles the document.
What stitching has to get right
Both of these are silent when wrong — nothing raises, you just get a document
that is subtly not the one you fed in. tests/test_azure_chunking.py pins both.
- Page numbers must be shifted everywhere, not just on the
pagesentries. Table cells and figure captions carry their ownboundingRegions, and a region left pointing at page 1 puts every highlight from that chunk on the cover page. - Span offsets must be rebased. Each element's
spansindex into its own chunk'scontentstring. Concatenate the strings without shifting the offsets and most spans point at the wrong text — measured at ~81% wrong on a real document. (FactFrame itself reads neithercontentnorspans; it renders from the positioned elements. The stitch is correct anyway, because the layout document is the contract.)
Caveats
- On a paid tier, leave chunking off. It is not an optimisation. The service costs roughly a fixed 84s per call plus ~0.11s per page, so on a 1380-page document five sequential chunks took ~573s against ~240s for a single call. Chunking wins only when a tier limit makes the single call impossible.
- Anything a provider computes per document is computed per chunk. A model that infers structure from statistics over the whole document sees only two pages at a time and may label a heading differently near a chunk boundary. Verified end to end on an 18-page datasheet processed two pages at a time: page numbering, tables, rows and figures all came out identical to a single-shot run, with two paragraphs reclassified as headings.
Layout
| Path | What |
|---|---|
src/factframe/ids.py |
stable, content-derived element ids |
src/factframe/geometry.py |
page geometry, boxes, page-percentage conversion |
src/factframe/layout.py |
the layout document contract + id assignment |
src/factframe/blocks.py |
grouping paragraphs into citable blocks |
src/factframe/markdown.py |
markdown rendering with provenance sentinels |
src/factframe/elements.py |
the element index: id → kind, page, box, text |
src/factframe/provenance.py |
citation → page region; citation checking |
src/factframe/ask.py |
sourced answers from an LLM, citations validated |
src/factframe/providers/azure.py |
Azure Document Intelligence + free-tier chunking |
src/factframe/providers/pymupdf.py |
local, offline, free extraction |
src/factframe/docx_convert.py, xlsx_convert.py |
docx and workbooks → the same artifacts, no provider |
src/factframe/retrieval.py |
search over a workbook: filter, BM25, cell pinning |
src/factframe/webstore.py |
slug/path rules shared by the web demo and the MCP server |
src/factframe/mcp_server.py |
the MCP server — convert, list, search, read, link |
src/factframe/viewer.py |
the read-only viewer the MCP server runs for its links |
src/factframe/install.py |
factframe-mcp install — registers the server with a client |
packages/factframe-mcp/ |
the launcher package behind uvx factframe-mcp |
web/src/ |
markdown-it anchor plugin, resolver, highlight |
web/demo/ |
the side-by-side viewer |
Every module's docstring explains not just what it does but why it does it that way — the design decisions are the part that took the longest to get right.
Design notes
- Ids are content-derived, never random. Re-converting the same PDF produces the same ids, so citations stored by an earlier run keep resolving.
- A citation is a bare id and nothing else. Ids are unique across kinds, so the kind is recovered from the index at read time. Every schema that stored the kind alongside the id eventually stored it wrong.
- The provider is the only impure stage. Everything after
layout.jsonis a pure function, so re-rendering after a code change costs nothing and re-runs need no OCR. - Every citation is checked before it is trusted. An invented citation is worse than a missing one: it presents as verified.
Run the tests with pytest.
Extracted from the FactFrame subsystem of datasheets.md and generalised to
arbitrary PDFs. Status: working end to end (Python pipeline, viewer, demo);
broader test coverage and docs/ still to come.
Metadata
Release files for factframe 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| factframe-0.1.1.tar.gz | 953.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| factframe-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.9 MB
Release files / factframe-0.1.1.tar.gz
| Download URL | factframe-0.1.1.tar.gz |
|---|---|
| Size | 953.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7aa7566dde6f1ab4a79ab576a7a3e0358ae1bccc9bb4965990cd1605c55be531
|
|
BLAKE2b-256 checksum How to use checksums |
68a5c485e5e227609b7b26ff9531dd99dd310c3396fa4882fd96f70159a390e4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency logRelease files / factframe-0.1.1-py3-none-any.whl
| Download URL | factframe-0.1.1-py3-none-any.whl |
|---|---|
| Size | 927.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1cbfc8c0baa96b843dc5c694f85aadfbf9d38abb9046265dbe5d85df105ad6b0
|
|
BLAKE2b-256 checksum How to use checksums |
333bd1075287fc816a545cd9c59ab9a8c2de6c54f7e4ad8324b7ea519b8457af
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency log