Skip to main content

FactFrame

PDF (and docx, xlsx) to markdown, where every block of text, every table row and every figure has a stable id — and any id resolves back to a rectangle on a page of the original PDF. An answer that cites an id is an answer a reader can check in one click.

pdf ──▶ provider ──▶ layout.json ──▶ annotate ──▶ document.md      (ids inline)
                                             └──▶ elements.json    (id → page + box)
                                             └──▶ manifest.json    (pages, counts)
{{para:a1b2c3 bbox=1.00,2.35,7.50,3.10}}
The device operates from a single supply.

{{table:9f8e7d bbox=1.00,3.40,7.50,5.90}}
| {{row:0a2f13 table=9f8e7d idx=0}} Symbol | Min | Max |
| --- | --- | --- |
| {{row:44c1a0 table=9f8e7d idx=1}} VIH | 2.0 | 5.5 |

Use it from your AI assistant (MCP)

FactFrame runs as a local MCP server: point it at folders of PDF, docx and xlsx files and your assistant can convert, search, read and cite them — with a link from every citation to the exact place in the original. Everything happens on your machine; no document leaves it. Free for use on your own computer.

One command, using uv:

uvx factframe-mcp install --root ~/Documents/contracts

That registers the server with the desktop clients it finds (Claude Desktop, Cursor, Windsurf) and tells you what to restart. Repeat --root for more folders. No uv yet?

# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh && uvx factframe-mcp install
# Windows (PowerShell)
irm https://astral.sh/uv/install.ps1 | iex; uvx factframe-mcp install

Claude Code:

claude mcp add factframe -- uvx factframe-mcp --root ~/Documents/contracts

Then ask your assistant to convert a folder and answer from it. Details and options are under MCP server below.

Not read: scanned PDFs with no text layer (there is no OCR in the local version). Need your team's Google Drive or SharePoint connected, or a secured deployment? Write to dany@factframe.tech.

Try it

pip install -e ".[local,dev]"

factframe convert your.pdf -o out/          # markdown + provenance index
factframe convert your.xlsx -o out/         # also .docx; needs the [xlsx] extra for workbooks
factframe show out/ 44c1a0                  # where did that element come from?
factframe ask out/ "What is the max supply voltage?"   # sourced answers (needs [openai])

# --llm anthropic for the Anthropic API instead of the default Azure OpenAI model:
factframe ask out/ "What is the max supply voltage?" --llm anthropic

Web demo

Side-by-side highlight demo — rendered markdown on the left, the source PDF on the right, click either to light up the other.

Install dependencies once:

cd web
npm install

The demo reads converted documents from web/demo/public/<slug>/ (or, for a folder of PDFs converted together, web/demo/public/<folder>/<slug>/) — document.md, elements.json, manifest.json, plus the source PDF itself as source.pdf. You can drop one in by hand:

factframe convert your.pdf -o web/demo/public/sample
cp your.pdf web/demo/public/sample/source.pdf

Then run it:

cd web
npm run demo

Vite serves it on http://localhost:5173/ and opens it in a browser. Pick a different port if that one is taken (npm run demo -- --port 5174). / lists every converted document; each opens at /doc/<slug> (or /doc/<folder>/<slug>), optionally with ?highlight=<id> to land scrolled and flashed on one element — a bookmarkable, shareable deep link.

You don't have to convert up front, either — the doc list has a Convert & view button for a single PDF or .xlsx and a Choose folder picker to convert a whole folder of them in one go. Both post to a dev-server-only /api/convert endpoint (see demo/vite.config.ts), which shells out to factframe convert (the same local/pymupdf provider, titled from the filename) into web/demo/public/<slugified-filename>/ (or, with a folder, one level deeper). Needs factframe importable from the project's .venv (or on $PATH) in whatever shell runs npm run demo — it only exists in the dev server, not in npm run build.

MCP server

A local, stdio-transport MCP server. Six tools:

Tool What it does
convert_folder convert every PDF, docx and xlsx in a folder and its subfolders; unchanged files are skipped on the next run
list_documents what is already converted
search_documents keyword search across documents (or one folder / one document), returning element ids
read_document a document's markdown, ids inline, a range of pages at a time
get_element one element's text by id, to check a citation
make_link a URL that opens the viewer scrolled and highlighted to one element

Options, on the command line the client runs:

Option Meaning
--root FOLDER a folder the server may read from; repeatable. Required: without one, convert_folder refuses to run
--store DIR where conversions are written. Default ~/.factframe/store
--mirror also write a markdown copy next to each source file (report.pdf.md)

FACTFRAME_ALLOWED_ROOTS (path-separator-separated), FACTFRAME_PUBLIC_DIR and FACTFRAME_MIRROR=1 are the same three as environment variables.

uvx factframe-mcp install writes the entry for you. To add it by hand, every client reads the same shape, from different files (claude_desktop_config.json, .cursor/mcp.json, .vscode/mcp.json, ...):

{
  "mcpServers": {
    "factframe": {
      "command": "/absolute/path/to/uvx",
      "args": ["factframe-mcp", "--root", "/absolute/path/to/your/folder"]
    }
  }
}

Use absolute paths: a desktop app does not inherit your shell's PATH, and some clients do not expand ~. VS Code wants "servers" instead of "mcpServers".

The viewer. make_link's URLs open a read-only viewer that the server itself runs on 127.0.0.1 (port 47200, or a free one if that is taken), started the first time a link is asked for. It lives as long as the server does, so a link opens on this computer while the client is running. Set FACTFRAME_WEB_BASE_URL to point links at a viewer hosted elsewhere instead.

From a checkout. An editable install (pip install -e ".[mcp,local,xlsx]") writes into web/demo/public/, so anything converted through the server shows up in a running npm run demo; run npm run viewer:build in web/ to give the built-in viewer something to serve.

Licensing. FactFrame is Apache-2.0. The local PDF reader is PyMuPDF, which is AGPL-3.0 (or commercially licensed by Artifex); that matters if you redistribute or host this as a service, not for running it on your own machine.

Configuration

Nothing is required to start. The default provider reads the PDF's own text layer locally — no account, no key, no per-page cost. Credentials only unlock the optional paths:

Variable Needed for Notes
AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT --provider azure e.g. https://<resource>.cognitiveservices.azure.com/
AZURE_DOCUMENT_INTELLIGENCE_KEY --provider azure one key, or several separated by commas
AZURE_DOCUMENT_INTELLIGENCE_PAGES_PER_CHUNK free-tier Azure keys set to 2 — see below
AZURE_OPENAI_ENDPOINT factframe ask (default --llm openai) e.g. https://<resource>.openai.azure.com/
AZURE_OPENAI_API_KEY factframe ask (default --llm openai)
AZURE_OPENAI_API_VERSION factframe ask (default --llm openai) defaults to 2025-04-01-preview (needed for the Responses API)
ANTHROPIC_API_KEY factframe ask --llm anthropic conversion and highlighting work without it
FACTFRAME_AUTH_USER Web demo & API Username for HTTP Basic Auth
FACTFRAME_AUTH_PASS Web demo & API Password for HTTP Basic Auth
FACTFRAME_HOST npm run demo (Vite dev server) Shell env var, not a .env entry — defaults to 127.0.0.1; set to 0.0.0.0 (or a specific address) to opt into binding all interfaces, e.g. for a remote dev session over a forwarded port. Same idea as HOST for npm run demo:serve (prod-server.mjs), which already defaults to 127.0.0.1.

factframe ask defaults to --llm openai, a model deployed on Azure OpenAI (gpt-5.4-mini by default) — pip install factframe[openai]. Pass --llm anthropic to use the Anthropic API instead (pip install factframe[anthropic]), and --model to override either adapter's default model (for Azure OpenAI, this is the deployment name).

Working with free-tier Azure keys

The free (F0) tier reads only the first two pages of any document you submit. There is no error: pages three onward are simply absent from the result, so a 40-page datasheet silently converts to a two-page one. If you are on a free key, this is the single most important thing to know about the provider.

The fix is to split the PDF into two-page pieces, analyse each one, and stitch the results back into a single document:

export AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT="https://<resource>.cognitiveservices.azure.com/"
export AZURE_DOCUMENT_INTELLIGENCE_KEY="key1,key2,key3"
export AZURE_DOCUMENT_INTELLIGENCE_PAGES_PER_CHUNK=2

factframe convert your.pdf -o out/ --provider azure

AZURE_DOCUMENT_INTELLIGENCE_KEY takes several comma-separated keys. Chunks are spread across them round-robin, and a retry always moves to the next key — so a key that has hit its rate or monthly page limit doesn't fail the same chunk twice. Three free keys give you three times the throughput and three times the monthly page allowance.

Everything lives in providers/azure.py: split_pdf cuts the PDF, shift_pages moves each chunk's page numbers to where they really belong (a chunk always comes back numbered from 1), and _combine reassembles the document.

What stitching has to get right

Both of these are silent when wrong — nothing raises, you just get a document that is subtly not the one you fed in. tests/test_azure_chunking.py pins both.

  • Page numbers must be shifted everywhere, not just on the pages entries. Table cells and figure captions carry their own boundingRegions, and a region left pointing at page 1 puts every highlight from that chunk on the cover page.
  • Span offsets must be rebased. Each element's spans index into its own chunk's content string. Concatenate the strings without shifting the offsets and most spans point at the wrong text — measured at ~81% wrong on a real document. (FactFrame itself reads neither content nor spans; it renders from the positioned elements. The stitch is correct anyway, because the layout document is the contract.)

Caveats

  • On a paid tier, leave chunking off. It is not an optimisation. The service costs roughly a fixed 84s per call plus ~0.11s per page, so on a 1380-page document five sequential chunks took ~573s against ~240s for a single call. Chunking wins only when a tier limit makes the single call impossible.
  • Anything a provider computes per document is computed per chunk. A model that infers structure from statistics over the whole document sees only two pages at a time and may label a heading differently near a chunk boundary. Verified end to end on an 18-page datasheet processed two pages at a time: page numbering, tables, rows and figures all came out identical to a single-shot run, with two paragraphs reclassified as headings.

Layout

Path What
src/factframe/ids.py stable, content-derived element ids
src/factframe/geometry.py page geometry, boxes, page-percentage conversion
src/factframe/layout.py the layout document contract + id assignment
src/factframe/blocks.py grouping paragraphs into citable blocks
src/factframe/markdown.py markdown rendering with provenance sentinels
src/factframe/elements.py the element index: id → kind, page, box, text
src/factframe/provenance.py citation → page region; citation checking
src/factframe/ask.py sourced answers from an LLM, citations validated
src/factframe/providers/azure.py Azure Document Intelligence + free-tier chunking
src/factframe/providers/pymupdf.py local, offline, free extraction
src/factframe/docx_convert.py, xlsx_convert.py docx and workbooks → the same artifacts, no provider
src/factframe/retrieval.py search over a workbook: filter, BM25, cell pinning
src/factframe/webstore.py slug/path rules shared by the web demo and the MCP server
src/factframe/mcp_server.py the MCP server — convert, list, search, read, link
src/factframe/viewer.py the read-only viewer the MCP server runs for its links
src/factframe/install.py factframe-mcp install — registers the server with a client
packages/factframe-mcp/ the launcher package behind uvx factframe-mcp
web/src/ markdown-it anchor plugin, resolver, highlight
web/demo/ the side-by-side viewer

Every module's docstring explains not just what it does but why it does it that way — the design decisions are the part that took the longest to get right.

Design notes

  • Ids are content-derived, never random. Re-converting the same PDF produces the same ids, so citations stored by an earlier run keep resolving.
  • A citation is a bare id and nothing else. Ids are unique across kinds, so the kind is recovered from the index at read time. Every schema that stored the kind alongside the id eventually stored it wrong.
  • The provider is the only impure stage. Everything after layout.json is a pure function, so re-rendering after a code change costs nothing and re-runs need no OCR.
  • Every citation is checked before it is trusted. An invented citation is worse than a missing one: it presents as verified.

Run the tests with pytest.

Extracted from the FactFrame subsystem of datasheets.md and generalised to arbitrary PDFs. Status: working end to end (Python pipeline, viewer, demo); broader test coverage and docs/ still to come.

Metadata

Release files for factframe 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for factframe 0.1.0
File Size Uploaded
factframe-0.1.0.tar.gz 949.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for factframe 0.1.0
File Interpreter ABI Platform
factframe-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.9 MB

Release files / factframe-0.1.0.tar.gz

Download URL factframe-0.1.0.tar.gz
Size 949.9 kB
Tags Source
SHA-256 checksum
How to use checksums
6cc18cbb8a280c53452582a4bc2a6a984dfef1676d6042789b3450b53e8a17ff
BLAKE2b-256 checksum
How to use checksums
f683e9719ab766e878d1f33f3a287aeee9213c64169cacaedaa195f439d8b9f7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / factframe-0.1.0-py3-none-any.whl

Download URL factframe-0.1.0-py3-none-any.whl
Size 925.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2cb2d2a1c8737fb48034111d327bc1e43c086ebd2cd7bf52e0ec5f24c4a916c1
BLAKE2b-256 checksum
How to use checksums
a27aab71412ebe84c8ce4c9d5446cd21fd22ce7b453e1954e988641facb098d6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page