Skip to main content

docspine

PyPI

A pure-Rust Word (.docx) parser with Python bindings (PyO3 / maturin, abi3-py311). A .docx file is OOXML — a zip archive of XML parts — and docspine walks word/document.xml directly to produce a structured, information-preserving model: paragraphs (styled runs), tables (rows, cells, merges, fills, nesting), and embedded pictures. Tables are a first-class focus. Embedded images can additionally be OCR'd locally, offline, and deterministically via the sibling ocrspine crate (PP-OCRv5 through tract-onnx — no cloud, no network), and an image that is a table can be reconstructed into a grid from its OCR word boxes. A parsed document can also be exported to PDF (to_pdf() / save_pdf()) with flowed layout and pagination through the shared pure-Rust pdf-typeset engine from pdfspine — no LibreOffice, no cloud converter.

docspine is the document-engine sibling of pdfspine (PDF) and pptspine (PowerPoint), all sharing the same ocrspine OCR core.

Capabilities

Area Status
Body blocks: paragraphs + tables in document order parsed
Paragraphs: runs, text, style name, alignment, list level parsed
Run styling: font, size, bold, italic, underline, color parsed
Tables: rows, cells, cell paragraphs parsed
Table merges: gridSpan (horizontal) parsed
Table merges: vMerge restart / continue (vertical) parsed
Nested tables (a table inside a cell) parsed
Cell shading/fill, cell width (dxa), table grid columns parsed
Row height, header rows parsed
Embedded pictures: r:embed rel → media name + raw bytes + EMU extent parsed
Image OCR (embedded pictures → words + boxes) working (ocr_image)
Image-table reconstruction from OCR boxes → grid working (reconstruct_image_table)
PDF export: to_pdf() / save_pdf() — flowed layout + pagination; per-section page geometry (sectPr), styles.xml + theme effective styles, numbering engine, table fidelity (borders/merges/margins, cross-page; vAlign top-only), inline/anchored images, defaultTabStop tab advance working
Legacy binary .doc (OLE/CFB) probe + typed downgrade (full body deferred)

Parsing is tolerant: unknown elements are skipped, missing attributes become None, and malformed input yields a typed DocError rather than a panic.

docx first; legacy .doc deferred

Modern .docx (OOXML) is the primary target. The old binary .doc is a Microsoft compound document (OLE/CFB): rebuilding its body from the binary FIB + piece table is large, fiddly, and shares almost nothing with the docx path. So docspine ships detection + a clean typed downgrade today (a .doc byte stream yields DocUnsupportedError, and probe_doc reports the CFB streams when built with the legacy-doc feature); full .doc body reconstruction is a follow-up, not a blocker.

Install

pip install docspine

docspine is on PyPI. OCR works out of the box: the PP-OCRv5 weights ship in the shared ocrspine-models data package — a runtime dependency pip pulls in automatically — so a plain pip install docspine finds the OCR models with no extra setup (it no longer needs a sibling ../ocrspine/models checkout or OCRSPINE_MODELS). To build from source instead, see below.

Build (from the package root)

uv venv .venv
VIRTUAL_ENV="$(pwd)/.venv" uv pip install maturin pytest
# Structural parsing needs no models. The OCR path resolves models from a
# sibling ../ocrspine/models by default (or set OCRSPINE_MODELS).
OCRSPINE_MODELS="$(cd ../ocrspine && pwd)/models" \
  VIRTUAL_ENV="$(pwd)/.venv" .venv/bin/maturin develop --release

Use from Python

import docspine

doc = docspine.open("report.docx")
print(doc.block_count)

for block in doc.body():            # list[dict], introspectable
    if block["kind"] == "paragraph":
        for run in block["runs"]:
            print(run["text"], run["bold"], run["color"])
    elif block["kind"] == "table":
        for row in block["rows"]:
            for cell in row["cells"]:
                print(cell["text"], "span", cell["grid_span"], cell["v_merge"])

# Run OCR on raw image bytes (PNG/JPEG), offline:
items = docspine.ocr_image(open("scan.png", "rb").read())
print(" ".join(i["text"] for i in items))

# Reconstruct a table that lives inside an image into a grid:
for table in docspine.reconstruct_image_table(open("table.png", "rb").read()):
    for cell in table["cells"]:
        print(cell["row"], cell["col"], cell["text"])

Export to PDF

doc = docspine.open("report.docx")
doc.save_pdf("report.pdf")         # flowed layout + pagination
pdf_bytes = doc.to_pdf()           # or in-memory bytes

# Optional: map a requested font family to a local font file (or to another
# installed family), layered on top of the built-in substitution table:
doc.save_pdf("report.pdf", font_map={"Calibri": "/path/to/Carlito.ttf"})

Per-section page geometry from sectPr is honored (paper size, orientation, margins per section). Rendering is deterministic and fully offline. Missing fonts degrade gracefully: an available face is substituted and a Python UserWarning is emitted once per warning kind — the export never fails on a missing font.

Rust workspace

crates/
  doc-core    domain model + geometry (twip/EMU) + typed DocError. No IO/zip/XML.
  doc-parse   OOXML reader: zip extract + quick-xml walk -> Document.
              Legacy binary .doc probing behind the `legacy-doc` feature.
  doc-ocr     image-OCR bridge over ocrspine (PaddleOcr) + image-table reconstruction.
  doc-render  docx -> PDF renderer over the shared pdf-typeset engine (from pdfspine).
  py-bindings PyO3 _core extension (the FFI chokepoint); `ocr` feature gates OCR.

Deferred / follow-up

  • Full legacy binary .doc (OLE/CFB / [MS-DOC]) body reconstruction (FIB, piece table, CHPX/PAPX). Today: detection + typed downgrade + probe_doc.
  • Richer styling, headers/footers, footnotes/endnotes, comments, fields, hyperlinks targets, SmartArt/charts.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

docspine-0.3.0.tar.gz (153.6 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

docspine-0.3.0-cp311-abi3-win_amd64.whl (10.4 MB view details)

Uploaded CPython 3.11+Windows x86-64

docspine-0.3.0-cp311-abi3-manylinux_2_28_aarch64.whl (9.5 MB view details)

Uploaded CPython 3.11+manylinux: glibc 2.28+ ARM64

docspine-0.3.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (10.5 MB view details)

Uploaded CPython 3.11+manylinux: glibc 2.17+ x86-64

docspine-0.3.0-cp311-abi3-macosx_11_0_arm64.whl (9.2 MB view details)

Uploaded CPython 3.11+macOS 11.0+ ARM64

docspine-0.3.0-cp311-abi3-macosx_10_12_x86_64.whl (10.1 MB view details)

Uploaded CPython 3.11+macOS 10.12+ x86-64

File details

Details for the file docspine-0.3.0.tar.gz.

File metadata

  • Download URL: docspine-0.3.0.tar.gz
  • Upload date:
  • Size: 153.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for docspine-0.3.0.tar.gz
Algorithm Hash digest
SHA256 502f405c2ca4fbff914a8224c59c7da55a6b44a72a45f944920266e77fed0161
MD5 c2305251985927e8fd3414c2b2201051
BLAKE2b-256 e6c27b0007cc2b6a3772ad01d55d1f40f218ef33b2920bc3c5e3437850929fd0

See more details on using hashes here.

File details

Details for the file docspine-0.3.0-cp311-abi3-win_amd64.whl.

File metadata

  • Download URL: docspine-0.3.0-cp311-abi3-win_amd64.whl
  • Upload date:
  • Size: 10.4 MB
  • Tags: CPython 3.11+, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for docspine-0.3.0-cp311-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 9c077852b497573c9ddc471f2e2ff77be326358b9de50b6165c2667804cfe9d4
MD5 c22216167482314a18436465fabe464c
BLAKE2b-256 cd2139d1122f22ff38e7ae16d9410aac281e7e905d30a3c4a3da5b44d9938c26

See more details on using hashes here.

File details

Details for the file docspine-0.3.0-cp311-abi3-manylinux_2_28_aarch64.whl.

File metadata

File hashes

Hashes for docspine-0.3.0-cp311-abi3-manylinux_2_28_aarch64.whl
Algorithm Hash digest
SHA256 6bb875723b7e3cf2905775f54f2258bc6a81a50d2a9e7246f56ea879d6188ee7
MD5 9898fc2f637c5ec36300a30d788e0eac
BLAKE2b-256 dda08d7f76f4c99e55c5d8952093cf162dcfb53313323f548c82790a3608e20c

See more details on using hashes here.

File details

Details for the file docspine-0.3.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for docspine-0.3.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 92a6d60432c760020a338408c274f9a74ebbad2980283d7a828771132a51a085
MD5 7133ffbb7b2944f79b53a2ba3bcc1833
BLAKE2b-256 eff37d6f77fca889e1e5c5acc4392a1397c6d27aded706d6eedaa0f65be24bbc

See more details on using hashes here.

File details

Details for the file docspine-0.3.0-cp311-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for docspine-0.3.0-cp311-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 cbd4bfcae1f68102248d7e342719ce1080c7d1500453998336f7ce2f7bbf8eb0
MD5 c4998cc32c701f5a0964dba4e4842fcb
BLAKE2b-256 62a845dc308aebd9da75612028b9ee867f3beb2cb35ecb853f834ac99ef9e064

See more details on using hashes here.

File details

Details for the file docspine-0.3.0-cp311-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for docspine-0.3.0-cp311-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 cf698e67f9cbc754d2d14a5edd6a392a0d4aaff61aa7b1f91352c26746dbe64a
MD5 6583d609c9ac11d74fbf33a60b9853ad
BLAKE2b-256 8fe2665004b8a30333944be6e426add61afc4c29c503d0c28d7227ba0a4c7bf8

See more details on using hashes here.

Release history Release notifications | RSS feed

0.5.1

6 files

0.5.0

6 files

0.4.0

6 files

This release

0.3.0 This release

6 files

0.2.0

6 files

0.1.1

6 files

0.1.0

5 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page