Skip to main content

docspine

PyPI

A pure-Rust Word (.docx) parser with Python bindings (PyO3 / maturin, abi3-py311). A .docx file is OOXML — a zip archive of XML parts — and docspine walks word/document.xml directly to produce a structured, information-preserving model: paragraphs (styled runs), tables (rows, cells, merges, fills, nesting), and embedded pictures. Tables are a first-class focus. Embedded images can additionally be OCR'd locally, offline, and deterministically via the sibling ocrspine crate (PP-OCRv5 through tract-onnx — no cloud, no network), and an image that is a table can be reconstructed into a grid from its OCR word boxes. A parsed document can also be exported to PDF (to_pdf() / save_pdf()) with flowed layout and pagination through the shared pure-Rust pdf-typeset engine from pdfspine — no LibreOffice, no cloud converter.

docspine is the document-engine sibling of pdfspine (PDF) and pptspine (PowerPoint), all sharing the same ocrspine OCR core.

Capabilities

Area Status
Body blocks: paragraphs + tables in document order parsed
Paragraphs: runs, text, style name, alignment, list level parsed
Run styling: font, size, bold, italic, underline, color parsed
Tables: rows, cells, cell paragraphs parsed
Table merges: gridSpan (horizontal) parsed
Table merges: vMerge restart / continue (vertical) parsed
Nested tables (a table inside a cell) parsed
Cell shading/fill, cell width (dxa), table grid columns parsed
Row height, header rows parsed
Embedded pictures: r:embed rel → media name + raw bytes + EMU extent parsed
Image OCR (embedded pictures → words + boxes) working (ocr_image)
Image-table reconstruction from OCR boxes → grid working (reconstruct_image_table)
PDF export: to_pdf() / save_pdf() — flowed layout + pagination; per-section page geometry (sectPr), styles.xml + theme effective styles, numbering engine, table fidelity (borders/merges/margins, cross-page; cell vAlign top/center/bottom), paragraph borders/shading, inline images + absolutely-positioned anchored images (no text wrap), hyperlinks as PDF link annotations, defaultTabStop tab advance working
Legacy binary .doc (OLE/CFB) probe + typed downgrade (full body deferred)

Parsing is tolerant: unknown elements are skipped, missing attributes become None, and malformed input yields a typed DocError rather than a panic.

docx first; legacy .doc deferred

Modern .docx (OOXML) is the primary target. The old binary .doc is a Microsoft compound document (OLE/CFB): rebuilding its body from the binary FIB + piece table is large, fiddly, and shares almost nothing with the docx path. So docspine ships detection + a clean typed downgrade today (a .doc byte stream yields DocUnsupportedError, and probe_doc reports the CFB streams when built with the legacy-doc feature); full .doc body reconstruction is a follow-up, not a blocker.

Install

pip install docspine

docspine is on PyPI. OCR works out of the box: the PP-OCRv5 weights ship in the shared ocrspine-models data package — a runtime dependency pip pulls in automatically — so a plain pip install docspine finds the OCR models with no extra setup (it no longer needs a sibling ../ocrspine/models checkout or OCRSPINE_MODELS). To build from source instead, see below.

Build (from the package root)

uv venv .venv
VIRTUAL_ENV="$(pwd)/.venv" uv pip install maturin pytest
# Structural parsing needs no models. The OCR path resolves models from a
# sibling ../ocrspine/models by default (or set OCRSPINE_MODELS).
OCRSPINE_MODELS="$(cd ../ocrspine && pwd)/models" \
  VIRTUAL_ENV="$(pwd)/.venv" .venv/bin/maturin develop --release

Use from Python

import docspine

doc = docspine.open("report.docx")
print(doc.block_count)

for block in doc.body():            # list[dict], introspectable
    if block["kind"] == "paragraph":
        for run in block["runs"]:
            print(run["text"], run["bold"], run["color"])
    elif block["kind"] == "table":
        for row in block["rows"]:
            for cell in row["cells"]:
                print(cell["text"], "span", cell["grid_span"], cell["v_merge"])

# Run OCR on raw image bytes (PNG/JPEG), offline:
items = docspine.ocr_image(open("scan.png", "rb").read())
print(" ".join(i["text"] for i in items))

# Reconstruct a table that lives inside an image into a grid:
for table in docspine.reconstruct_image_table(open("table.png", "rb").read()):
    for cell in table["cells"]:
        print(cell["row"], cell["col"], cell["text"])

Export to PDF

doc = docspine.open("report.docx")
doc.save_pdf("report.pdf")         # flowed layout + pagination
pdf_bytes = doc.to_pdf()           # or in-memory bytes

# Optional: map a requested font family to a local font file (or to another
# installed family), layered on top of the built-in substitution table:
doc.save_pdf("report.pdf", font_map={"Calibri": "/path/to/Carlito.ttf"})

Per-section page geometry from sectPr is honored (paper size, orientation, margins per section). Rendering is deterministic and fully offline. Missing fonts degrade gracefully: an available face is substituted and a Python UserWarning is emitted once per warning kind — the export never fails on a missing font.

Rust workspace

crates/
  doc-core    domain model + geometry (twip/EMU) + typed DocError. No IO/zip/XML.
  doc-parse   OOXML reader: zip extract + quick-xml walk -> Document.
              Legacy binary .doc probing behind the `legacy-doc` feature.
  doc-ocr     image-OCR bridge over ocrspine (PaddleOcr) + image-table reconstruction.
  doc-render  docx -> PDF renderer over the shared pdf-typeset engine (from pdfspine).
  py-bindings PyO3 _core extension (the FFI chokepoint); `ocr` feature gates OCR.

Deferred / follow-up

  • Full legacy binary .doc (OLE/CFB / [MS-DOC]) body reconstruction (FIB, piece table, CHPX/PAPX). Today: detection + typed downgrade + probe_doc.
  • Richer styling, headers/footers, footnotes/endnotes, comments, fields, hyperlinks targets, SmartArt/charts.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

docspine-0.4.0.tar.gz (158.7 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

docspine-0.4.0-cp311-abi3-win_amd64.whl (10.4 MB view details)

Uploaded CPython 3.11+Windows x86-64

docspine-0.4.0-cp311-abi3-manylinux_2_28_aarch64.whl (9.5 MB view details)

Uploaded CPython 3.11+manylinux: glibc 2.28+ ARM64

docspine-0.4.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (10.5 MB view details)

Uploaded CPython 3.11+manylinux: glibc 2.17+ x86-64

docspine-0.4.0-cp311-abi3-macosx_11_0_arm64.whl (9.2 MB view details)

Uploaded CPython 3.11+macOS 11.0+ ARM64

docspine-0.4.0-cp311-abi3-macosx_10_12_x86_64.whl (10.1 MB view details)

Uploaded CPython 3.11+macOS 10.12+ x86-64

File details

Details for the file docspine-0.4.0.tar.gz.

File metadata

  • Download URL: docspine-0.4.0.tar.gz
  • Upload date:
  • Size: 158.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for docspine-0.4.0.tar.gz
Algorithm Hash digest
SHA256 bd40a16cd41d39090bc05961005f325c0f67520898daf1a9055b1916ca59d982
MD5 3227ab3068749db66eae160284bbf997
BLAKE2b-256 4a49e87b822b88c1299dd5904d583deec1b32d4ab5d9db61e842abcaf6a53b2f

See more details on using hashes here.

File details

Details for the file docspine-0.4.0-cp311-abi3-win_amd64.whl.

File metadata

  • Download URL: docspine-0.4.0-cp311-abi3-win_amd64.whl
  • Upload date:
  • Size: 10.4 MB
  • Tags: CPython 3.11+, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for docspine-0.4.0-cp311-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 6ef7a647378ff6843db7a99380779fd92d35457fd8f870d6af9e6cf147e85638
MD5 7c794ab6e62e2c55f7895c06e7f542db
BLAKE2b-256 fb89853ee2e8a07f438af74ba7572105d4381671ed65d548567d1b42cfd549ae

See more details on using hashes here.

File details

Details for the file docspine-0.4.0-cp311-abi3-manylinux_2_28_aarch64.whl.

File metadata

File hashes

Hashes for docspine-0.4.0-cp311-abi3-manylinux_2_28_aarch64.whl
Algorithm Hash digest
SHA256 8c8863ee2c8bce309c7c368ea82602227f412958ad573597f33a9805b27354d0
MD5 52ce07f5dd38040772d1dea475b51d24
BLAKE2b-256 aa9b83497fac66cd4dc11316fd36e1b99e2ff5eefea46132e87034522378ad4f

See more details on using hashes here.

File details

Details for the file docspine-0.4.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for docspine-0.4.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 a57ea08ce897c7344f08603ef84b5d01ae7fbe5e0fbbec91ac21d2935b7ff784
MD5 09494761a669a41aa36a6651453f66f5
BLAKE2b-256 d2fe542d7d73b5b63216e8463394ce7e0b62e53a525c8a1c64154c3c9bae2b98

See more details on using hashes here.

File details

Details for the file docspine-0.4.0-cp311-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for docspine-0.4.0-cp311-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 fa1ed8e08232a4f41a15af912cdbf312596f5f3fe58cf046b4f6ea197d4b6dfa
MD5 570e665c6d231090cc9483dd630606c2
BLAKE2b-256 02b7422bfc3dfdd3a0e32544075f0cfa49e36875f35a71e9b851f987457671b4

See more details on using hashes here.

File details

Details for the file docspine-0.4.0-cp311-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for docspine-0.4.0-cp311-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 3bc1ad44d3fd8698e88d1024ea4a8170ca5ac013cb6a9489ffd65f5d3167b611
MD5 7e98b21481409a3ff08b31558843e17d
BLAKE2b-256 12311e08a767326c4bafa8bbee59ec9bf7c41dbbec2f8008f58ce674ff49a05f

See more details on using hashes here.

Release history Release notifications | RSS feed

0.5.1

6 files

0.5.0

6 files

This release

0.4.0 This release

6 files

0.3.0

6 files

0.2.0

6 files

0.1.1

6 files

0.1.0

5 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page