Skip to main content

betteroffice-docx

Read, edit, lay out, and rasterize DOCX documents from Python. python-docx reads a document and writes one back; this also paginates it — page boxes, a display list, PNG pages — because the Rust BetterOffice DOCX core is compiled into the wheel: no Word, no LibreOffice subprocess, no COM.

pip install betteroffice-docx

The distribution is hyphenated, the module is not: import betteroffice_docx.

Read a document

from betteroffice_docx import Document

document = Document.open_path("report.docx")

print(document.structure())          # Structure(body_paragraphs=42, body_tables=3, sections=2)
for paragraph in document:
    print(paragraph.id, paragraph.style, repr(paragraph.text))

for table in document.tables():
    for row in table.rows:
        print([cell.text for cell in row.cells])

paragraphs() walks the body in document order and descends into table cells and content controls, so a cell paragraph is reachable both ways. document[key] and document.paragraph(key) take either a w14:paraId or a body index. Everything a read returns is a value, not a live view — read again after an edit.

Page geometry is in twips — 1440 to the inch. The module exports TWIPS_PER_INCH and TWIPS_PER_POINT.

section = document.sections()[0]
print(section.page_width, section.page_height, section.margin_left)
print(document.headers()[0].text)

Edit text

edit = document.replace_text("11111111", "Edited from Python")
print(edit.para_id, edit.start, edit.end)

document.save_path("report-edited.docx")

replace_text rewrites one paragraph and keeps its style, alignment, and the run formatting it already had. The engine rebuilds the paragraph from a single run, so a paragraph that mixes runs — half bold, a hyperlink, a field — raises UnsupportedEditError rather than flattening the formatting you did not ask it to touch. An unknown w14:paraId raises KeyError.

replace_text uses non-None entries from paragraph_ids. Repeated IDs receive fresh identities that are saved after editing.

Write

save() serializes the edited model; reopening it returns the edited text.

document = Document.open_path("report.docx")
document.replace_text(document.paragraph_ids[0], "New first line")
print(Document.open(document.save()).paragraph(0).text)

Saving is deterministic. The engine has no clock, so timestamps come from document.timestamp — the epoch until you set one — and the same input plus the same edits produce the same bytes. save(now=..., update_modified_date=True, modified_by=...) overrides that for one call.

Saving rebuilds the ZIP container and preserves untouched retained parts.

Export structured content

content = document.export_structured(revision_view="accepted", stories=["body", "comments"])
for block in content["stories"][0]["blocks"]:
    print(block["kind"], block["anchor"])

markdown = document.export_markdown(revision_view="markup")
print(markdown["markdown"])

export_structured returns plain dicts in the same camelCase schema the JavaScript and Rust APIs produce: ordered stories of paragraphs, headings, list items, tables and content controls, each with the location it was read from, and diagnostics for everything omitted or not represented. revision_view is required (accepted, original, or markup); only the body is exported unless stories selects headers, footers, footnotes, endnotes or comments. Fields export cached results; images export alt text and relationship metadata. max_blocks and max_bytes stop the export at a whole block and set truncated. export_markdown and render_docx_markdown(content) add a <!-- docx-export:N --> marker per block, mapped to its anchor in anchors. Content with no location of its own is anchored {"kind": "unlocated", "story": ..., "reason": ...}, never as a paragraph.

List content controls

for control in document.list_content_controls()["controls"]:
    print(control["tag"], control["controlType"], control["value"])

matches = document.find_content_controls({"kind": "tag", "tag": "customer.name"})

Controls come in document order with the same camelCase fields the JavaScript and Rust APIs produce: the control's id, w:id, type, tag, alias, lock and placeholder state, whether it is data-bound, its placement, anchor, parent, current value and effective lock. stories defaults to every category, and max_controls or max_bytes refuse with ExportError rather than returning a partial list. Tags, aliases and ooxmlIds match exactly.

Lay a document out

Layout is a two-stage contract. Something else measures text — the browser, or ooxml-text — and the engine paginates the measured blocks and compiles them into a display list:

layout = document.layout(open("layout-input.json").read())
print(len(layout), layout.pages)
layout.write("layout.json")

pages = layout.display_list
print(len(pages), pages.primitives)

layout() takes the envelope as a dict or as a JSON string, and returns the page boxes (layout.json, layout.to_dict()) beside the display list that paints them.

Rasterize

Register a face for each text family before rasterizing:

from pathlib import Path

document.register_font("Carlito", Path("Carlito-Regular.ttf").read_bytes())
document.register_font("Carlito", Path("Carlito-Bold.ttf").read_bytes(), bold=True)

png = document.render_png(layout.display_list, 0)
png.write("page-0.png")
print(len(png), png.skipped_images)

Text whose family has no chain raises RenderError naming the chain it wanted — missing font chain for `calibri|0|0` — so a missing face is loud rather than silently blank.

png.skipped_images counts unresolved references. Register image bytes under the relationship ID and owning part:

document.register_image(
    "rId9", Path("header.png").read_bytes(), scope="header_footer", part="rId7"
)

Images the display list already carries as data: URLs need no registration. A page past MAX_PIXMAP_DIM per side or MAX_PIXMAP_PIXELS in area is refused before any surface is allocated.

Compared with python-docx

python-docx betteroffice-docx
Read paragraphs, tables, sections yes yes
Write text back to a file yes yes, single-run paragraphs
Paginate (page boxes, display list) no yes
Rasterize pages to PNG no yes
Engine pure Python Rust, compiled

python-docx is a far broader authoring library. If what you need is pagination, page images, or an engine that reads what Word actually wrote, that is the gap this fills.

API

Document.open(data) / open_path(path) open from bytes or a path
document.structure() paragraph, table, section, and note counts
document.paragraphs() / tables() / sections() body content
document.headers() / footers() header and footer stories
document[key] / document.paragraph(key) one paragraph by ID or index
document.paragraph_ids / text body IDs, and the whole text
document.warnings / template_variables what the parser found
document.replace_text(para_id, text) rewrite one paragraph
document.export_structured(...) / export_markdown(...) read-only structured content or Markdown
document.list_content_controls(...) / find_content_controls(query) the document's content controls
render_docx_markdown(content) render exported content as Markdown
document.author / origin / timestamp how an edit is attributed and stamped
document.layout(input) paginate a measured envelope
document.register_font / register_image raster resources
document.render_png(display_list, page) rasterize one page
document.save() / save_path(path) serialize to DOCX

Errors raise DocxError or a more specific subclass: ParseError, EditError, UnsupportedEditError, LayoutError, RenderError, ExportError (export or content-control read options the engine refuses; its failure is the refusal as a dict). An unknown paragraph ID raises KeyError, an out-of-range index IndexError, and a bad argument — an unknown parse limit, an unknown image scope, malformed font bytes — ValueError.

Parser bounds can be tightened for untrusted input:

Document.open_path("untrusted.docx", limits={"max_paragraphs": 5_000, "max_tables": 500})

An unknown limit name raises ValueError rather than being ignored.

Threads

A Document is not pinned to a thread: the engine's document type is Send and Sync, so opening on one thread and dropping on another is fine. Parsing, layout, rasterization, and saving release the GIL for their duration, so several documents genuinely proceed in parallel.

Status

Pre-1.0: the API may change between minor versions. Text replacement preserves paragraph and run formatting.

Wheels are built for Linux (x86_64, aarch64), macOS (arm64, x86_64), and Windows (x86_64) against the stable ABI for CPython 3.9 and up.

Apache-2.0.

Metadata

Release files for betteroffice-docx 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for betteroffice-docx 0.2.0
File Size Uploaded
betteroffice_docx-0.2.0.tar.gz 2.4 MB Details

Built distributions (wheels)

Table of built distributions (wheels) for betteroffice-docx 0.2.0
File
betteroffice_docx-0.2.0-cp39-abi3-win_amd64.whl CPython 3.9 abi3 Windows x86-64 Details
betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.9 abi3 Linux glibc 2.17+ x86-64 Details
betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl CPython 3.9 abi3 Linux glibc 2.17+ ARM64 Details
betteroffice_docx-0.2.0-cp39-abi3-macosx_11_0_arm64.whl CPython 3.9 abi3 macOS 11.0+ ARM64 Details
betteroffice_docx-0.2.0-cp39-abi3-macosx_10_12_x86_64.whl CPython 3.9 abi3 macOS 10.12+ x86-64 Details

Total release size: 37.2 MB

Release files / betteroffice_docx-0.2.0.tar.gz

Download URL betteroffice_docx-0.2.0.tar.gz
Size 2.4 MB
Tags Source
SHA-256 checksum
How to use checksums
0470807cf448637c3281ed356a709ae6da1c6dd6c11633a04d74aeede5a82b12
BLAKE2b-256 checksum
How to use checksums
ca3880326ffb70e153e71be8e2bd32cef86739fe8fcfb3c12ab9fc4715a47e43
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / betteroffice_docx-0.2.0-cp39-abi3-win_amd64.whl

Download URL betteroffice_docx-0.2.0-cp39-abi3-win_amd64.whl
Size 7.5 MB
Tags CPython 3.9 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
23b4557147dff0018da311b893ae2449e2d037a1105393162c8f17d14ad87290
BLAKE2b-256 checksum
How to use checksums
52a053eb808800dc3f9bdde7f0216ed940fc0c30c5f2b7cdd6fa8746e9a7df8b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 7.1 MB
Tags CPython 3.9 Linux glibc 2.17+ x86-64 abi3
SHA-256 checksum
How to use checksums
9f946e2f59f301b5318774ca540456f751a448f99c42cca8e4c9db58278afa90
BLAKE2b-256 checksum
How to use checksums
2210a0d628769d565a15b943e2d9a2a16bbfe8eb439d3fb05e4949cd1c0ab922
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl

Download URL betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Size 6.7 MB
Tags CPython 3.9 Linux glibc 2.17+ ARM64 abi3
SHA-256 checksum
How to use checksums
a17d937b357eeaff6167a7f156dbfa6a0ac738d2e637e4c4f169d53da7cad7c1
BLAKE2b-256 checksum
How to use checksums
e2ab0008800d8d944dc03dba0665d296114a38f6e212590f1e2746ddff865e23
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / betteroffice_docx-0.2.0-cp39-abi3-macosx_11_0_arm64.whl

Download URL betteroffice_docx-0.2.0-cp39-abi3-macosx_11_0_arm64.whl
Size 6.6 MB
Tags CPython 3.9 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
9f2d22f8fdc778eaa962d5ba6b4773ec9ef54e3f536000d8e5fcb3f8936ca360
BLAKE2b-256 checksum
How to use checksums
50fe7f22fa348c3c9843336d2b392ef0b4af64b1013b6943db0b416dcfed244f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / betteroffice_docx-0.2.0-cp39-abi3-macosx_10_12_x86_64.whl

Download URL betteroffice_docx-0.2.0-cp39-abi3-macosx_10_12_x86_64.whl
Size 7.0 MB
Tags CPython 3.9 abi3 macOS 10.12+ x86-64
SHA-256 checksum
How to use checksums
2aa4a27cdc34c7b5f198ce0916ae84e50265cb13301eb7068e064b379573807d
BLAKE2b-256 checksum
How to use checksums
ee37de2f5267014a67acccd4fc3ef439ddc29e01a7a825649a6eac28508c249f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

6 release files

0.1.0

6 release files

0.0.2

6 release files

0.0.1

6 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page