betteroffice-docx
Read, edit, lay out, and rasterize DOCX documents from Python. python-docx
reads a document and writes one back; this also paginates it — page boxes, a
display list, PNG pages — because the Rust
BetterOffice DOCX core is compiled into the wheel:
no Word, no LibreOffice subprocess, no COM.
pip install betteroffice-docx
The distribution is hyphenated, the module is not: import betteroffice_docx.
Read a document
from betteroffice_docx import Document
document = Document.open_path("report.docx")
print(document.structure()) # Structure(body_paragraphs=42, body_tables=3, sections=2)
for paragraph in document:
print(paragraph.id, paragraph.style, repr(paragraph.text))
for table in document.tables():
for row in table.rows:
print([cell.text for cell in row.cells])
paragraphs() walks the body in document order and descends into table cells
and content controls, so a cell paragraph is reachable both ways.
document[key] and document.paragraph(key) take either a w14:paraId or a
body index. Everything a read returns is a value, not a live view — read again
after an edit.
Page geometry is in twips — 1440 to the inch. The module exports
TWIPS_PER_INCH and TWIPS_PER_POINT.
section = document.sections()[0]
print(section.page_width, section.page_height, section.margin_left)
print(document.headers()[0].text)
Edit text
edit = document.replace_text("11111111", "Edited from Python")
print(edit.para_id, edit.start, edit.end)
document.save_path("report-edited.docx")
replace_text rewrites one paragraph and keeps its style, alignment, and the
run formatting it already had. The engine rebuilds the paragraph from a single
run, so a paragraph that mixes runs — half bold, a hyperlink, a field — raises
UnsupportedEditError rather than flattening the formatting you did not ask it
to touch. An unknown w14:paraId raises KeyError.
replace_text uses non-None entries from paragraph_ids. Repeated IDs
receive fresh identities that are saved after editing.
Write
save() serializes the edited model; reopening it returns the edited text.
document = Document.open_path("report.docx")
document.replace_text(document.paragraph_ids[0], "New first line")
print(Document.open(document.save()).paragraph(0).text)
Saving is deterministic. The engine has no clock, so timestamps come from
document.timestamp — the epoch until you set one — and the same input plus the
same edits produce the same bytes. save(now=..., update_modified_date=True, modified_by=...) overrides that for one call.
Saving rebuilds the ZIP container and preserves untouched retained parts.
Export structured content
content = document.export_structured(revision_view="accepted", stories=["body", "comments"])
for block in content["stories"][0]["blocks"]:
print(block["kind"], block["anchor"])
markdown = document.export_markdown(revision_view="markup")
print(markdown["markdown"])
export_structured returns plain dicts in the same camelCase schema the
JavaScript and Rust APIs produce: ordered stories of paragraphs, headings, list
items, tables and content controls, each with the location it was read from, and
diagnostics for everything omitted or not represented. revision_view is
required (accepted, original, or markup); only the body is exported unless
stories selects headers, footers, footnotes, endnotes or comments.
Fields export cached results; images export alt text and relationship metadata.
max_blocks and max_bytes stop the export at a whole block and set
truncated. export_markdown and render_docx_markdown(content) add a
<!-- docx-export:N --> marker per block, mapped to its anchor in anchors.
Content with no location of its own is anchored
{"kind": "unlocated", "story": ..., "reason": ...}, never as a paragraph.
List content controls
for control in document.list_content_controls()["controls"]:
print(control["tag"], control["controlType"], control["value"])
matches = document.find_content_controls({"kind": "tag", "tag": "customer.name"})
Controls come in document order with the same camelCase fields the JavaScript and
Rust APIs produce: the control's id, w:id, type, tag, alias, lock and
placeholder state, whether it is data-bound, its placement, anchor, parent,
current value and effective lock. stories defaults to every category, and
max_controls or max_bytes refuse with ExportError rather than returning a
partial list. Tags, aliases and ooxmlIds match exactly.
Lay a document out
Layout is a two-stage contract. Something else measures text — the browser, or
ooxml-text — and the engine paginates the measured blocks and compiles them
into a display list:
layout = document.layout(open("layout-input.json").read())
print(len(layout), layout.pages)
layout.write("layout.json")
pages = layout.display_list
print(len(pages), pages.primitives)
layout() takes the envelope as a dict or as a JSON string, and returns the
page boxes (layout.json, layout.to_dict()) beside the display list that
paints them.
Rasterize
Register a face for each text family before rasterizing:
from pathlib import Path
document.register_font("Carlito", Path("Carlito-Regular.ttf").read_bytes())
document.register_font("Carlito", Path("Carlito-Bold.ttf").read_bytes(), bold=True)
png = document.render_png(layout.display_list, 0)
png.write("page-0.png")
print(len(png), png.skipped_images)
Text whose family has no chain raises RenderError naming the chain it wanted
— missing font chain for `calibri|0|0` — so a missing face is loud rather
than silently blank.
png.skipped_images counts unresolved references. Register image bytes under
the relationship ID and owning part:
document.register_image(
"rId9", Path("header.png").read_bytes(), scope="header_footer", part="rId7"
)
Images the display list already carries as data: URLs need no registration.
A page past MAX_PIXMAP_DIM per side or MAX_PIXMAP_PIXELS in area is refused
before any surface is allocated.
Compared with python-docx
python-docx |
betteroffice-docx |
|
|---|---|---|
| Read paragraphs, tables, sections | yes | yes |
| Write text back to a file | yes | yes, single-run paragraphs |
| Paginate (page boxes, display list) | no | yes |
| Rasterize pages to PNG | no | yes |
| Engine | pure Python | Rust, compiled |
python-docx is a far broader authoring library. If what you need is
pagination, page images, or an engine that reads what Word actually wrote, that
is the gap this fills.
API
Document.open(data) / open_path(path) |
open from bytes or a path |
document.structure() |
paragraph, table, section, and note counts |
document.paragraphs() / tables() / sections() |
body content |
document.headers() / footers() |
header and footer stories |
document[key] / document.paragraph(key) |
one paragraph by ID or index |
document.paragraph_ids / text |
body IDs, and the whole text |
document.warnings / template_variables |
what the parser found |
document.replace_text(para_id, text) |
rewrite one paragraph |
document.export_structured(...) / export_markdown(...) |
read-only structured content or Markdown |
document.list_content_controls(...) / find_content_controls(query) |
the document's content controls |
render_docx_markdown(content) |
render exported content as Markdown |
document.author / origin / timestamp |
how an edit is attributed and stamped |
document.layout(input) |
paginate a measured envelope |
document.register_font / register_image |
raster resources |
document.render_png(display_list, page) |
rasterize one page |
document.save() / save_path(path) |
serialize to DOCX |
Errors raise DocxError or a more specific subclass: ParseError,
EditError, UnsupportedEditError, LayoutError, RenderError,
ExportError (export or content-control read options the engine refuses; its failure is the refusal
as a dict). An unknown
paragraph ID raises KeyError, an out-of-range index IndexError, and a bad
argument — an unknown parse limit, an unknown image scope, malformed font bytes
— ValueError.
Parser bounds can be tightened for untrusted input:
Document.open_path("untrusted.docx", limits={"max_paragraphs": 5_000, "max_tables": 500})
An unknown limit name raises ValueError rather than being ignored.
Threads
A Document is not pinned to a thread: the engine's document type is Send and
Sync, so opening on one thread and dropping on another is fine. Parsing,
layout, rasterization, and saving release the GIL for their duration, so several
documents genuinely proceed in parallel.
Status
Pre-1.0: the API may change between minor versions. Text replacement preserves paragraph and run formatting.
Wheels are built for Linux (x86_64, aarch64), macOS (arm64, x86_64), and Windows (x86_64) against the stable ABI for CPython 3.9 and up.
Links
- BetterOffice — the project
- Documentation
- Source —
bindings/python-docx - betteroffice-docx on crates.io — the engine this wraps
Apache-2.0.
Metadata
Release files for betteroffice-docx 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| betteroffice_docx-0.2.0.tar.gz | 2.4 MB | Details |
Built distributions (wheels)
| File | Reset | |||
|---|---|---|---|---|
| betteroffice_docx-0.2.0-cp39-abi3-win_amd64.whl | CPython 3.9 | abi3 | Windows x86-64 | Details |
| betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl | CPython 3.9 | abi3 | Linux glibc 2.17+ x86-64 | Details |
| betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl | CPython 3.9 | abi3 | Linux glibc 2.17+ ARM64 | Details |
| betteroffice_docx-0.2.0-cp39-abi3-macosx_11_0_arm64.whl | CPython 3.9 | abi3 | macOS 11.0+ ARM64 | Details |
| betteroffice_docx-0.2.0-cp39-abi3-macosx_10_12_x86_64.whl | CPython 3.9 | abi3 | macOS 10.12+ x86-64 | Details |
Total release size: 37.2 MB
Release files / betteroffice_docx-0.2.0.tar.gz
| Download URL | betteroffice_docx-0.2.0.tar.gz |
|---|---|
| Size | 2.4 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0470807cf448637c3281ed356a709ae6da1c6dd6c11633a04d74aeede5a82b12
|
|
BLAKE2b-256 checksum How to use checksums |
ca3880326ffb70e153e71be8e2bd32cef86739fe8fcfb3c12ab9fc4715a47e43
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / betteroffice_docx-0.2.0-cp39-abi3-win_amd64.whl
| Download URL | betteroffice_docx-0.2.0-cp39-abi3-win_amd64.whl |
|---|---|
| Size | 7.5 MB |
| Tags | CPython 3.9 Windows x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
23b4557147dff0018da311b893ae2449e2d037a1105393162c8f17d14ad87290
|
|
BLAKE2b-256 checksum How to use checksums |
52a053eb808800dc3f9bdde7f0216ed940fc0c30c5f2b7cdd6fa8746e9a7df8b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
| Download URL | betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl |
|---|---|
| Size | 7.1 MB |
| Tags | CPython 3.9 Linux glibc 2.17+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
9f946e2f59f301b5318774ca540456f751a448f99c42cca8e4c9db58278afa90
|
|
BLAKE2b-256 checksum How to use checksums |
2210a0d628769d565a15b943e2d9a2a16bbfe8eb439d3fb05e4949cd1c0ab922
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
| Download URL | betteroffice_docx-0.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl |
|---|---|
| Size | 6.7 MB |
| Tags | CPython 3.9 Linux glibc 2.17+ ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
a17d937b357eeaff6167a7f156dbfa6a0ac738d2e637e4c4f169d53da7cad7c1
|
|
BLAKE2b-256 checksum How to use checksums |
e2ab0008800d8d944dc03dba0665d296114a38f6e212590f1e2746ddff865e23
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / betteroffice_docx-0.2.0-cp39-abi3-macosx_11_0_arm64.whl
| Download URL | betteroffice_docx-0.2.0-cp39-abi3-macosx_11_0_arm64.whl |
|---|---|
| Size | 6.6 MB |
| Tags | CPython 3.9 abi3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
9f2d22f8fdc778eaa962d5ba6b4773ec9ef54e3f536000d8e5fcb3f8936ca360
|
|
BLAKE2b-256 checksum How to use checksums |
50fe7f22fa348c3c9843336d2b392ef0b4af64b1013b6943db0b416dcfed244f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / betteroffice_docx-0.2.0-cp39-abi3-macosx_10_12_x86_64.whl
| Download URL | betteroffice_docx-0.2.0-cp39-abi3-macosx_10_12_x86_64.whl |
|---|---|
| Size | 7.0 MB |
| Tags | CPython 3.9 abi3 macOS 10.12+ x86-64 |
|
SHA-256 checksum How to use checksums |
2aa4a27cdc34c7b5f198ce0916ae84e50265cb13301eb7068e064b379573807d
|
|
BLAKE2b-256 checksum How to use checksums |
ee37de2f5267014a67acccd4fc3ef439ddc29e01a7a825649a6eac28508c249f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log