An import-compatible, agent-safe fork of python-docx designed to prevent silent corruption when editing existing Word documents.
paper-docx is a strict-superset hard fork of python-docx for safely inspecting, editing, reviewing, and composing existing Microsoft Word (.docx) documents. It keeps python-docx's package layer, XML mapping, and object model. It adds typed inspection of what a document actually contains, edits that survive Word's run fragmentation, and refusals in place of guesses.
import docx # the import name is unchanged — see "Drop-in by design"
Every added operation either does exactly what it claims or refuses atomically. Gates such as Restrict-Editing are checked before anything is touched, and compound edits capture the live package first and restore it if they raise. A refusal raises a typed PaperRefusal and leaves the document byte-for-byte unchanged in memory and on disk.
Why paper-docx exists
python-docx is excellent at creating documents. Its lossless package layer, disciplined XML mapping, and years of absorbed edge cases are why this fork builds on it.
The harder problem is changing a contract or other real-world document without losing formatting, revisions, fields, or content outside the body. Hand-edited XML can produce silent corruption: a file that opens fine and is quietly wrong. An agent cannot eyeball the result, so it needs the document's structure and every edit outcome as typed, machine-readable data. It also needs the library to refuse rather than guess.
Quick start
Create a native Word redline from two document versions:
import tempfile, docx
from docx.package import compare
tmp = tempfile.mkdtemp()
a = docx.Document()
a.add_paragraph("Payment is due within thirty calendar days of the invoice date.")
a.save(f"{tmp}/v1.docx")
b = docx.Document()
b.add_paragraph("Payment is due within thirty business days of the invoice date.")
b.save(f"{tmp}/v2.docx")
result = compare(f"{tmp}/v1.docx", f"{tmp}/v2.docx", author="Reviewer")
[(r.revision_type, r.text) for r in result.document.revisions]
# [('deletion', 'calendar'), ('insertion', 'business')]
result.document.revisions.accept_all()
result.document.paragraphs[0].text
# 'Payment is due within thirty business days of the invoice date.'
compare emits markup Word renders as tracked changes. Before returning, it accepts and rejects private copies and verifies both outcomes. If a difference cannot be represented safely as a redline, such as a style or package-part change, it raises a typed refusal instead of returning an incomplete result.
What paper-docx adds
Reading and editing one document
docx.storytraverses the body, headers, footers, footnotes, endnotes, comments, tracked insertions, content controls, and text boxes. Callers can view the document as it stands, before pending revisions, or all at once.docx.searchfinds exact text by default across Word's run fragmentation, with explicit normalized matching when wanted. A returnedSpancan replace the matched text while preserving unaffected runs, emit the replacement as a tracked change, or anchor a comment.docx.blocksinserts, deletes, or replaces whole paragraphs relative to an exact string, live span, or owner-bound live paragraph block, as plain edits or as a tracked change.docx.tableops/docx.numberingprovide explicit exact-or-normalized top-level table lookup plus cell, row, and list edits. Table lookup never synthesizes a match across cells, and copied rows accept only simple uniform templates rather than flattening ambiguous or unsupported structures.docx.controlsfills content controls with the correct value type and clears placeholder state so Word treats them as filled.docx.bookmarks/docx.fieldscreate bookmarks over a span and author page numbers, dates, cross-references, captions, and tables of contents as fields with placeholder results.docx.notes/docx.linksanchor real footnotes, endnotes, and hyperlinks to a matched span, so they compose withdocx.searchrather than needing their own targeting.Drawing.replace_pictureswaps the image behind a drawing in place, leaving its size, position, and identity alone.docx.formattingresolves effective formatting through document defaults, styles, and direct formatting, with provenance for each value.
Selection and preservation
Search is exact by default across Word's ordinary run fragmentation; normalized
matching is an explicit policy. find_text(..., near=...) orders the complete
candidate set for inspection; find_one() retains ordinary zero/one/many
resolution and accepts no contextual-ranking keyword. A live Span identifies
selected characters and a live Block identifies one attached document element.
Serialized Anchor values are location evidence, not mutation authority; after a
reload, reacquire a live target. Operations that require one paragraph refuse
cross-paragraph targets. See the
search, story, and
block references for the complete contracts.
Ordinary untracked replacement considers every maximal exact prefix/suffix
alignment. It preserves unchanged affixes and edits one proved structural
region only when those alignments agree on the changed interval and writable
destination; ambiguous repeated affixes or insertion boundaries refuse with
guidance to re-find the intended substring. Mixed formatting across the
consumed text is not itself a refusal: nonempty replacement text takes the
complete direct w:rPr of the run holding the first consumed character, and
the differently formatted runs it consumes collapse into that format.
Untouched affixes and boundary fragments keep their own runs. Crossing
distinct inline wrapper owners, or moving a marker, still refuses. Every
successful text-changing direct Span.replace() consumes that span; re-find
before another operation. Direct no-ops and atomically refused or rolled-back
operations leave the supplied span reusable. There is no separate
topology-preservation mode or character-capacity allocator. Formatting inspection
composes an explicit search with format_of(span), or passes a run or paragraph
directly. See
the replacement,
formatting, and
revision references for details.
Tracked replacement uses the same unique localization and the same start-run
rule: inserted revision text takes the direct formatting of the run holding the
first consumed character, while deletion markup retains each source run's own
formatting. Replacing bold/italic Alpha with Omega authors a bold Omega,
because the changed interval starts in the bold run. A changed interval that
would cross distinct inline wrapper owners still refuses and leaves the
document unchanged.
Live blocks and spans captured from historical revision views remain useful for inspection, but mutation destinations must be reacquired from the current view. Cross-document composition applies that rule to its destination and refuses before importing anything when insertion after a paragraph, table, or block content control would remain inside an open complex-field result.
Reviewing
doc.revisionsenumerates and resolves tracked changes across every part: insertions, deletions, run and paragraph format changes, table-row revisions, and moves. Lists unresolvable markup by name.docx.comments/docx.commentopsread comments and delete one by identity, keeping the modern comment identity parts consistent.docx.protectionreads, sets, and respects Restrict-Editing. Mutating operations refuse on a protected document unless the caller explicitly overrides; the protection setting stays in the document.
Working across documents
docx.package.comparegenerates a native tracked-change redline from two documents, with the accept/reject round-trip shown above.docx.package.patch_save/diff_package/text_diffkeeps unchanged parts byte-identical and reports changed parts and text.diagnoseexplains why an unreadable file cannot be opened.docx.compositioncopies formatted content between documents, reconciles styles, numbering, media, hyperlinks, and bookmarks, and reports every part touched.docx.errorsexposes typed refusals, distinct from programmer errors.
Safety contract
Callers can catch PaperRefusal separately from programmer errors, which remain plain ValueError or TypeError. The subclasses say what went wrong: DocumentProtectedError, MalformedPackageError, TargetNotFoundError, AmbiguousTargetError, UnsupportedStructureError, RelationshipPolicyError, and BoundaryViolationError. Comparison and rewrite paths preserve meaningful whitespace, including trailing spaces inside runs.
Drop-in by design
Only the distribution and repository are renamed. The importable package stays docx. This is the same distribution/import split as Pillow (pip install pillow, import PIL), and it preserves existing code, snippets, and model priors.
- GitHub repository / PyPI distribution:
paper-docx - Python import:
docx - Fork sentinel:
docx.__paper_version__ = "0.2.0"
Installation
Install from PyPI:
python -m pip uninstall -y python-docx paper-docx
python -m pip install paper-docx
The clean uninstall is required when migrating from python-docx. Both distributions use the frozen docx import package, and pip cannot safely overlay or uninstall two distributions that own the same files.
Confirm the install:
paper-docx-doctor
Pip does not treat paper-docx as satisfying another package's declared dependency on python-docx. That dependency will reinstall upstream and overwrite shared docx files. Replace or remove the dependency, or run that package in a separate environment.
In a controlled deployment, a constraint containing python-docx<0 makes pip reject direct or transitive attempts to install upstream. The constraint must be applied to every install in that environment.
Documentation
The Sphinx docs extend the upstream python-docx documentation to cover the fork's additions: start with docs/user/paper-additions.rst and the docs/api/paper-*.rst reference pages. Inherited python-docx behavior works as documented at the python-docx documentation.
Testing
- Upstream's pytest and behave suites run on every commit to check compatibility with existing behavior.
- A frozen, hash-pinned fixture corpus spans generated and LibreOffice-authored documents.
- The contract harness checks refusal atomicity and validates the fixture corpus with a headless LibreOffice load smoke.
Contributing
Contributions are welcome — see CONTRIBUTING.md for the engineering discipline this fork runs on. The short version: the upstream suite must remain green; persistence changes need saved-and-reopened assertions and exact package-delta checks; guarded refusals must be atomic and leave bytes unchanged; and a refusal message must name what was found, why it is unsafe, and what to do about it.
Useful non-code contributions include real-world fixtures authored by desktop Word under the provenance rules in CONTRIBUTING.md.
Community
- Bugs and feature requests: GitHub Issues
- Questions and ideas: GitHub Discussions
- Security: see SECURITY.md
Acknowledgments
paper-docx exists because python-docx's package layer and XML mapping are excellent. Thanks to Steve Canny and the python-docx contributors for the work this project builds on. Upstream python-docx lives at github.com/python-openxml/python-docx.
Citation
If you reference paper-docx in research or writing:
@software{paper_docx,
title = {paper-docx: an agent-first structure editor for Word documents},
author = {{Paper Instruments, Inc.}},
year = {2026},
url = {https://github.com/paper-instruments/paper-docx}
}
Cite it as a fork of python-docx by Steve Canny and contributors.
License
MIT, inherited from python-docx. Original work © Steve Canny and the python-docx contributors; fork additions © Paper Instruments, Inc. This fork preserves the upstream license and attribution. See LICENSE.
Release files for paper-docx 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| paper_docx-0.2.0.tar.gz | 7.3 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| paper_docx-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 7.7 MB
Release files / paper_docx-0.2.0.tar.gz
| Download URL | paper_docx-0.2.0.tar.gz |
|---|---|
| Size | 7.3 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
034a961dc42eb44d7676e3a5af7ded485e89eb97851f8ccb24c481355b2c9d7d
|
|
BLAKE2b-256 checksum How to use checksums |
70dcbdc50405502a1a92707284005c1fab3e9e61fd3a7894a18bd80e24ac1920
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency logRelease files / paper_docx-0.2.0-py3-none-any.whl
| Download URL | paper_docx-0.2.0-py3-none-any.whl |
|---|---|
| Size | 439.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d61b6c1e23c2e50364ca9ebf913b08121bba12dc9119ffd0047ad927ff29ace0
|
|
BLAKE2b-256 checksum How to use checksums |
f8c6b54eb65833f5c59e10be4d8f3f422eefba440aa805cca345f80410383975
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency log