DOCX to Markdown for Python
Convert Word documents to Markdown with individual pages, headers, footers, comments, and tracked changes. Use the full document for text processing or attach page references to search results and AI answers.
The package wraps the @docx-editor.dev/docx-to-markdown converter as a self-contained executable. No Node.js installation is required.
PyPI · Documentation · Changelog · Try the demo
Install
pip install docx-to-markdown
Requires Python 3.11 or later. Wheels include the converter and bundled fonts.
| Platform | Supported systems |
|---|---|
| Linux | x64 and arm64, glibc 2.17 or newer |
| macOS | x64 and arm64, macOS 13 or later |
| Windows | x64 |
x64 CPUs must support SSE4.2. The package version matches the @docx-editor.dev/docx-to-markdown npm release it wraps.
Convert a file
from docx_to_markdown import convert
result = convert("contract.docx")
print(result.markdown)
for page in result.pages:
print(f"Page {page.number}", page.markdown)
print(page.header_markdown, page.footer_markdown)
result.markdown contains the continuous document body. Headers and footers stay in result.pages. convert also accepts bytes, pathlib.Path objects, and binary file objects:
with open("contract.docx", "rb") as source:
result = convert(source)
Page breaks depend on fonts, document features, and revision mode, so they can differ from Microsoft Word. Store the document version with page citations and inspect result.warnings for omitted or approximated content.
Convert many files
Each convert call starts a converter process. Use Converter to reuse one process across a batch of files:
from pathlib import Path
from docx_to_markdown import Converter
with Converter() as converter:
for path in Path("contracts").glob("*.docx"):
result = converter.convert(path)
result.write(Path("output") / path.stem)
Options given to Converter are defaults for every call. Pass the same keywords to converter.convert to override them per file. One Converter is safe to share between threads; use one per thread for parallelism.
Command line
docx-to-markdown contract.docx # Markdown to stdout
docx-to-markdown contract.docx -o contract.md
docx-to-markdown contract.docx --json # full result as JSON
docx-to-markdown *.docx --bundle out/ --images # document.md, document.json, media/ per file
docx-to-markdown contract.docx --font fonts/ # family, weight, style read from each file
docx-to-markdown contract.docx --font carlito/Carlito-Regular.ttf:Calibri
python -m docx_to_markdown runs the same tool. Warnings go to stderr; -q hides them.
Fonts
Page breaks depend on font metrics. The package ships metric-compatible substitutes for Word's default fonts: Calibri, Cambria, Times New Roman, Arial, Courier New, and Century Gothic. Documents that use other fonts paginate approximately unless you supply the font files.
Point fonts at a directory or file. Family, weight, and style are read from each file, so a folder of licensed fonts is one argument:
from docx_to_markdown import convert
result = convert("contract.docx", fonts="fonts/")
Pass a list to combine several families or sources. Entries can be folders, files, or faces, and the first entry that serves a face wins:
from docx_to_markdown import convert, font_family
result = convert(
"contract.docx",
fonts=["brand-fonts/", "extra/Roboto-Bold.ttf", *font_family("Aptos", "Aptos.ttf")],
)
When the file's own family name differs from the name the document uses, register it under the document's name:
from docx_to_markdown import convert, font_files
result = convert("contract.docx", fonts=font_files("carlito/", family="Calibri"))
Or name each face yourself:
from docx_to_markdown import convert, font_family
aptos = font_family(
"Aptos",
"fonts/Aptos.ttf",
bold="fonts/Aptos-Bold.ttf",
italic="fonts/Aptos-Italic.ttf",
bold_italic="fonts/Aptos-BoldItalic.ttf",
)
result = convert("contract.docx", fonts=aptos)
assert result.fonts_complete
Your fonts take precedence over the bundled substitutes. A font file the converter cannot read is reported in result.font_errors and the conversion continues without it. Pass font_policy="strict" to fail instead of approximating.
Google Fonts fallback
google_fonts=True fetches families the local fonts cannot serve from a pinned Google Fonts catalog. It needs network access. The catalog is a closed set of families that ship four static faces, so it does not include variable-only families such as Roboto, Open Sans, or Lato. google_font_families() returns the list, and result.missing_fonts tells you what a document still needs:
from docx_to_markdown import convert, google_font_families
result = convert("report.docx", google_fonts=True)
still_missing = [f for f in result.missing_fonts if f not in google_font_families()]
Supply anything still missing as font files.
result.font_resolution reports which face measured each family. result.fonts_complete is True when every family measured with all of its faces. Check it before you rely on page numbers.
Images
Enable image extraction and save a folder with Markdown, JSON, and image files:
from docx_to_markdown import convert
result = convert("report.docx", images=True)
result.write("out/") # document.md, document.json, media/
result.media holds each unique image with its bytes, pixel size, and page occurrences. Pass images="html" to keep displayed sizes in <img> tags.
Comments and tracked changes
Use display_mode="proposed" when indexing the document's proposed text:
from docx_to_markdown import convert
result = convert("contract.docx", display_mode="proposed")
for comment in result.comments:
print(comment.author, comment.text)
display_mode selects how tracked changes are projected. "all-markup" (the default) keeps every insertion and deletion visible, "proposed" shows the document as if every change were accepted, and "original" as if every change were rejected. Comments and tracked changes with their Markdown offsets are in result.review_artifacts and result.review_bindings.
This setting changes the exported view; it does not accept or reject changes in the source DOCX file.
Typed end to end
Every argument and every field of the result is typed, and the package ships py.typed. Pages, warnings, images, font resolution, comments, tracked changes, and Markdown bindings are frozen dataclasses with snake_case fields, so your editor can complete and check calls like these:
from docx_to_markdown import PageProjection
for change in result.tracked_changes: # list[TrackedChange]
print(change.change, change.author, change.text)
for binding in result.review_bindings: # list[ReviewBinding]
if isinstance(binding.projection, PageProjection):
print(binding.projection.page_number, binding.ranges[0].start)
result.raw keeps the converter's complete JSON for anything the records do not carry.
Result
| Field | Content |
|---|---|
markdown |
The full logical document |
pages |
Page: number, markdown, header_markdown, footer_markdown, comments, tracked_changes |
warnings |
ExportWarning with a stable code, a message, and a page number when known |
media |
MediaAsset with bytes, pixel size, and ImageOccurrence placements |
font_resolution |
FontResolution: which face measured each family, with missing and complete |
font_errors |
Font files that could not be admitted |
review_artifacts |
Comment and TrackedChange records, also split as comments and tracked_changes |
review_bindings |
ReviewBinding: where each artifact sits in the Markdown, in UTF-16 offsets |
pagination |
Pagination: layout revision and display mode |
raw |
The converter's complete JSON result |
write(directory) saves document.md, document.json, and media/, the same layout as the Node.js package's writeMarkdownBundle.
Errors
ConversionError carries a stable code such as docx-unreadable, timeout, or runtime-missing, and a message. A missing input path raises FileNotFoundError.
Build from source
Versions and the shared changelog are managed by Changesets. For Python package changes, select @docx-editor.dev/docx-to-markdown when running bun changeset and describe the Python behavior in the release note. The release workflow builds and publishes Python wheels from the same version tag as the npm packages. See the release guide for the publishing flow.
The runtime is a Bun executable built from this repository. From the repository root:
bun install
bun run build:packages
cd python/docx-to-markdown
bun run typecheck:runtime
bun run build:runtime
uv build --wheel
bun run build:runtime compiles runtime/main.ts for the current machine and copies the packaged fonts and license texts into src/docx_to_markdown/_vendor/. Pass --target bun-linux-x64 to cross-compile. Set DOCX_TO_MARKDOWN_RUNTIME to run the package against an executable built elsewhere.
License
The package is licensed under the Apache License, Version 2.0. LICENSE and NOTICE ship in the wheel's metadata.
The executable is a compiled bundle. Everything it contains keeps its own license, and every text travels with the wheel in docx_to_markdown/_vendor/licenses/:
THIRD_PARTY_NOTICES.mdlists each bundled npm package with its license text. The list is generated from the bundle graph at build time, and a package without a license text fails the build.bun-LICENSE.mdcovers the Bun runtime (MIT) and the libraries it links, including JavaScriptCore under the LGPL.harfbuzz-COPYING.txtcovers the HarfBuzz text shaper.OFL-*.txt,LICENSE-Liberation.txt,GUST-FONT-LICENSE.txt, andLPPL-1.3c.txtcover the bundled font files.
Release files for docx-to-markdown 2.21.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distributions (wheels)
| File | Reset | |||
|---|---|---|---|---|
| docx_to_markdown-2.21.0-py3-none-win_amd64.whl | Python 3 | none | Windows x86-64 | Details |
| docx_to_markdown-2.21.0-py3-none-manylinux_2_17_x86_64.whl | Python 3 | none | Linux glibc 2.17+ x86-64 | Details |
| docx_to_markdown-2.21.0-py3-none-manylinux_2_17_aarch64.whl | Python 3 | none | Linux glibc 2.17+ ARM64 | Details |
| docx_to_markdown-2.21.0-py3-none-macosx_13_0_x86_64.whl | Python 3 | none | macOS 13.0+ x86-64 | Details |
| docx_to_markdown-2.21.0-py3-none-macosx_13_0_arm64.whl | Python 3 | none | macOS 13.0+ ARM64 | Details |
Total release size: 191.6 MB
Release files / docx_to_markdown-2.21.0-py3-none-win_amd64.whl
| Download URL | docx_to_markdown-2.21.0-py3-none-win_amd64.whl |
|---|---|
| Size | 44.7 MB |
| Tags | Python 3 Windows x86-64 |
|
SHA-256 checksum How to use checksums |
d4adb47effbdd4263dc27d01fbc6e4fa563915586644eb9437dc8d8a930bb597
|
|
BLAKE2b-256 checksum How to use checksums |
2ad72868e7e0a665127cfe2b5b3bd21a151ebcba343f73d49845f2b421133652
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.
Transparency logRelease files / docx_to_markdown-2.21.0-py3-none-manylinux_2_17_x86_64.whl
| Download URL | docx_to_markdown-2.21.0-py3-none-manylinux_2_17_x86_64.whl |
|---|---|
| Size | 41.5 MB |
| Tags | Linux glibc 2.17+ x86-64 Python 3 |
|
SHA-256 checksum How to use checksums |
4d0d7e905f49b9a1fa14d2f5d825655a261850b9fc25ad8f1c857dde1c223aaa
|
|
BLAKE2b-256 checksum How to use checksums |
fb4d4cf95b8307d6127305df2a5c119c701c61c9d600dbf3854425e663157e26
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.
Transparency logRelease files / docx_to_markdown-2.21.0-py3-none-manylinux_2_17_aarch64.whl
| Download URL | docx_to_markdown-2.21.0-py3-none-manylinux_2_17_aarch64.whl |
|---|---|
| Size | 41.5 MB |
| Tags | Linux glibc 2.17+ ARM64 Python 3 |
|
SHA-256 checksum How to use checksums |
2fd883d71024976d0f4b2afe4116d860020c0cd0733793938194ba50755206e9
|
|
BLAKE2b-256 checksum How to use checksums |
022f9b61f28e1d2c2704aaf2cec0671e2f4ef4e1efc7a99618e50f6e0aaf8cbd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.
Transparency logRelease files / docx_to_markdown-2.21.0-py3-none-macosx_13_0_x86_64.whl
| Download URL | docx_to_markdown-2.21.0-py3-none-macosx_13_0_x86_64.whl |
|---|---|
| Size | 33.3 MB |
| Tags | Python 3 macOS 13.0+ x86-64 |
|
SHA-256 checksum How to use checksums |
768fbfa8fdd1c564551c617c07cd0c7d4e361bdb4fccc2dceca05afff33e4d4c
|
|
BLAKE2b-256 checksum How to use checksums |
6c43d27ceafaa553b4508799922a50ac99b8d4ed40bdea1a60d6517c10a2757c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.
Transparency logRelease files / docx_to_markdown-2.21.0-py3-none-macosx_13_0_arm64.whl
| Download URL | docx_to_markdown-2.21.0-py3-none-macosx_13_0_arm64.whl |
|---|---|
| Size | 30.6 MB |
| Tags | Python 3 macOS 13.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
0fda9c7927e5550d269a2c87240a52bc9a2a4c26fabc7be060ea564f123f3db4
|
|
BLAKE2b-256 checksum How to use checksums |
cff7d11c0ba3bd4d93566f6e072d6d9ede841569b34e8185075e48ccaff6b12e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.
Transparency log