Skip to main content

Get the parsed microsoft word document in a hierarchical tree structure.

Project description

mswordtree

Parse your whole word document in a hierarchical tree structure. The document content will be listed down as Heading and its children as subheading/paragraph/table etc.

Install the library using following comand

pip install mswordtree

Use the following code to parse your word document in a tree structure

from mswordtree import GetWordDocTree
root = GetWordDocTree('test.docx')

Now you can iterate over all objects of the document by using the following code

for item in root.Items:
    print('Type: {} -> Content {}\n'.format(item.Type, item.Content))

To make the json use the following code

from mswordtree import ToString
ToString([root])

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Files for mswordtree, version 0.1.1.3
Filename, size File type Python version Upload date Hashes
Filename, size mswordtree-0.1.1.3-py3-none-any.whl (6.0 kB) File type Wheel Python version py3 Upload date Hashes View hashes
Filename, size mswordtree-0.1.1.3.tar.gz (8.5 kB) File type Source Python version None Upload date Hashes View hashes

Supported by

Elastic Elastic Search Pingdom Pingdom Monitoring Google Google BigQuery Sentry Sentry Error logging AWS AWS Cloud computing DataDog DataDog Monitoring Fastly Fastly CDN SignalFx SignalFx Supporter DigiCert DigiCert EV certificate StatusPage StatusPage Status page