Get the parsed microsoft word document in a hierarchical tree structure.
Project description
mswordtree
Parse your whole word document in a hierarchical tree structure. The document content will be listed down as Heading and its children as subheading/paragraph/table etc.
Install the library using following comand
pip install mswordtree
Use the following code to parse your word document in a tree structure
from mswordtree import GetWordDocTree
root = GetWordDocTree('test.docx')
Now you can iterate over all objects of the document by using the following code
for item in root.Items:
print('Type: {} -> Content {}\n'.format(item.Type, item.Content))
To make the json use the following code
from mswordtree import ToString
ToString([root])
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
mswordtree-0.1.1.2.tar.gz
(68.0 kB
view hashes)
Built Distribution
Close
Hashes for mswordtree-0.1.1.2-py3-none-any.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 9ec92f4a5a25bc2b8c3a2b295ec74afb1ddd596ea01b212398cdaa42dd33ad11 |
|
MD5 | d76b8a23949e632000b64c4997cf3083 |
|
BLAKE2b-256 | 71b32e5beadcf59a1eba873e6fc3faa38497554ee804470087b21ec7b02f07a2 |