This project is forked from ankushshah89/python-docx2txt. A new feature is added: extract the hyperlinks and its corresponding texts.
It is a pure python-based utility to extract text from docx files. The code is taken and adapted from python-docx. It can however also extract text from header, footer and hyperlinks. It can now also extract images.
How to install?
pip install docxpy
How to run?
From command line:
# extract text
docx2txt file.docx
# extract text and images
docx2txt -i /tmp/img_dir file.docx
From python:
import docxpy
file = 'file.docx'
# extract text
text = docxpy.process(file)
# extract text and write images in /tmp/img_dir
text = docxpy.process(file, "/tmp/img_dir")
# if you want the hyperlinks
doc = docxpy.DOCReader(file)
doc.process() # process file
hyperlinks = doc.data['links']
Metadata
Release files for docxpy 0.8.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| docxpy-0.8.5.tar.gz | 4.1 kB | Details |
Release files / docxpy-0.8.5.tar.gz
| Download URL | docxpy-0.8.5.tar.gz |
|---|---|
| Size | 4.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7949c5b8f6a1b749d1449f4590a3ddc6a3c16d62944b548df2efba52bad3d857
|
|
BLAKE2b-256 checksum How to use checksums |
de39d3c28e3ef0637237356306d3e7916cf9d4deddc2c7517b16765b4bdb7b13
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |