Python package for converting xml and epubs to text files
Project description
epub conversion
Create text corpuses using epubs and wiki dumps. This is a python package with a Converter for epub and xml (wiki dumps) to text, lines, or Python generators.
Usage:
Epub usage
Book by book
To convert epubs to text files, usage is straightforward. First create a converter object:
converter = Converter("my_ebooks_folder/")
Then using this converter let's concatenate all the text within the ebooks into a single mega text file:
converter.convert("my_succinct_text_file.gz")
Line by line
You can also proceed line by line:
from epub_conversion.utils import open_book, convert_epub_to_lines
book = open_book("twilight.epub")
lines = convert_epub_to_lines(book)
Wikidump usage
Redirections
Suppose you are interested in all redirections in a given Wikipedia dump file that is still compressed, then you can access the dump as follows:
wiki = epub_conversion.wiki_decoder.almost_smart_open("enwiki.bz2")
Taking this dump as our input let us now use a generator to output all pairs of title
and redirection title
in this dump:
redirections = {redirect_from:redirect_to
for redirect_from, redirect_to in epub_conversion.wiki_decoder.get_redirection_list(wiki)
}
Page text
Suppose you are interested in the lines within each page's text section only, then:
for line in epub_conversion.wiki_decoder.convert_wiki_to_lines(wiki):
process_line( line )
See Also:
- Wikipedia NER a Python module that uses
epub_conversion
to process Wikipedia dumps and output only the lines that contain page to page links, with the link anchor texts extracted, and all markup removed.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
File details
Details for the file epub-conversion-1.0.15.tar.gz
.
File metadata
- Download URL: epub-conversion-1.0.15.tar.gz
- Upload date:
- Size: 6.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/3.1.1 pkginfo/1.5.0.1 requests/2.22.0 setuptools/42.0.2 requests-toolbelt/0.9.1 tqdm/4.42.0 CPython/3.7.6
File hashes
Algorithm | Hash digest | |
---|---|---|
SHA256 | c61a5460050952bb6cc294e6971c7cfbd711099c8828a8d6e45f4557f0bc8d6f |
|
MD5 | 9cd483f66b0e0c7a4dbf551a53240b80 |
|
BLAKE2b-256 | 25f2d95f96a5476532cf06c477647c2efb678724abc5698ff658f2a315010714 |