EPUB to Text Converter
A professional, high-performance EPUB conversion library for extracting and converting EPUB files to multiple formats (Text, Markdown, JSON). Supports both single-file and batch processing with parallel execution.
Features:
- 📚 Extract chapters, images, and metadata from EPUB files
- 📝 Export to multiple formats: Text, Markdown, JSON, HTML
- 🔄 Batch process multiple EPUB files in parallel
- 🖼️ Extract and link images with proper paths
- 📊 Get detailed book information and statistics
- 🎯 Simple CLI and comprehensive Python API
Installation
pip install epub-to-text
Or from source:
pip install https://github.com/thinh-vu/epub_to_text.git
Quick Start
Command Line
# Convert to markdown chapters with images
epub-to-text your_book.epub --chapters-markdown --extract-images
# Convert to all formats
epub-to-text book.epub --all -o output/
# Show book information
epub-to-text your_book.epub --info
# Batch process EPUBs with parallel execution
epub-to-text /your_epub_folder_path --batch --all --parallel
Python API
from epub_to_text import EpubProcessor
# Basic usage
processor = EpubProcessor('book.epub', 'output/')
summary = processor.get_summary()
processor.export_chapters_markdown()
processor.extract_images()
from epub_to_text import BatchProcessor
# Batch processing
batch = BatchProcessor(max_workers=4)
result = batch.process_batch(
'/epub/folder',
'./output',
{'chapters_markdown': True, 'extract_images': True},
recursive=True,
parallel=True
)
Documentation
Complete documentation available in the docs/ folder:
- Quick Start - Get started in minutes
- API Reference - Complete class/method documentation
- Architecture Guide - System design and patterns
- Integration Guide - AI agent integration patterns
- Advanced Usage - Custom processors and optimization
CLI Options
Usage: epub-to-text [OPTIONS] <file_or_directory>
Options:
--single-text Export entire book as text
--single-markdown Export entire book as markdown
--chapters-text Export each chapter as text files
--chapters-markdown Export each chapter as markdown files
--json Export as JSON with metadata
--all Export in all formats
--extract-images Extract and save images
--batch Process multiple EPUBs
--recursive Search subdirectories
--parallel Process files in parallel
--max-workers N Number of parallel workers (default: 4)
--info Show book information only
--verbose Detailed output
-o, --output DIR Output directory (default: ./exported_books)
Output Structure
Single file mode:
output/
├── book.md # Complete book as markdown
├── book.txt # Complete book as text
└── book.json # Structured data
Chapter-wise mode:
output/
└── Book_Title/
├── 01_Introduction.md
├── 02_Chapter_Two.md
├── 03_Conclusion.md
└── images/
├── cover.jpg
└── diagram1.png
Project Structure
epub_to_text/
├── __init__.py # Package initialization
├── cli.py # Command-line interface
├── reader.py # EPUB file reading
├── extractor.py # Content extraction
├── converter.py # Format conversion
├── processor.py # Single-file processing
└── batch_processor.py # Batch processing
Key Classes
| Class | Purpose |
|---|---|
EpubProcessor |
High-level single-file processing |
BatchProcessor |
Batch processing with parallel support |
EpubExtractor |
Extract chapters, images, metadata |
ContentConverter |
Format conversion utilities |
EpubReader |
Low-level EPUB file reading |
Requirements
- Python 3.10+
- ebooklib >= 0.17.1
- beautifulsoup4 >= 4.9.0
Use Cases
- Knowledge Base: Extract EPUB content for building AI training datasets
- Content Analysis: Process multiple books for NLP tasks
- Digital Library: Convert EPUB collections to searchable text/markdown
- Accessibility: Generate alternative formats from EPUB books
- Content Preservation: Archive book content in multiple formats
Contributing
- Fork the repository
- Create a feature branch
- Make your changes
- Submit a pull request
License
This project is licensed under the MIT License - see the LICENSE file for details.
Support
For issues and questions:
- Check the documentation
- Review API Reference and Architecture Guide
- Search existing issues
Acknowledgments
- ebooklib - EPUB parsing
- BeautifulSoup - HTML/XML parsing
Metadata
Release files for epub-to-text 2.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| epub_to_text-2.0.0.tar.gz | 15.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| epub_to_text-2.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 31.8 kB
Release files / epub_to_text-2.0.0.tar.gz
| Download URL | epub_to_text-2.0.0.tar.gz |
|---|---|
| Size | 15.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e72bd00699fee05ef1509d868fbf2ab53c9109ba64134c22c27ccd1c0465b37c
|
|
BLAKE2b-256 checksum How to use checksums |
d63cfc00b1e4b79b2ce4127453bb3ef06e5e49b4da77ca6a87385a01ebcf8247
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.9
|
Release files / epub_to_text-2.0.0-py3-none-any.whl
| Download URL | epub_to_text-2.0.0-py3-none-any.whl |
|---|---|
| Size | 16.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8622d120c3b2a1ee37e8672edd87306ca0c0450382f2bda5733455fa95cae17d
|
|
BLAKE2b-256 checksum How to use checksums |
e9bafcf41939b3a29d201882cfcb0165c858af4cd3e6ce536a8328bd5e27bd16
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.9
|