Skip to main content

A Python library for converting .docx files to DAISY 3 digital books with document preparation checks

Project description

DOCX2DAISY

DOCX2DAISY is a Python library that converts Microsoft Word (.docx) documents into DAISY 3 compliant digital books. The output is a ZIP package containing all necessary files for a DAISY 3 book, including text and images. In addition, DOCX2DAISY provides document preparation checks to help ensure that your Word document is properly structured for conversion.

Features

  • Comprehensive .docx Conversion
    Converts all text content (headings, paragraphs, lists, etc.) from .docx files into a structured DTBook XML format.

  • Image Extraction and Embedding
    Extracts images from the .docx (if present) and includes them in the DAISY package. Images are referenced in the DTBook XML via mediaobject elements with appropriate alternative text.

  • Document Preparation Checks
    Provides functions to verify that your document is properly prepared for conversion, checking for:

    • Essential metadata (Title, Author, Language)
    • Proper use of heading styles (e.g., Heading 1, Heading 2)
    • Alternative text for images
    • Presence of tables (and a note to provide descriptions, if needed)
  • Standards Compliance
    Generates the necessary DAISY 3 files:

    • DTBook XML (book.xml): Contains formatted text and images.
    • NCX (book.ncx): Provides navigation (table of contents) with proper DOCTYPE declaration.
    • SMIL (book.smil): Includes synchronization information with an empty <audio> tag for each navigation point and a DOCTYPE declaration.
    • OPF (book.opf): The package manifest linking all resources.
  • ZIP Packaging
    All generated files, including extracted images (if any), are bundled into a single ZIP package for easy distribution and use with DAISY readers.

Requirements

Install the dependency via pip:

pip install python-docx

Installation

To install the library locally, clone the repository and use pip:

git clone https://github.com/yourusername/DOCX2DAISY.git
cd DOCX2DAISY
pip install .

Usage

Document Preparation Checks

Before conversion, you can run a check on your .docx file to ensure proper preparation:

from daisy_converter import check_docx_preparation

check_docx_preparation("sample.docx")

This will output warnings if essential metadata, headings, alternative text for images, or table descriptions are missing.

Conversion Example

Below is an example script (tests/test_convert.py) demonstrating how to use DOCX2DAISY:

from daisy_converter import docx_to_daisy, check_docx_preparation
import zipfile
import os

input_docx = "sample.docx"       # Path to your .docx file (ensure it exists)
output_zip = "sample_daisy.zip"  # Desired output ZIP package name

# 1. Run document preparation checks
print("Running document preparation checks:")
check_docx_preparation(input_docx)
print("-" * 40)

# 2. Convert the DOCX to a DAISY 3 package
docx_to_daisy(input_docx, output_zip)
print(f"Converted '{input_docx}' to DAISY 3 book: '{output_zip}'.")

# 3. Verify the contents of the generated ZIP file
if os.path.exists(output_zip):
    with zipfile.ZipFile(output_zip, 'r') as z:
        files = z.namelist()
        print("\nFiles in the ZIP package:")
        for f in files:
            print(" -", f)
        
        required_files = {"book.xml", "book.ncx", "book.smil", "book.opf"}
        missing = required_files - set(files)
        if missing:
            print("\nMissing required files:", missing)
        else:
            print("\nAll required files are present.")

        image_files = [f for f in files if f.startswith("images/") and not f.endswith("/")]
        if image_files:
            print("\nImage files found:", image_files)
        else:
            print("\nNo images found in the package.")
else:
    print("Conversion failed: Output ZIP file not found.")

Project Structure

DOCX2DAISY/
├── docx2daisy/
│   ├── __init__.py
│   └── daisy_converter.py
├── tests/
│   └── test_convert.py
├── README.md
├── LICENSE
└── setup.py

Contributing

Contributions, bug reports, and feature requests are welcome. Please open an issue or submit a pull request on GitHub.

License

This project is licensed under the MIT License. See the LICENSE file for details.

Acknowledgments

This library leverages the python-docx library for processing Word documents and adheres to DAISY 3 specifications for digital talking books.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

docx2daisy-0.2.0.tar.gz (8.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

docx2daisy-0.2.0-py3-none-any.whl (9.0 kB view details)

Uploaded Python 3

File details

Details for the file docx2daisy-0.2.0.tar.gz.

File metadata

  • Download URL: docx2daisy-0.2.0.tar.gz
  • Upload date:
  • Size: 8.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for docx2daisy-0.2.0.tar.gz
Algorithm Hash digest
SHA256 0af64730ba3d60af78e77a63d8e1463eac456d0fce25a870463c9b1274816fa7
MD5 b0e1471daa78263b93af3834283fa364
BLAKE2b-256 8aa855e057f35dc7fb0256065829c1fa03800bb1bd94ce9fba6555c0bb2d9451

See more details on using hashes here.

File details

Details for the file docx2daisy-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: docx2daisy-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 9.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for docx2daisy-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 54cfe690295c3a4eeda1e62509796a995947fbcaa38cdaea267f68394e91f526
MD5 bd400a1f731a84e7c5042b757e9602d6
BLAKE2b-256 dedf506278f3cfbb2e183cda7176292f7c682d47f8865ae7a8be9a92a68fcf53

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page