Skip to main content

A Python library for extracting metadata from DOCX file footers using parallel processing

Project description

DOCX Footer Extractor

A Python library for extracting metadata from DOCX file footers using parallel processing.

Features

  • Extract key-value pairs from DOCX file footers
  • Process multiple files in parallel for better performance
  • Support for both folder processing and specific file lists
  • Extract metadata from footer text and tables
  • Python 3.9+ compatibility
  • Thread-safe processing with error handling

Installation

pip install docx_footer_extractor

Quick Start

from docx_footer_extractor import DocxFooterExtractor

# Create extractor instance
extractor = DocxFooterExtractor(max_workers=4)

# Process all DOCX files in a folder
results = extractor.extract("./documents")

# Process specific files
file_list = ["doc1.docx", "doc2.docx", "folder/doc3.docx"]
results = extractor.extract(file_list)

# Results format
for result in results:
    filename = result['filename']
    metadata = result['metadata']
    print(f"{filename}: {metadata}")

Usage

Using the Class

from docx_footer_extractor import DocxFooterExtractor

extractor = DocxFooterExtractor(max_workers=4)

# Process folder
results = extractor.extract("./my_documents")

# Process specific files
results = extractor.extract([
    "document1.docx",
    "path/to/document2.docx"
])

# Save results to file
extractor.save_results_to_file(results, "output.txt")

Output Format

The library returns a list of dictionaries with the following structure:

python[
    {
        "filename": "document1.docx",
        "metadata": {
            "Author": "John Doe",
            "Version": "1.0",
            "Date": "2025-01-15"
        }
    },
    {
        "filename": "document2.docx",
        "metadata": {
            "Title": "Report",
            "Department": "Sales"
        }
    }
]

Requirements

Python 3.9+
python-docx>=0.8.11

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

docx_footer_extractor-1.0.0.tar.gz (7.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

docx_footer_extractor-1.0.0-py3-none-any.whl (7.7 kB view details)

Uploaded Python 3

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page