Skip to main content

High-performance HWP/HWPX document extraction library

Project description

unhwp

High-performance Python library for extracting HWP/HWPX Korean word processor documents to Markdown.

Installation

pip install unhwp

Quick Start

import unhwp

# Simple conversion
markdown = unhwp.to_markdown("document.hwp")
print(markdown)

# Extract plain text
text = unhwp.extract_text("document.hwp")

# Full parsing with images
with unhwp.parse("document.hwp") as result:
    print(result.markdown)
    print(f"Sections: {result.section_count}")
    print(f"Paragraphs: {result.paragraph_count}")

    # Save images
    for img in result.images:
        img.save(f"output/{img.name}")

Features

  • Fast: Native Rust library with zero-copy parsing
  • Complete: Extracts text, tables, images, and document structure
  • Clean Output: Optional cleanup pipeline for polished Markdown
  • Format Support: HWP 5.0, HWPX, and HWP 3.x (legacy)

API Reference

Functions

to_markdown(path) -> str

Convert an HWP/HWPX document to Markdown.

markdown = unhwp.to_markdown("document.hwp")

to_markdown_with_cleanup(path, cleanup_options=None) -> str

Convert with optional cleanup.

markdown = unhwp.to_markdown_with_cleanup(
    "document.hwp",
    cleanup_options=unhwp.CleanupOptions.aggressive()
)

extract_text(path) -> str

Extract plain text content.

text = unhwp.extract_text("document.hwp")

parse(path, render_options=None) -> ParseResult

Parse a document with full access to content and images.

with unhwp.parse("document.hwp") as result:
    print(result.markdown)
    print(result.text)
    for img in result.images:
        print(img.name, len(img.data))

detect_format(path) -> int

Detect the document format.

fmt = unhwp.detect_format("document.hwp")
if fmt == unhwp.FORMAT_HWP5:
    print("HWP 5.0 format")
elif fmt == unhwp.FORMAT_HWPX:
    print("HWPX format")

Classes

RenderOptions

Options for Markdown rendering.

opts = unhwp.RenderOptions(
    include_frontmatter=True,
    image_path_prefix="images/",
    preserve_line_breaks=False,
)

CleanupOptions

Options for output cleanup.

# Presets
opts = unhwp.CleanupOptions.minimal()
opts = unhwp.CleanupOptions.default()
opts = unhwp.CleanupOptions.aggressive()
opts = unhwp.CleanupOptions.disabled()

# Custom
opts = unhwp.CleanupOptions(
    enabled=True,
    preset=1,
    detect_mojibake=True,
)

Constants

  • FORMAT_UNKNOWN - Unknown format
  • FORMAT_HWP5 - HWP 5.0 binary format
  • FORMAT_HWPX - HWPX XML format
  • FORMAT_HWP3 - HWP 3.x legacy format

Platform Support

  • Windows (x64)
  • Linux (x64)
  • macOS (x64, ARM64)

License

MIT License - see LICENSE for details.

Links

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

unhwp-0.1.12-py3-none-any.whl (2.9 MB view details)

Uploaded Python 3

File details

Details for the file unhwp-0.1.12-py3-none-any.whl.

File metadata

  • Download URL: unhwp-0.1.12-py3-none-any.whl
  • Upload date:
  • Size: 2.9 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for unhwp-0.1.12-py3-none-any.whl
Algorithm Hash digest
SHA256 41eb1e0500bc57ad37d5e7893baad4f4a89711bbfc9e556ff3a4a8356beea8a5
MD5 e44ac42113820c9cf23a216ba23973c3
BLAKE2b-256 bfe6ec18bdb121a5f17ce46a204af69c37be1e2a9ff96227cc84a893f7a8e6f6

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page