Skip to main content

Xyslice

Xyslice extracts text, tables, and images from PDF and Office documents. It uses recursive XY cut analysis to split document regions along whitespace gaps and assemble text in reading order.

The PyPI distribution is named xyslice. The Python package is named peppermint.

Installation

Python 3.12 or newer is required.

pip install xyslice

Extract text from a PDF

from peppermint.extraction.pdfxycut import extract_blocks

blocks = list(extract_blocks("example.pdf"))
for block in blocks:
    if block["type"] == "text":
        print(block["page"], block["text"])

PDF blocks include their type, bounding box, and page number. Text blocks contain extracted text, table blocks contain cell rows, and image blocks contain Base64 image data. PDF page numbers start at 1.

Supported formats

Format Extraction function Return value
PDF peppermint.extraction.pdfxycut.extract_blocks Iterator of blocks
Word DOCX peppermint.extraction.docxycut.extract_blocks List of page groups containing blocks
PowerPoint PPTX peppermint.extraction.pptxycut.extract_blocks Iterator of blocks with slide numbers in page
Excel XLSX and CSV peppermint.extraction.xlsxcut.extract_blocks Iterator of table and image blocks with sheet information

Each function accepts a file path. For example, extract spreadsheet tables with:

from peppermint.extraction.xlsxcut import extract_blocks

for block in extract_blocks("example.xlsx"):
    if block["type"] == "table":
        for row in block["rows"]:
            print(row)

Layout features and training

The optional training dependencies provide pandas, a Parquet engine, scikit-learn, and model serialization:

pip install "xyslice[training]"
python -m peppermint.features.build example.pdf --out features.parquet

Feature extraction produces geometry, font, spacing, alignment, and border features. The training function in peppermint.models.train_layout_classifier expects a Parquet dataset containing a label column. No pretrained classifier is included.

Limitations

Xyslice is an early release. Layout extraction uses heuristics, so results depend on document structure and formatting. Scanned PDF text requires OCR before extraction. DOCX and PPTX layout positions are estimated from document properties rather than rendered by Microsoft Office. Legacy binary Office formats such as DOC, PPT, and XLS are not supported.

License

Xyslice is licensed under AGPL-3.0-or-later. The release includes the full license text. Its PDF extraction dependency, PyMuPDF, is offered under AGPL or commercial licensing; see the PyMuPDF licensing documentation.

Metadata

Release files for xyslice 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for xyslice 0.1.0
File Size Uploaded
xyslice-0.1.0.tar.gz 34.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for xyslice 0.1.0
File Interpreter ABI Platform
xyslice-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 74.0 kB

Release files / xyslice-0.1.0.tar.gz

Download URL xyslice-0.1.0.tar.gz
Size 34.6 kB
Tags Source
SHA-256 checksum
How to use checksums
8fee2f07877fadde238cb826e3ad625e412ef2d4022f9fc99221ca809737f34a
BLAKE2b-256 checksum
How to use checksums
9d049857b66347e9f8067746693636735e86968d301acc10cf934d5ca389f36d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.5

Release files / xyslice-0.1.0-py3-none-any.whl

Download URL xyslice-0.1.0-py3-none-any.whl
Size 39.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
fc4ac4d9bcddc19abfa7c240c7834af5888c4989beb034f040e820f03cc346d6
BLAKE2b-256 checksum
How to use checksums
f5e20cce379261fd57147a9f219aa5ef6a6c9338c1496a058b415ecbcaaadef6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.5

Release history Release notifications | RSS feed

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page