Skip to main content

Xyslice

Xyslice extracts text, tables, and images from PDF and Office documents. It uses recursive XY cut analysis to split document regions along whitespace gaps and assemble text in reading order.

The PyPI distribution is named xyslice. The Python package is named peppermint.

Installation

Python 3.12 or newer is required.

pip install xyslice

Extract text from a PDF

from peppermint.extraction.pdfxycut import extract_blocks

blocks = list(extract_blocks("example.pdf"))
for block in blocks:
    if block["type"] == "text":
        print(block["page"], block["text"])

PDF blocks include their type, bounding box, and page number. Text blocks contain extracted text, table blocks contain cell rows, and image blocks contain Base64 image data. PDF page numbers start at 1.

Supported formats

Format Extraction function Return value
PDF peppermint.extraction.pdfxycut.extract_blocks Iterator of blocks
Word DOCX peppermint.extraction.docxycut.extract_blocks List of page groups containing blocks
PowerPoint PPTX peppermint.extraction.pptxycut.extract_blocks Iterator of blocks with slide numbers in page
Excel XLSX and CSV peppermint.extraction.xlsxcut.extract_blocks Iterator of table and image blocks with sheet information

Each function accepts a file path. For example, extract spreadsheet tables with:

from peppermint.extraction.xlsxcut import extract_blocks

for block in extract_blocks("example.xlsx"):
    if block["type"] == "table":
        for row in block["rows"]:
            print(row)

Layout features and training

The optional training dependencies provide pandas, a Parquet engine, scikit-learn, and model serialization:

pip install "xyslice[training]"
python -m peppermint.features.build example.pdf --out features.parquet

Feature extraction produces geometry, font, spacing, alignment, and border features. The training function in peppermint.models.train_layout_classifier expects a Parquet dataset containing a label column. No pretrained classifier is included.

Limitations

Xyslice is an early release. Layout extraction uses heuristics, so results depend on document structure and formatting. Scanned PDF text requires OCR before extraction. DOCX and PPTX layout positions are estimated from document properties rather than rendered by Microsoft Office. Legacy binary Office formats such as DOC, PPT, and XLS are not supported.

License

Xyslice is licensed under AGPL-3.0-or-later. The release includes the full license text. Its PDF extraction dependency, PyMuPDF, is offered under AGPL or commercial licensing; see the PyMuPDF licensing documentation.

Metadata

Release files for xyslice 0.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for xyslice 0.1.3
File Size Uploaded
xyslice-0.1.3.tar.gz 34.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for xyslice 0.1.3
File Interpreter ABI Platform
xyslice-0.1.3-py3-none-any.whl Python 3 none any Details

Total release size: 74.0 kB

Release files / xyslice-0.1.3.tar.gz

Download URL xyslice-0.1.3.tar.gz
Size 34.6 kB
Tags Source
SHA-256 checksum
How to use checksums
c23f4e51ccb7e5e59d2601748e05aa4a550113e6539ec10447a11bdf597b1f25
BLAKE2b-256 checksum
How to use checksums
daeb41ea95a04092dc788e2408c8fc116946eb936573540489988f219e6d19e0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.5

Release files / xyslice-0.1.3-py3-none-any.whl

Download URL xyslice-0.1.3-py3-none-any.whl
Size 39.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7068e1b97c7f74edc8b58b2f68785d7704207caf10737878c92efa3645409a81
BLAKE2b-256 checksum
How to use checksums
a0374fc6fed9a080fc8335cd11a3ecee03c419cb552a2ef2b0716aaac1cff789
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.5

Release history Release notifications | RSS feed

This release

0.1.3 This release

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page