Xyslice
Xyslice extracts text, tables, and images from PDF and Office documents. It uses recursive XY cut analysis to split document regions along whitespace gaps and assemble text in reading order.
The PyPI distribution is named xyslice. The Python package is named peppermint.
Installation
Python 3.12 or newer is required.
pip install xyslice
Extract text from a PDF
from peppermint.extraction.pdfxycut import extract_blocks
blocks = list(extract_blocks("example.pdf"))
for block in blocks:
if block["type"] == "text":
print(block["page"], block["text"])
PDF blocks include their type, bounding box, and page number. Text blocks contain extracted text, table blocks contain cell rows, and image blocks contain Base64 image data. PDF page numbers start at 1.
Supported formats
| Format | Extraction function | Return value |
|---|---|---|
peppermint.extraction.pdfxycut.extract_blocks |
Iterator of blocks | |
| Word DOCX | peppermint.extraction.docxycut.extract_blocks |
List of page groups containing blocks |
| PowerPoint PPTX | peppermint.extraction.pptxycut.extract_blocks |
Iterator of blocks with slide numbers in page |
| Excel XLSX and CSV | peppermint.extraction.xlsxcut.extract_blocks |
Iterator of table and image blocks with sheet information |
Each function accepts a file path. For example, extract spreadsheet tables with:
from peppermint.extraction.xlsxcut import extract_blocks
for block in extract_blocks("example.xlsx"):
if block["type"] == "table":
for row in block["rows"]:
print(row)
Layout features and training
The optional training dependencies provide pandas, a Parquet engine, scikit-learn, and model serialization:
pip install "xyslice[training]"
python -m peppermint.features.build example.pdf --out features.parquet
Feature extraction produces geometry, font, spacing, alignment, and border features. The training function in peppermint.models.train_layout_classifier expects a Parquet dataset containing a label column. No pretrained classifier is included.
Limitations
Xyslice is an early release. Layout extraction uses heuristics, so results depend on document structure and formatting. Scanned PDF text requires OCR before extraction. DOCX and PPTX layout positions are estimated from document properties rather than rendered by Microsoft Office. Legacy binary Office formats such as DOC, PPT, and XLS are not supported.
License
Xyslice is licensed under AGPL-3.0-or-later. The release includes the full license text. Its PDF extraction dependency, PyMuPDF, is offered under AGPL or commercial licensing; see the PyMuPDF licensing documentation.
Metadata
Release files for xyslice 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| xyslice-0.1.3.tar.gz | 34.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| xyslice-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 74.0 kB
Release files / xyslice-0.1.3.tar.gz
| Download URL | xyslice-0.1.3.tar.gz |
|---|---|
| Size | 34.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c23f4e51ccb7e5e59d2601748e05aa4a550113e6539ec10447a11bdf597b1f25
|
|
BLAKE2b-256 checksum How to use checksums |
daeb41ea95a04092dc788e2408c8fc116946eb936573540489988f219e6d19e0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.8.5
|
Release files / xyslice-0.1.3-py3-none-any.whl
| Download URL | xyslice-0.1.3-py3-none-any.whl |
|---|---|
| Size | 39.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7068e1b97c7f74edc8b58b2f68785d7704207caf10737878c92efa3645409a81
|
|
BLAKE2b-256 checksum How to use checksums |
a0374fc6fed9a080fc8335cd11a3ecee03c419cb552a2ef2b0716aaac1cff789
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.8.5
|