Pyntagma is a Python library for creating and managing complex data structures with ease. Its name is derived from the Greek word 'Syntagma', meaning 'composition', symbolizing that this package fits for semi-structured documents
Project description
Pyntagma
Pyntagma is a Python library for creating and managing complex data structures with ease. Its name is derived from the Greek word 'Syntagma', meaning 'composition', symbolizing that this package fits for semi-structured documents.
Features
- PDF Document Processing: Extract and analyze text, words, and lines from PDF documents
- Multi-file Document Support: Handle documents that span multiple PDF files
- Precise Positioning: Track exact coordinates and positions of text elements
- Type-safe Design: Built with Pydantic models for robust data validation
- Silent PDF Processing: Suppresses verbose logging during PDF operations
- Flexible Cropping: Extract specific regions from PDF pages
Installation
Install Pyntagma using uv (recommended):
uv add git+https://github.com/MarcellGranat/pyntagma.git
Quick Start
Basic Document Processing
from pyntagma import Document
from pathlib import Path
# Create a document from one or more PDF files
doc = Document(files=[
Path("document-part1.pdf"),
Path("document-part2.pdf")
])
# Access pages
print(f"Total pages: {len(doc.pages)}")
# Get the first page
page = doc.pages[0]
print(f"Page dimensions: {page.width} x {page.height}")
# Extract words and lines
words = page.words
lines = page.lines
print(f"Found {len(words)} words and {len(lines)} lines")
Working with Text Elements
# Access word properties
for word in page.words[:5]: # First 5 words
print(f"'{word.text}' at position ({word.x0}, {word.top})")
print(f"Word dimensions: {word.x1 - word.x0} x {word.bottom - word.top}")
# Access line properties
for line in page.lines[:3]: # First 3 lines
print(f"Line: '{line.text}'")
print(f"Line words: {len(line.words)}")
Position-based Operations
from pyntagma import Position, HorizontalCoordinate, VerticalCoordinate
# Create custom positions
position = Position(
x0=HorizontalCoordinate(page=page, value=100),
x1=HorizontalCoordinate(page=page, value=200),
top=VerticalCoordinate(page=page, value=50),
bottom=VerticalCoordinate(page=page, value=80)
)
# Check if one position contains another
word_position = page.words[0].position
if position.contains(word_position):
print("Word is within the specified region")
PDF Cropping
from pyntagma import Crop
# Define a crop region
crop = Crop(
path=Path("document.pdf"),
page_number=0,
x0=100.0,
x1=400.0,
top=50.0,
bottom=200.0,
padding=10,
resolution=300
)
# Use the crop for further processing
print(f"Crop region: {crop}")
API Reference
Core Classes
Document
Represents a multi-file PDF document.
Properties:
files: list[Path]- List of PDF files comprising the documentpages: list[Page]- All pages across all filesn_pages: int- Total number of pages
Page
Represents a single page within a document.
Properties:
path: Path- Path to the PDF file containing this pagefile_page_number: int- Page number within the file (0-indexed)page_number: int- Page number within the document (0-indexed)words: list[Word]- All words on the pagelines: list[Line]- All lines on the pageheight: float- Page height in pointswidth: float- Page width in points
Word
Represents a single word with position information.
Properties:
page: Page- The page containing this wordtext: str- The word textx0, x1: float- Horizontal boundariestop, bottom: float- Vertical boundariesposition: Position- Position object for spatial operationsline: Line- The line containing this word
Line
Represents a line of text with position information.
Properties:
page: Page- The page containing this linetext: str- The complete line textx0, x1: float- Horizontal boundariestop, bottom: float- Vertical boundariesposition: Position- Position object for spatial operationswords: list[Word]- Words within this line
Position
Represents a rectangular region on a page.
Properties:
x0, x1: HorizontalCoordinate- Horizontal boundariestop, bottom: VerticalCoordinate- Vertical boundaries
Methods:
contains(other: Position) -> bool- Check if this position contains another
Crop
Defines a rectangular region for extraction from a PDF page.
Properties:
path: Path- PDF file pathpage_number: int- Target page numberx0, x1: float- Horizontal boundariestop, bottom: float- Vertical boundariespadding: int = 0- Padding around the crop regionresolution: int = 600- Output resolution for image extraction
Coordinate System
Pyntagma uses the standard PDF coordinate system:
- Origin (0,0) is at the bottom-left corner
- X-axis extends horizontally to the right
- Y-axis extends vertically upward
- All measurements are in points (1/72 inch)
Utility Functions
words_of_line(line: Line) -> list[Word]- Extract words from a lineline_of_word(word: Word) -> Line- Find the line containing a wordsilent_pdfplumber(path, **kwargs)- Context manager for silent PDF processing
Development
Setting up the Development Environment
# Clone the repository
git clone <repository-url>
cd pyntagma
# Install dependencies with uv
uv sync --group test
# Run tests
uv run pytest
# Run tests with coverage
uv run pytest --cov=src/pyntagma --cov-report=html
Running Tests
The test suite includes comprehensive tests for all major functionality:
# Run all tests
uv run pytest
# Run specific test file
uv run pytest tests/test_document.py
# Run with verbose output
uv run pytest -v
Test Coverage
Current test coverage: 80%
Coverage breakdown:
__init__.py: 100%document.py: 90%pdf_reader.py: 75%position.py: 77%
To generate coverage reports:
uv run pytest --cov=src/pyntagma --cov-report=html
# Open htmlcov/index.html in your browser
Requirements
- Python 3.9+
- pydantic >= 2.0.0
- pdfplumber >= 0.9.0
Development Requirements
- pytest >= 7.0.0
- pytest-cov >= 4.0.0
Contributing
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Make your changes
- Add tests for new functionality
- Ensure all tests pass (
uv run pytest) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
License
This project is licensed under the MIT License - see the LICENSE file for details.
Author
MarcellGranat - granatcellmar98@gmail.com
Acknowledgments
- Built with Pydantic for robust data validation
- PDF processing powered by pdfplumber
- Uses uv for fast Python package management
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pyntagma-0.0.1.tar.gz.
File metadata
- Download URL: pyntagma-0.0.1.tar.gz
- Upload date:
- Size: 34.7 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.7.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6a014c8e564874726f6d747e91e3b9c4197bbb39fc7b0454120e4dcf85dd46a1
|
|
| MD5 |
9fc3ffff4706050e32a5b3cc132dc0f8
|
|
| BLAKE2b-256 |
b6e4f80123a4305d01cff5492874304ec541ce6d9bf8acb878b98a66d58a8689
|
File details
Details for the file pyntagma-0.0.1-py3-none-any.whl.
File metadata
- Download URL: pyntagma-0.0.1-py3-none-any.whl
- Upload date:
- Size: 9.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.7.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
42df587c97b3d263a8417225c3761bfbbb4e8659d71c1f77f96e498b3a6b4524
|
|
| MD5 |
91dc5707b68c87ef32a58dffa0381b4f
|
|
| BLAKE2b-256 |
8e2e2bb33faee1159bd50958441e1cad892508c2f6a5d555b3aa243ea33081c6
|