Swarmauri Parser PyPDFTK
Form-field parser for Swarmauri built on PyPDFTK. Extracts PDF AcroForm field metadata and returns it as Swarmauri Document content.
Features
- Calls
pypdftk.dump_data_fieldsto extract field key/value pairs. - Emits a single
Documentwith newline-delimitedkey: valuetext andmetadata['source']set to the PDF path. - Returns an empty list when no form fields exist or when parsing fails (logs the error).
Prerequisites
- Python 3.10 or newer.
- PyPDFTK plus the
pdftk/pdftk-javabinary available on the system path. Install operating-system packages: e.g.,apt install pdftk-javaor downloadpdftkfor macOS/Windows. - Read access to the PDF file path you provide.
Installation
# pip
pip install swarmauri_parser_pypdftk
# poetry
poetry add swarmauri_parser_pypdftk
# uv (pyproject-based projects)
uv add swarmauri_parser_pypdftk
Quickstart
from swarmauri_parser_pypdftk import PyPDFTKParser
parser = PyPDFTKParser()
documents = parser.parse("forms/enrollment.pdf")
for doc in documents:
print(doc.metadata["source"])
print(doc.content)
Example output:
source: forms/enrollment.pdf
GivenName: John
FamilyName: Doe
BirthDate: 1990-01-01
Handling Missing Fields
parser = PyPDFTKParser()
docs = parser.parse("forms/plain.pdf")
if not docs:
print("No form fields detected or parsing failed.")
Tips
- Ensure
pdftkis installed and available onPATH; PyPDFTK delegates to the binary. - For encrypted PDFs, remove or provide the password before parsing;
pdftkcannot dump fields from password-protected documents without credentials. - Combine with other Swarmauri parsers to extract both structured form data (
PyPDFTKParser) and free-form text (PyPDF2ParserorFitzPdfParser).
Want to help?
If you want to contribute to swarmauri-sdk, read up on our guidelines for contributing that will help you get started.
Metadata
Release files for swarmauri_parser_pypdftk 0.9.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| swarmauri_parser_pypdftk-0.9.0.tar.gz | 7.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| swarmauri_parser_pypdftk-0.9.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 15.5 kB
Release files / swarmauri_parser_pypdftk-0.9.0.tar.gz
| Download URL | swarmauri_parser_pypdftk-0.9.0.tar.gz |
|---|---|
| Size | 7.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
baf0efbfd75723f51d256e4efddc50f550f335a9fc63552789b1b4590829c5dc
|
|
BLAKE2b-256 checksum How to use checksums |
782d4143416abc0c08a202acf95d222723448b86fc1ae149eb6a15a52c157fad
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.0 {"installer":{"name":"uv","version":"0.11.0","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / swarmauri_parser_pypdftk-0.9.0-py3-none-any.whl
| Download URL | swarmauri_parser_pypdftk-0.9.0-py3-none-any.whl |
|---|---|
| Size | 8.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
fb2593675ef5e7605eb41e202cebcd69e274f8f9962dc4a8cda61e39590159b1
|
|
BLAKE2b-256 checksum How to use checksums |
668dde17f88d7e5167c7161b8c7da960710b048215a1e2bfe6ba10a3bdb0d242
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.0 {"installer":{"name":"uv","version":"0.11.0","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|