Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

polytext

polytext

PyPI - Version PyPI Build PyPI - Downloads PyPI Downloads PyPI - Python Version

Doc Utils

A Python package for document conversion and text extraction.

Features

  • Convert various document formats (DOCX, ODT, PPT, etc.) to PDF
  • Extract text from PDF, Markdown, IMAGE, and audio files
  • Support for both local files and S3/GCS cloud storage
  • Multiple PDF parsing backends (PyPDF, PyMuPDF)
  • Transcribe audio & video files (local or cloud) to text/markdown
  • Extract YouTube video transcripts
  • Extract text from URLs

Installation

# Library only – assumes system requirements are already present
pip install polytext

Heads-up: Polytext’s PDF generator relies on [WeasyPrint] under the hood.
The PyPI wheel contains only Python code; you still need WeasyPrint’s native libraries (Pango, Cairo, GDK-PixBuf, HarfBuzz, Fontconfig) installed at the OS level.

System requirements

Requirement Notes macOS (Homebrew) Ubuntu / Debian
Python Supported on 3.11 – 3.13
WeasyPrint still requires its native libraries
brew install python@3.11 sudo apt install python3.11
WeasyPrint – native stack installs Pango, Cairo, etc. brew install weasyprint sudo apt install weasyprint
LibreOffice used for Office → PDF conversion brew install --cask libreoffice sudo apt install libreoffice

Usage

Converting Documents to PDF

from polytext import convert_to_pdf, ConversionError

try:
    # Convert a document to PDF
    pdf_path = convert_to_pdf('input.docx', 'output.pdf')
    print(f"PDF saved to: {pdf_path}")
except ConversionError as e:
    print(f"Conversion failed: {e}")

Features that require the API key for Google Gemini are:

  • audio
  • video
  • image
  • youtube
from polytext.loader.base import BaseLoader

llm_api_key = "your_google_gemini_api_key"  # Set your Google Gemini API key here

# Instantiate the loader 
loader = BaseLoader(llm_api_key=llm_api_key)

Text or Markdown Extraction

from polytext.loader.base import BaseLoader

markdown_output = False # Change if you want to extract text as markdown
source = "local" # Change to "cloud" if you want to extract from cloud storage (s3 or GCS)

# Instantiate the loader (optionally set markdown_output, llm_api_key, etc.)
loader = BaseLoader(markdown_output=markdown_output, source=source)

# Extract text from a local file
result = loader.get_text(input_list=["/path/to/document.docx"])
print(result["text"])
# Extract text from cloud file
result = loader.get_text(input_list=["s3://your-bucket/path/to/document.docx"])
print(result["text"])

# Extract text from a markdown file (local)
result = loader.get_text(input_list=["/path/to/document.md"])
print(result["text"])
# Extract text from cloud file
result = loader.get_text(input_list=["s3://your-bucket/path/to/document.md"])
print(result["text"])

# Extract text from an audio file (local)
result = loader.get_text(input_list=["/path/to/audio.mp3"])
print(result["text"])
# Extract text from cloud file
result = loader.get_text(input_list=["s3://your-bucket/path/to/audio.mp3"])
print(result["text"])

# Extract text from a video file (local)
result = loader.get_text(input_list=["/path/to/video.mp4"])
print(result["text"])
# Extract text from cloud file
result = loader.get_text(input_list=["s3://your-bucket/path/to/video.mp4"])
print(result["text"])

# Extract text from Image (local)
result = loader.get_text(input_list=["/path/to/image.jpg"])
print(result["text"])
# Extract text from cloud file
result = loader.get_text(input_list=["s3://your-bucket/path/to/image.jpg"])
print(result["text"])

# Extract transcript from a YouTube video
result = loader.get_text(input_list=["https://www.youtube.com/watch?v=xxxx"])
print(result["text"])

# Extract text from a URL
result = loader.get_text(input_list=["https://www.domain-name.com/path"])
print(result["text"])

S3 authentication

By default, Polytext uses the standard boto3 credential chain when loading s3:// inputs (environment variables, AWS profiles, IAM roles, and other boto3-supported providers).

For runtimes that need to assume an AWS role through Google OIDC, STS web identity authentication can be enabled explicitly:

from polytext.loader.base import BaseLoader

loader = BaseLoader(
    aws_auth_mode="sts_web_identity",
    aws_role_arn="arn:aws:iam::111122223333:role/ExampleRole",
    aws_region="eu-central-1",
    aws_role_session_name="polytext-session",
    gcp_id_token_audience="example-gcp-audience",
)

The same configuration can also come from environment variables:

POLYTEXT_AWS_AUTH_MODE=sts_web_identity
AWS_ROLE_ARN=arn:aws:iam::111122223333:role/ExampleRole
AWS_REGION=eu-central-1
AWS_ROLE_SESSION_NAME=polytext-session
GCP_ID_TOKEN_AUDIENCE=example-gcp-audience
GOOGLE_APPLICATION_CREDENTIALS=/absolute/path/to/service_account.json

Polytext uses the temporary STS credentials only to create the S3 client. It does not export them to os.environ and does not reset boto3's global session.

License

MIT Licence

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

polytext-0.2.8b1.tar.gz (131.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

polytext-0.2.8b1-py3-none-any.whl (111.5 kB view details)

Uploaded Python 3

File details

Details for the file polytext-0.2.8b1.tar.gz.

File metadata

  • Download URL: polytext-0.2.8b1.tar.gz
  • Upload date:
  • Size: 131.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for polytext-0.2.8b1.tar.gz
Algorithm Hash digest
SHA256 2b7f4bbf74b54aa06fdd55749afdad6af02c695b1aeaa2369e19813bc8531e9a
MD5 c600e45c40f87de9191e9153a8e8265c
BLAKE2b-256 16c81f6ff0694965c1b3118f81e2eeed763220e2029f3624974e57b37f203da1

See more details on using hashes here.

File details

Details for the file polytext-0.2.8b1-py3-none-any.whl.

File metadata

  • Download URL: polytext-0.2.8b1-py3-none-any.whl
  • Upload date:
  • Size: 111.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for polytext-0.2.8b1-py3-none-any.whl
Algorithm Hash digest
SHA256 9ddf2f5fb2031831b5d08f99ecfee7024187e635d4c83d0584d6877b93d0d7fa
MD5 cbdd7fbaa5b6551e40f1c39dffc19ee8
BLAKE2b-256 c860a79d8b26918c9bfc3d6b5e5ff8681e46bc9a780c3838898bb6a2a9e41213

See more details on using hashes here.

Release history Release notifications | RSS feed

0.2.8

2 files

This release

0.2.8b1 This release

2 files

0.2.7

2 files

0.2.6

2 files

0.2.5

2 files

0.2.4

2 files

0.2.3

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.4

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page