Skip to main content

A Python client for the ReadPDFs API

Project description

ReadPDFs

A Python client for the ReadPDFs API that allows you to process PDF files and convert them to markdown.

Installation

pip install readpdfs

Usage

Basic Client Usage

from readpdfs import ReadPDFs

# Initialize the client
client = ReadPDFs(api_key="your_api_key")

# Process a PDF from a URL
result = client.process_pdf(pdf_url="https://example.com/document.pdf")

# Process a local PDF file
result = client.process_pdf(file_path="path/to/local/document.pdf")

# Process from file content
with open("document.pdf", "rb") as f:
    content = f.read()
    result = client.process_pdf(file_content=content, filename="document.pdf")

# Fetch markdown content
markdown = client.fetch_markdown(url="https://api.readpdfs.com/documents/123/markdown")

# Get user documents
documents = client.get_user_documents(clerk_id="user_123")

FastAPI Integration

from fastapi import FastAPI, File, UploadFile, HTTPException
from readpdfs import ReadPDFs
from typing import Optional

app = FastAPI()
client = ReadPDFs(api_key="your_api_key")

@app.post("/process-pdf")
async def process_pdf(
    pdf_url: Optional[str] = None,
    file: Optional[UploadFile] = File(None),
    quality: str = "standard"
):
    try:
        if pdf_url and file:
            raise HTTPException(
                status_code=400,
                detail="Provide either pdf_url or file, not both"
            )
            
        if pdf_url:
            result = client.process_pdf(pdf_url=pdf_url, quality=quality)
        elif file:
            content = await file.read()
            result = client.process_pdf(
                file_content=content,
                filename=file.filename,
                quality=quality
            )
        else:
            raise HTTPException(
                status_code=400,
                detail="Either pdf_url or file must be provided"
            )
            
        return result
        
    except Exception as e:
        raise HTTPException(status_code=500, detail=str(e))

Features

  • Process PDFs from URLs, local files, or file content
  • Convert PDFs to markdown
  • Fetch markdown content
  • Retrieve user documents
  • Configurable processing quality
  • FastAPI integration support

Requirements

  • Python 3.7+
  • requests library

For FastAPI integration:

  • fastapi
  • python-multipart
  • uvicorn

API Examples

cURL

# Process PDF from URL
curl -X POST "http://localhost:8000/process-pdf?pdf_url=https://example.com/document.pdf&quality=high"

# Upload PDF file
curl -X POST "http://localhost:8000/process-pdf?quality=high" \
  -H "Content-Type: multipart/form-data" \
  -F "file=@/path/to/local/document.pdf"

Python Requests

import requests

# URL method
response = requests.post(
    "http://localhost:8000/process-pdf",
    params={"pdf_url": "https://example.com/document.pdf", "quality": "high"}
)

# File upload method
with open("document.pdf", "rb") as f:
    response = requests.post(
        "http://localhost:8000/process-pdf",
        params={"quality": "high"},
        files={"file": f}
    )

result = response.json()

License

This project is licensed under the MIT License - see the LICENSE file for details.


This update:
1. Added the new file content processing method
2. Included a complete FastAPI integration example
3. Added API examples using cURL and Python requests
4. Updated the requirements section to include FastAPI-related packages
5. Reorganized the usage section to separate basic client usage from FastAPI integration

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

readpdfs-0.1.3.tar.gz (4.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

readpdfs-0.1.3-py3-none-any.whl (4.9 kB view details)

Uploaded Python 3

File details

Details for the file readpdfs-0.1.3.tar.gz.

File metadata

  • Download URL: readpdfs-0.1.3.tar.gz
  • Upload date:
  • Size: 4.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.10.6

File hashes

Hashes for readpdfs-0.1.3.tar.gz
Algorithm Hash digest
SHA256 e1823c6e6bf2c2812d0a0faeaedf3b82b2047ea418b54581ba8142f60930ec61
MD5 f8d8c4ba3880aa1a2e7af6a4ae874c24
BLAKE2b-256 0abc9ea1ec012019f19b12e529679ca3688b09cedfb188116a692f30dfb3ae30

See more details on using hashes here.

File details

Details for the file readpdfs-0.1.3-py3-none-any.whl.

File metadata

  • Download URL: readpdfs-0.1.3-py3-none-any.whl
  • Upload date:
  • Size: 4.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.10.6

File hashes

Hashes for readpdfs-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 720dd46a42393c2cb6623e026074d2cd7a429e1ec3e4d90d31560a148e8a64d1
MD5 0d82ecd8eb71b3016b2e750f550929c5
BLAKE2b-256 748f9070fce7b42ab877988bd007308084c94b48b430880cbd3c0d7e726327fa

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page