openparser-sdk
Parse and extract structured data from documents with the OpenParser API.
Install the PyPI distribution openparser-sdk; import the client as openparser.
Install
pip install openparser-sdk
Use Python 3.10 or newer. Create an API key in the OpenParser dashboard.
Quick Start
import os
from pathlib import Path
from openparser import OpenParserClient
client = OpenParserClient(api_key=os.environ["OPENPARSER_API_KEY"])
result = client.parse.sync(
{"ocr_model": "paddleocr-vl-1.6", "output_format": "openparser@1"},
file=Path("invoice.pdf"),
)
print(result.page_count)
Set OPENPARSER_API_KEY to create the client without constructor arguments.
Set OPENPARSER_BASE_URL to use a different API origin.
Parse
The API holds synchronous requests for up to 300 seconds. It returns the result when processing finishes or a durable job reference when the wait expires.
# Sync: hold the connection until the result or timeout.
parsed = client.parse.sync(
{"ocr_model": "paddleocr-vl-1.6"},
file=Path("document.pdf"),
)
# Async: create a job and poll it later.
accepted = client.parse.async_(
{"ocr_model": "paddleocr-vl-1.6"},
file=Path("document.pdf"),
)
job = client.wait_for_job(accepted.id)
# Reuse a file-pool upload instead of inline bytes.
uploaded = client.files.upload(Path("document.pdf"))
parsed = client.parse.sync({"ocr_model": "paddleocr-vl-1.6", "file_id": uploaded.id})
The SDK adds an Idempotency-Key header to every parse and extract request. Pass
idempotency_key= when you need to control retries.
Extract
extracted = client.extract.sync(
{
"ocr_model": "paddleocr-vl-1.6",
"llm_model": "openai/gpt-4.1-mini",
"schema": {
"type": "object",
"properties": {"total": {"type": "number"}},
},
},
file=Path("invoice.pdf"),
)
suggested = client.extract.suggest_schema(
{
"parse_job_id": "opj_...",
"hint": "Invoice number, vendor, and total",
}
)
Grounding and lineage
Set "grounding": "field" to receive verified citations and a lineage@1
derivation DAG:
extracted = client.extract.sync(
{
"ocr_model": "mistral-ocr-4",
"llm_model": "openai/gpt-5.6-terra",
"grounding": "field",
"schema": {
"type": "object",
"properties": {"total": {"type": "number"}},
},
},
file=Path("invoice.pdf"),
)
lineage = extracted.to_dict().get("lineage")
if lineage is not None:
print(lineage["outputs"])
The graph connects each output value to its source evidence and the operations that produced it. Evidence carries the closest recognition confidence supplied by the selected OCR model. Applications can append normalization, calculation, inference, and human-review activities without replacing the original machine result.
OpenParserClient is synchronous. Methods named async_ submit durable jobs
without waiting for processing to finish.
Jobs
jobs_page = client.jobs.list(status="succeeded", limit=25)
first_job_id = jobs_page.data[0].id
job = client.jobs.get("opj_...")
parse_result = client.jobs.result("opj_...", format="openparser@1")
source_bytes = client.jobs.source("opj_...")
jobs.result() returns parse representations. Extract output is available on
the job returned by jobs.get().
Files
uploaded = client.files.upload(Path("contract.pdf"))
metadata = client.files.get(uploaded.id)
content = client.files.download(uploaded.id)
client.files.delete(uploaded.id)
Uploads accept pathlib.Path, a file handle, or {"content": bytes, "filename": str, "mime_type": str?}.
Models
ocr_models = client.models.list_ocr()
llm_models = client.models.list_llm(mode="search", q="claude")
first_ocr_model = ocr_models.data[0]
first_llm_model = llm_models.data[0]
Pipelines
pipeline = client.pipelines.create(
{
"name": "invoice-extract",
"ocr_model": "paddleocr-vl-1.6",
"llm_model": "anthropic/claude-sonnet-4",
"schema": {
"type": "object",
"properties": {"vendor": {"type": "string"}},
"required": ["vendor"],
},
}
)
listed = client.pipelines.list()
first_pipeline = listed.items[0]
current = client.pipelines.get(pipeline.id)
updated = client.pipelines.update(pipeline.id, {"name": "invoice-v2"})
client.pipelines.delete(pipeline.id)
Errors
Every non-2xx response raises a typed subclass of OpenParserError:
| HTTP | Class |
|---|---|
| 400 | OpenParserValidationError |
| 401 | OpenParserAuthError |
| 402 | OpenParserPaymentRequiredError |
| 403 | OpenParserForbiddenError |
| 404 | OpenParserNotFoundError |
| 409 | OpenParserConflictError |
| 413 | OpenParserLimitExceededError |
| 415 | OpenParserUnsupportedMediaError |
| 422 | OpenParserUnprocessableError |
| 429 | OpenParserRateLimitError |
| 503 | OpenParserServiceUnavailableError |
| 504 | OpenParserGatewayTimeoutError |
| 5xx | OpenParserServerError |
Each error exposes the API response through error.envelope: code, message,
request_id, and retryable.
Development
After changing the OpenAPI specification, regenerate the client and run its checks:
packages/openparser-sdk-python/scripts/codegen.sh
packages/openparser-sdk-python/scripts/check-codegen.sh
uv run --project packages/openparser-sdk-python pytest
License
Release files for openparser-sdk 1.0.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| openparser_sdk-1.0.3.tar.gz | 100.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| openparser_sdk-1.0.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 376.4 kB
Release files / openparser_sdk-1.0.3.tar.gz
| Download URL | openparser_sdk-1.0.3.tar.gz |
|---|---|
| Size | 100.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f79bf53dbef23441d7ed6fa991665c951332029e170ecbf634c52f13c1e5b7c0
|
|
BLAKE2b-256 checksum How to use checksums |
7bc6e9b027d4c3f1448e312ff5e489fea3c753755925b1fb7fbf8a5bad155bb8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 31, 2026.
Transparency logRelease files / openparser_sdk-1.0.3-py3-none-any.whl
| Download URL | openparser_sdk-1.0.3-py3-none-any.whl |
|---|---|
| Size | 276.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6245f88cbf03628f9b10a0138ad89724f508fd92c8c988cfef0eac5b6963194a
|
|
BLAKE2b-256 checksum How to use checksums |
fdfb750df794aa65ecfa01ad170a691a6d484dfd46525b8b983bf9739cbfca8f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 31, 2026.
Transparency log