Docuparser
Docuparser is an extensible Python library for document intelligence. It uses Docling for document conversion and provides optional extensions for PII masking, embeddings, structured table extraction, summaries, and knowledge graphs.
Requirements
- Python 3.10 or newer
- A document supported by Docling
- An LLM endpoint for summaries, table extraction, and graph relationships
- Optional feature dependencies for embeddings, PII analysis, and graphs
Installation
Create and activate a virtual environment:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
Install the core package:
python -m pip install -e .
Install optional features as needed:
python -m pip install -e ".[pii]"
python -m pip install -e ".[embeddings]"
python -m pip install -e ".[graph]"
python -m pip install -e ".[all]"
The package has a regex fallback for common PII patterns when the pii extra
is not installed. The embeddings and graph extensions require their
corresponding extras.
LLM Configuration
Docuparser accepts an OpenAI-compatible configuration dictionary. Set the API key in the environment rather than committing it to source control:
$env:LLM_API_KEY = "your-api-key"
from docuparser import Docuparser
parser = Docuparser(
file_path="sample.pdf",
llm={
"base_url": "https://api.openai.com/v1",
"model_name": "gpt-4o-mini",
"temperature": 0.0,
},
)
You can provide base_url, model_name, api_key, temperature, and
extra_headers for a compatible LLM endpoint. A custom BaseLLMClient,
callable, or compatible LangChain chat model can also be supplied.
Basic Usage
Parsing is lazy when an extension or get_markdown() is first used. Call
parse() explicitly when you want to perform conversion up front.
from docuparser import Docuparser
parser = Docuparser("sample.pdf")
parser.parse()
print(parser.get_markdown())
The parsed state is available through parser.doc_context:
markdown: exported document textraw_document: the underlying Docling documenttables: extracted raw tableschunks: generated chunks and vectorsartifacts: extension outputsmetadata: document metadata
PII Masking
from docuparser import Docuparser
parser = Docuparser("sample.pdf")
parser.mask_pii()
print(parser.get_markdown())
To restrict masking to selected entity types:
parser.mask_pii(entities=["EMAIL_ADDRESS", "PHONE_NUMBER", "US_SSN"])
Supported configured entity names include PERSON, EMAIL_ADDRESS,
PHONE_NUMBER, CREDIT_CARD, CRYPTO, IP_ADDRESS, and US_SSN.
Embeddings
Install .[embeddings] first. The extension uses Docling's HybridChunker
and Sentence Transformers.
from docuparser import Docuparser
parser = Docuparser("sample.pdf")
chunks = parser.to_embeddings(
model_name="all-MiniLM-L6-v2",
chunk_size=512,
device="cpu",
)
for chunk in chunks:
print(chunk["chunk_id"], len(chunk["vector"]))
The model may be downloaded the first time it is used. Use device="cuda"
in a CUDA-capable PyTorch environment for GPU execution.
Structured Tables
Table extraction requires an LLM and a Pydantic schema:
from pydantic import BaseModel
from docuparser import Docuparser
class InvoiceItem(BaseModel):
item: str
amount: float
currency: str = "USD"
parser = Docuparser("invoice.pdf", llm={
"base_url": "https://api.openai.com/v1",
"model_name": "gpt-4o-mini",
})
tables = parser.get_clean_tables(InvoiceItem)
for table in tables:
for row in table:
print(row.model_dump())
Rows that fail Pydantic validation are skipped. Results are also available at
parser.doc_context.artifacts["structured_tables"].
Summaries
summary = parser.get_summary(max_characters=12000)
print(summary)
This requires an LLM. Without one, the extension stores an explanatory
message in the summary artifact.
Knowledge Graphs
Install .[graph] first. GLiNER extracts entities, and the configured LLM
identifies relationships between those entities.
from docuparser import Docuparser
parser = Docuparser("report.pdf", llm={
"base_url": "https://api.openai.com/v1",
"model_name": "gpt-4o-mini",
})
graph = parser.get_knowledge_graph(
labels=["Organization", "Person", "Location", "Product"]
)
print(graph.to_dict())
for query in graph.to_cypher():
print(query)
to_cypher() returns MERGE and MATCH statements suitable for review or
submission to a compatible graph database. Review generated queries before
executing them against a database.
Pipelines and Custom Extensions
Built-in extensions can be composed with pipe() or run_pipeline():
from docuparser import Docuparser
parser = Docuparser("sample.pdf", llm={
"base_url": "https://api.openai.com/v1",
"model_name": "gpt-4o-mini",
})
parser.run_pipeline(["pii_masking", "summary"])
print(parser.doc_context.artifacts["summary"])
Custom extensions implement BaseExtension and return the updated
ParsedDocument:
from docuparser import BaseExtension, Docuparser, ParsedDocument
class WordCountExtension(BaseExtension):
def run(self, doc: ParsedDocument, llm=None) -> ParsedDocument:
doc.artifacts["word_count"] = len(doc.markdown.split())
return doc
parser = Docuparser("sample.pdf")
parser.pipe(WordCountExtension())
print(parser.doc_context.artifacts["word_count"])
Registered extensions can be inspected with:
from docuparser import ExtensionRegistry
print(ExtensionRegistry.list_available())
Development
Run the test suite from the project directory:
python -m pytest
Tests use mocks for external services where possible. Real LLM, Docling, GLiNER, and embedding model downloads may require network access and system resources.
License
Docuparser is open-source software released under the Apache License 2.0.
Copyright 2026.
Release files for docuparser 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| docuparser-0.1.0.tar.gz | 16.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| docuparser-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 33.6 kB
Release files / docuparser-0.1.0.tar.gz
| Download URL | docuparser-0.1.0.tar.gz |
|---|---|
| Size | 16.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b45b8afb3b3ef2a688411ad4d02e5bd5733575087c28ecf98d5949ed2520a28e
|
|
BLAKE2b-256 checksum How to use checksums |
75f1d1a42656deb3b313dd28e43e67f06962747a7981810da5b88555d3895154
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / docuparser-0.1.0-py3-none-any.whl
| Download URL | docuparser-0.1.0-py3-none-any.whl |
|---|---|
| Size | 17.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7e8e7a821796af86f8c148b4370a0292063b1d3c5d7881de79ee222d3fc858f8
|
|
BLAKE2b-256 checksum How to use checksums |
ef76046bd5a43f02c26a35d9c28fa4100115042399246838da7a053c10928a6c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|