vlmparse
[[📜 arXiv]] | [Dataset (🤗Hugging Face)] | [pypi] | [vlmparse] | [Benchmark] | [Leaderboard]
A unified wrapper for Vision Language Models (VLM) and OCR solutions to parse PDF documents into Markdown.
Features:
- ⚡ Async/concurrent processing for high throughput
- 🐳 Automatic Docker server management for local models
- 🔄 Unified interface across all VLM/OCR providers
- 📊 Built-in result visualization with Streamlit
Supported Converters:
- Open Source Small VLMs:
LightOnOCR-1B-1025,LightOnOCR-2-1B,MinerU2.5-2509-1.2B,HunyuanOCR,PaddleOCR-VL-1.5,granite-docling-258M,olmOCR-2-7B-1025-FP8,dots.ocr,dots.ocr-1.5,dots.mocr,chandra,chandra-ocr-2,DeepSeek-OCR,DeepSeek-OCR-2,Nanonets-OCR2-3B,GLM-OCR,FireRed-OCR,OCRVerse,Qianfan-OCR - Open Source Generalist VLMs: such as the Qwen family.
- Pipelines:
docling - Proprietary LLMs:
gemini,gpt
Installation
Simplest solution with only the cli:
uv tool install vlmparse
If you want to run the granite-docling model or use the streamlit viewing app:
uv tool install vlmparse[docling_core,st_app]
If you prefer cloning the repository and using the local version:
uv sync
With optional dependencies:
uv sync --all-extras
Activate the virtual environment:
source .venv/bin/activate
CLI Usage
Note that you can bypass the previous installation step and just add uvx before each of the commands below.
Convert PDFs
With a general VLM (requires setting your api key as an environment variable):
vlmparse convert "*.pdf" -o ./output --model gemini-2.5-flash-lite
Convert with auto deployment of a small vlm (or any huggingface VLM model, requires a gpu + docker installation):
vlmparse convert "*.pdf" -o ./output --model nanonets/Nanonets-OCR2-3B
Deploy a local model server
Deployment (requires a gpu + docker installation):
- You need a gpu dedicated for this.
- Check that the port is not used by another service.
vlmparse serve lightonocr2 --port 8000 --gpu 1
then convert:
vlmparse convert "*.pdf" -o ./output --uri http://localhost:8000/v1
You can also list all running servers:
vlmparse list
You can get a list of registered models and their capabilities (ocr, ocr_layout, table, image_description) and inline image description with:
vlmparse registry
Show logs of a server (if only one server is running, the container name is not needed):
vlmparse log <container_name>
Stop a server (if only one server is running, the container name is not needed):
vlmparse stop <container_name>
View conversion results with Streamlit
vlmparse view ./output
Configuration
Set API keys as environment variables:
export GOOGLE_API_KEY="your-key"
export OPENAI_API_KEY="your-key"
Python API
Client interface:
from vlmparse.registries import converter_config_registry
# Get a converter configuration
config = converter_config_registry.get("gemini-2.5-flash-lite")
client = config.get_client()
# Convert a single PDF
document = client("path/to/document.pdf")
print(document.to_markdown())
# Batch convert multiple PDFs
documents = client.batch(["file1.pdf", "file2.pdf"])
Docker server interface:
from vlmparse.registries import docker_config_registry
config = docker_config_registry.get("lightonocr")
server = config.get_server()
server.start()
# Client calls...
server.stop()
Converter with automatic server management:
from vlmparse.converter_with_server import ConverterWithServer
with ConverterWithServer(model="mineru25") as converter_with_server:
documents = converter_with_server.parse(inputs=["file1.pdf", "file2.pdf"], out_folder="./output")
Note that if you pass an uri of a vllm server to ConverterWithServer, the model name is inferred automatically and no server is started.
Credits
This work was realised by members of Probayes and OpenValue, two subsidiaries of La Poste.
Metadata
Release files for vlmparse 0.1.28
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vlmparse-0.1.28.tar.gz | 117.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vlmparse-0.1.28-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 254.0 kB
Release files / vlmparse-0.1.28.tar.gz
| Download URL | vlmparse-0.1.28.tar.gz |
|---|---|
| Size | 117.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
68eef0bb3b0bd2d962bb00b114847b0d1cc0553db253e9a8f080d6edb4917464
|
|
BLAKE2b-256 checksum How to use checksums |
b590552089124fcb6f6af0c6b765559ef1165720f938b66a9904900dd52e7254
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / vlmparse-0.1.28-py3-none-any.whl
| Download URL | vlmparse-0.1.28-py3-none-any.whl |
|---|---|
| Size | 136.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
be5e89d1c5b061a5ff860eff5c30f0f0ab788a79460a0441b0d540446ff6f342
|
|
BLAKE2b-256 checksum How to use checksums |
6d7a970e8ac70b8eab61a0c162d3b9eb9590c93d52138a57e2dd71b5c18ca329
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|