Open-Source Pre-Processing Tools for Unstructured Data
The unstructured-inference repo contains hosted model inference code for layout parsing models.
These models are invoked via API as part of the partitioning bricks in the unstructured package.
Requires Python >=3.11, <3.14.
Installation
Package
pip install unstructured-inference
Detectron2
Detectron2 is required for using models from the layoutparser model zoo but is not automatically installed with this package. For MacOS and Linux, build from source with:
pip install 'git+https://github.com/facebookresearch/detectron2.git@57bdb21249d5418c130d54e2ebdc94dda7a4c01a'
Other install options can be found in the Detectron2 installation guide.
Windows is not officially supported by Detectron2, but some users are able to install it anyway. See discussion here for tips on installing Detectron2 on Windows.
Development Setup
This project uses uv for dependency management.
# Clone and install all dependencies (including dev/test/lint groups)
git clone https://github.com/Unstructured-IO/unstructured-inference.git
cd unstructured-inference
make install
Run make help for a full list of available targets.
Getting Started
To get started with the layout parsing model, use the following commands:
from unstructured_inference.inference.layout import DocumentLayout
layout = DocumentLayout.from_file("sample-docs/loremipsum.pdf")
print(layout.pages[0].elements)
Once the model has detected the layout and OCR'd the document, the text extracted from the first
page of the sample document will be displayed.
You can convert a given element to a dict by running the .to_dict() method.
Models
The inference pipeline operates by finding text elements in a document page using a detection model, then extracting the contents of the elements using direct extraction (if available), OCR, and optionally table inference models.
We offer several detection models including Detectron2 and YOLOX.
Using a non-default model
When doing inference, an alternate model can be used by passing the model object to the ingestion method via the model parameter. The get_model function can be used to construct one of our out-of-the-box models from a keyword, e.g.:
from unstructured_inference.models.base import get_model
from unstructured_inference.inference.layout import DocumentLayout
model = get_model("yolox")
layout = DocumentLayout.from_file("sample-docs/layout-parser-paper.pdf", detection_model=model)
Using your own model
Any detection model can be used for in the unstructured_inference pipeline by wrapping the model in the UnstructuredObjectDetectionModel class. To integrate with the DocumentLayout class, a subclass of UnstructuredObjectDetectionModel must have a predict method that accepts a PIL.Image.Image and returns a list of LayoutElements, and an initialize method, which loads the model and prepares it for inference.
Security Policy
See our security policy for information on how to report security vulnerabilities.
Learn more
| Section | Description |
|---|---|
| Unstructured Community Github | Information about Unstructured.io community projects |
| Unstructured Github | Unstructured.io open source repositories |
| Company Website | Unstructured.io product and company info |
Release files for unstructured-inference 1.6.13
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| unstructured_inference-1.6.13.tar.gz | 50.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| unstructured_inference-1.6.13-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 107.6 kB
Release files / unstructured_inference-1.6.13.tar.gz
| Download URL | unstructured_inference-1.6.13.tar.gz |
|---|---|
| Size | 50.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4efa2f3f517370aa09166c669cdc1addee71531e95bfbec7e6e3be66be90c147
|
|
BLAKE2b-256 checksum How to use checksums |
2ece1e58008b8416884edfa6d18d65b635cde892189512d2e7531f91553b34d0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 11, 2026.
Transparency logRelease files / unstructured_inference-1.6.13-py3-none-any.whl
| Download URL | unstructured_inference-1.6.13-py3-none-any.whl |
|---|---|
| Size | 57.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
73b727260e4a43133276727c371b57f7d4d92c68e8a7529a28e9576755356524
|
|
BLAKE2b-256 checksum How to use checksums |
01c06598b1420605382e7ed879ff7e1d51e12cc7a82e46971c7a61c3740c77db
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 11, 2026.
Transparency log