llama-index-readers-hexread
A LlamaIndex reader for HexRead. It
converts PDFs and images to Markdown through the HexRead API and returns LlamaIndex Documents.
Nothing is processed locally.
API access requires a paid HexRead plan. The free trial is web only, so no API key can be issued for it. Keys are created and revoked in your HexRead dashboard.
Install
pip install llama-index-readers-hexread
Python 3.10+. Pulls in hexread and llama-index-core.
Quick start
from llama_index.readers.hexread import HexReadReader
docs = HexReadReader().load_data("report.pdf") # one Document per page
Set HEXREAD_API_KEY, or pass api_key= to the reader. The API client is built on the first
conversion, never at construction, so a reader can be created before a key is available.
file is a path, raw bytes, or an open binary file.
Reader options
| Option | Default | Effect |
|---|---|---|
api_key |
environment, then CLI credential | key for this reader |
model |
auto |
parser to request; naming one requires a plan that allows it |
lang |
unset | OCR language hint, passed through to the parser |
split |
"page" |
"page" for one Document per page, "file" for one per document |
base_url |
https://api.hexread.com/v1 |
API base URL |
client |
built on first use | an existing HexRead or AsyncHexRead to convert with |
extra_metadata |
{} |
extra keys merged into every Document's metadata |
A reader holding a HexRead serves the sync methods; one holding an AsyncHexRead serves
aload_data. Using the wrong one raises TypeError with the fix in the message.
Methods
| Method | Returns | Notes |
|---|---|---|
load_data(file, extra_info=None) |
list[Document] |
converts one file |
lazy_load_data(file, extra_info=None) |
Iterator[Document] |
converts when the first Document is pulled |
aload_data(file, extra_info=None) |
list[Document] |
async; needs an AsyncHexRead or no injected client |
load_data_many(files, *, max_workers=2, extra_info=None, raise_on_error=False) |
list[Document] |
several files, a few at a time, in input order |
extra_info is merged underneath the reader's own keys, so a SimpleDirectoryReader's
file_name and creation_date survive while source and page stay authoritative.
load_data_many keeps the batch alive when one file fails: by default the failure is logged and
that file's Documents are skipped. Pass raise_on_error=True to abort instead. Keep max_workers
at or below the concurrency your plan allows, or the API answers 429.
reader = HexReadReader()
docs = reader.load_data_many(["a.pdf", "b.pdf", "scan.png"], max_workers=2)
Async, with a client the reader does not own and so does not close:
from hexread import AsyncHexRead
from llama_index.readers.hexread import HexReadReader
async with AsyncHexRead() as client:
docs = await HexReadReader(client=client).aload_data("report.pdf")
With no injected client, aload_data builds one for the call and closes it before returning.
Document metadata
| Key | Value |
|---|---|
source |
the path that was converted |
page |
page index, 0-based (absent when split="file") |
page_label |
page number as a string, 1-based (absent when split="file") |
total_pages |
pages in the converted document |
model |
parser that produced the Markdown |
route_reason |
why the auto router picked that parser (absent when a model was requested) |
parser |
always hexread |
Each Document also gets a stable id_: "<source>:<page index>" per page, or "<source>" when
split="file", so re-ingesting a document updates in place instead of duplicating.
A whole directory
HexReadReader works as a SimpleDirectoryReader file extractor:
from llama_index.core import SimpleDirectoryReader
from llama_index.readers.hexread import HexReadReader
reader = HexReadReader()
extractor = {ext: reader for ext in (".pdf", ".png", ".jpg", ".jpeg", ".tiff", ".webp")}
docs = SimpleDirectoryReader("./contracts", file_extractor=extractor).load_data()
License
Licensed under the MIT License, © HexWorld Solutions GmbH.
Source, issues and the core client: github.com/HexWorldEU/hexread-python.
Release files for llama-index-readers-hexread 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llama_index_readers_hexread-0.1.0.tar.gz | 6.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llama_index_readers_hexread-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 13.7 kB
Release files / llama_index_readers_hexread-0.1.0.tar.gz
| Download URL | llama_index_readers_hexread-0.1.0.tar.gz |
|---|---|
| Size | 6.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a2eb60c9b3a3a3d53635fc0ab86da4ff0cf76df3ad0f9b656cd32a44b1962b43
|
|
BLAKE2b-256 checksum How to use checksums |
706fb214523f5eaf4cf26e2d2094820a7bd1fde7292b2f1320aca0a9b9e23640
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 5, 2026.
Transparency logRelease files / llama_index_readers_hexread-0.1.0-py3-none-any.whl
| Download URL | llama_index_readers_hexread-0.1.0-py3-none-any.whl |
|---|---|
| Size | 7.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
60372eb4c7999b8713bbf53b7e7f018a4a4c7baaa87126fa96b0e30224daf547
|
|
BLAKE2b-256 checksum How to use checksums |
d3c3bf1d51444394e302a305730982d13c5b85e4d2575c45c042b00e9300e3a2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 5, 2026.
Transparency log