paperless-epub-parser
An EPUB parser plugin for Paperless-ngx —
makes .epub files consumable and full-text searchable.
No fork of Paperless-ngx. The upstream project exposes a supported parser
plugin mechanism: any Python package that advertises a class under the
paperless_ngx.parsers entry-point group is discovered automatically at
startup and competes with the built-in parsers on score. This package is that,
and nothing more.
Why
Paperless-ngx has no built-in EPUB support. Neither Apache Tika (whose Paperless mime list is hardcoded to Office formats) nor Gotenberg/LibreOffice (which has an EPUB export filter but no import filter) can open one. This plugin supplies the missing parser, so an EPUB library becomes searchable alongside everything else in the archive.
Scope is deliberately narrow: text for search. No archive PDF, no cover-art thumbnail, no page count. See "Design decisions" below.
Install
The container's site-packages is root-owned and the image is ephemeral, so a
hand-run pip install inside a running container does not persist. Build a thin
derived image instead:
FROM ghcr.io/paperless-ngx/paperless-ngx:3.1.3
RUN pip install --no-cache-dir paperless-epub-parser
Then point the webserver service at it:
services:
webserver:
build: .
# ... rest of your existing service definition
⚠️ Do not add a USER instruction to this Dockerfile. It is the obvious
thing to write and it will break your instance. See "The USER trap" below.
⚠️ Pin the base image tag. The upstream latest tag moves; a rebuild months
later would silently install the plugin into a different Paperless version than
the one you tested against.
Verify discovery in the startup log:
Loaded third-party parser 'EPUB' v1.1.0 by David Newman (entrypoint: 'epub').
The USER trap
If you write the natural-looking Dockerfile — USER root, install, USER paperless — the container will not start. It exits cleanly, forever:
Restarts=3, climbing · ExitCode=0 · no error output
That signature is the confusing kind. A clean exit with no message reads like "nothing to do", not "broken", so it is easy to spend a while looking in the wrong place.
The cause: the official image declares no USER at all.
docker inspect ghcr.io/paperless-ngx/paperless-ngx:3.1.3 --format '{{.Config.User}}'
# -> (empty)
It runs /init (s6-overlay) as root and drops privileges internally, to
the uid/gid given by USERMAP_UID / USERMAP_GID. Setting USER paperless
yourself pre-empts that handoff, and init exits.
So: install your package and stop. Add no USER line; the base image owns
privilege dropping.
If you already have a crashing instance: Restarts climbing, ExitCode=0, no
errors — remove the USER instruction and rebuild.
Design decisions
Text extraction is delegated to markitdown.
Its EpubConverter already implements the genuinely fiddly parts: locating the
OPF package document via META-INF/container.xml, resolving manifest hrefs
(including percent-encoded ones), walking the spine in reading order rather
than filename order, and converting each XHTML chapter to Markdown with headings
and tables preserved. Reimplementing that would duplicate a maintained upstream.
EPUB conversion lives in markitdown's base install — the converter subclasses
HtmlConverter and needs only beautifulsoup4, which is a core dependency.
There is no [epub] extra and none is required, so this plugin depends on bare
markitdown and deliberately does not pull in the PDF/DOCX/PPTX/XLSX extras
(Paperless already handles those formats natively).
No archive PDF. EPUB is the readable artifact; it belongs in an e-reader, not
in a PDF viewer. can_produce_archive and requires_pdf_rendition are both
False, so the consumer skips PDF generation entirely.
The thumbnail is rendered text, not cover art. This follows the precedent set
by Paperless-ngx's own TextDocumentParser. Cover extraction would mean adding
ebooklib plus, for the common case of an SVG or fixed-layout cover, an SVG
rasteriser — two dependencies and a whole class of failure (missing, oversized,
malformed, DRM-wrapped covers) in exchange for a nicer grid tile.
get_page_count() returns None. EPUB is a reflowable format. It has no
pages. Any integer would be an invention.
get_date() returns None. EPUB dc:date is unreliable in practice — often
a build timestamp rather than a publication date, and absent altogether from the
fixtures this plugin was validated against. Paperless's own date parser reads the
filename and extracted text and does better than a blind metadata grab.
Fully local. The parser declares uses_remote_service = False and never
contacts anything off-box.
Known upstream defect, and the correction
markitdown's EpubConverter locates the spine with
getElementsByTagName("itemref") and the manifest with
getElementsByTagName("item"). Both match on the qualified name, so on an
OPF written with namespace prefixes — <opf:itemref>, <opf:item> — those
lookups return zero elements. The converter then finds no spine, emits only
the metadata block, and returns a few hundred characters for a book of several
hundred thousand. Nothing raises.
That syntax is legal EPUB 2 and is emitted by several publisher toolchains. In a 795-book library it affected exactly one book — which is precisely why it was worth fixing: a hole that size is invisible when you read the output.
This plugin detects that specific condition before delegating, and takes a local
spine-walking path instead (_epub.py). Detection is the precise condition, not
a heuristic on output length, so books markitdown handles correctly are never
rerouted. Only the spine walk is replaced — each document is still converted by
markitdown's own HtmlConverter, so HTML-to-Markdown behaviour is unchanged.
Measured on the affected book: 203 → 157,367 characters, with all other books byte-identical.
Non-conformant EPUBs detected as application/zip
Some EPUBs are rejected by Paperless before this — or any — parser is consulted:
ConsumerError: book.epub: Unsupported mime type application/zip
The book is fine. Its container is not. The OCF specification requires a
mimetype entry as the first record in the archive, stored uncompressed, so
that a reader can identify the format from the opening bytes. Plenty of real
files violate this — some list a directory entry first, some put mimetype
third.
Typesetting it correctly matters because Paperless identifies files with
libmagic, which reads the leading bytes and reports application/zip for a
misordered archive. The consumer then asks the parser registry for a handler for
application/zip, and — note the ordering — a parser whose declared MIME types
omit the detected type is skipped outright, so its score() is never called.
That ordering is why this plugin declares application/zip at all. It is not a
claim to handle ZIP files; it is the only way to be offered the file. The claim
is then gated inside score(), which opens the archive and requires both
structural markers together:
- a
META-INF/container.xmlentry, and - a
mimetypeentry whose content is exactlyapplication/epub+zip.
The mimetype entry's position is deliberately ignored, since position is
precisely the thing that varies.
Because application/zip is a busy type, the claim is scored 1, not 10 — an
archive must never out-rank a parser that genuinely handles whatever the file is.
An ODT, for instance, also carries a mimetype entry; only its value differs,
and it is declined. docx, xlsx, jar, apk and plain zip are all declined.
If your archive does not need ZIP files handled by anything else, this is transparent. If you run a plugin that legitimately handles ZIP archives, it will out-score this one and win.
Failure behaviour
- Malformed / unreadable file →
ParseErrornaming the file, so Paperless reports a document error rather than an "unexpected error". - DRM-protected or fixed-layout EPUB → no extractable flow text, which is
raised as a
ParseErrorwith a message that says so. It is not silently filed as an empty document. - Thumbnail generation never raises. The consumer calls it unguarded, so a failure there would abort an otherwise successful consumption.
Development
uv venv .venv --python 3.12
uv pip install --python .venv/bin/python -e ".[dev]"
.venv/bin/python -m pytest -v
The suite stubs the conversion seam (_extract_epub), so it runs without
markitdown, Paperless-ngx, or Django installed. It covers the registry
contract, the protocol surface, the failure paths, and the thumbnail guarantee.
Real conversion is verified separately against genuine EPUBs — see below.
Verification
Unit tests prove the contract. They do not prove that markitdown extracts
useful text from a real book. To check that, and to confirm end-to-end discovery
by Paperless:
# 1. Extraction, standalone
uv venv /tmp/v && uv pip install --python /tmp/v/bin/python markitdown
/tmp/v/bin/python -c "
from markitdown import MarkItDown
r = MarkItDown().convert('book.epub')
print(len(r.text_content or r.markdown or ''))
"
# 2. Discovery, via the container's startup log
docker compose logs webserver | grep 'Loaded third-party parser'
# 3. End-to-end: drop an EPUB into consume/ and watch it become searchable
cp book.epub /path/to/paperless/consume/
Compatibility
- Paperless-ngx 3.1.3 (the parser registry landed in the 3.x series; this
plugin targets the
ParserProtocolinterface as it exists in 3.1.3). - Python 3.11+ (the image ships 3.14;
markitdownsupports 3.10–3.14).
Reporting problems
Please include the book that misbehaved, or at least:
- the exact error from
docker compose logs webserver - the output of
file yourbook.epub - the first few entries of
unzip -l yourbook.epub
Those three lines distinguish "the container is non-conformant" from "the text extraction failed", which are different problems with different fixes.
Licence
MIT — see LICENSE.
Copyright (c) 2026 David Newman and Guan.
Release files for paperless-epub-parser 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| paperless_epub_parser-1.1.0.tar.gz | 26.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| paperless_epub_parser-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 45.4 kB
Release files / paperless_epub_parser-1.1.0.tar.gz
| Download URL | paperless_epub_parser-1.1.0.tar.gz |
|---|---|
| Size | 26.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
422278489368e309a60abc304464b41518843ffb0799ad9bb5b2ea87329742fd
|
|
BLAKE2b-256 checksum How to use checksums |
847026779ad26156d4b2c6945700162bcb4d71aa2a7cac5a6fed0ebe9152dd09
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency logRelease files / paperless_epub_parser-1.1.0-py3-none-any.whl
| Download URL | paperless_epub_parser-1.1.0-py3-none-any.whl |
|---|---|
| Size | 18.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3e159ef4fde81742486bccfc388b09aa5beb54ea976c0f9a0f88dc400d2ca57f
|
|
BLAKE2b-256 checksum How to use checksums |
f3015d6c453b96f4bfe31eda4b9c27b6263b5c132e3e8ff82e5f3c98884507d2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency log