Deterministic EPUB normalizer and tooling.
Project description
SimpleBook
Status: active (deterministic EPUB normalizer + tools).
Quick start
source ./activate
regtest --list
simplebook tests/epubs/the-hobbit.epub --preview -o /tmp/the-hobbit.json
simplebook tests/epubs/the-hobbit.epub --validate
gen-docs
Install (pip)
python -m pip install .
Editable install for development:
python -m pip install -e .[dev]
CLI
Run the normalizer with the installed command or the dev wrapper:
simplebook path/to/book.epub --preview --out /tmp/book.json
python -m simplebook path/to/book.epub --validate
scripts/simplebook.py path/to/book.epub
Flags:
--previewoutputs chapter structure/chunks only (omits paragraphs).--outwrites JSON to a file instead of stdout.--validatevalidates output against the JSON schema.
Aliases and completions are loaded by source ./activate (see devtools/aliases.sh).
Generated docs
Docs are generated into docs/generated. Regenerate anytime with:
gen-docs
Purpose
Provide a deterministic, no‑LLM ingest + structuring pipeline that turns an EPUB into canonical book structure (sections/chapters/chunks) plus normalized artifacts.
Owns
- Container resolution + OPF parsing (EPUB2/3)
- TOC extraction (NCX + nav)
- Spine resolution + reading order
- Canonical text extraction (HTML/XHTML → text)
- Section/Chapter inference (deterministic)
- Chunking (paragraph/segment splitting)
- Asset index + manifest normalization
- Raw artifact capture (OPF/NCX/nav, source file map)
Inputs
- EPUB file (zip)
Outputs
- Normalized manifest JSON (stable schema)
- Structured content: sections → chapters → chunks
- Raw artifact capture (OPF, NCX, nav)
- Asset map (images/css/fonts) + cover reference
Plan (deterministic, no LLMs)
-
Resolve package
- Read
META-INF/container.xml→ locate OPF. - Parse OPF 2.0/3.0 metadata, manifest, spine.
- Read
-
Build reading order
- Resolve spine items to hrefs; verify existence.
- Collect candidate content files (.html/.xhtml).
-
Extract navigation
- Prefer EPUB3 nav (
toc.xhtml), fallback to NCX. - Normalize to a single TOC list with labels + hrefs.
- Prefer EPUB3 nav (
-
Extract text deterministically
- Parse XHTML/HTML.
- Strip non-content elements (nav, scripts, styles, hidden).
- Normalize whitespace; preserve basic structure (headings, paragraphs).
-
Structure + chunk
- Deterministically segment into sections (toc entries + headings).
- Derive chapters from toc labels/heading levels when present.
- Chunk text into stable sizes with paragraph boundaries.
-
Normalize artifacts
- Canonical manifest with stable paths and metadata.
- Record raw artifacts for debugging.
- Build asset map (images/css/fonts), cover reference.
Non-goals
- No LLM usage (observations/extractions handled elsewhere).
- No styling or formatting preservation beyond basic structure.
- No “pretty” output for external ebook standards.
Data model plan (minimal, public repo)
Goal: make this a standalone, general-purpose Python module with a small, stable output surface. It is not a server or service — just classes and pure functions that can be imported by FantasyWorldBible (or any other app).
Minimal output model (serialization only)
NormalizedBook
metadata(title, language, identifiers)chapters[](ordered, real chapters only)toc[](mirrors chapters; chapter-only TOC)chunks[](paragraph chunks, numbered per chapter)artifacts(raw OPF/NCX/nav + container for debugging)
Public surface (ebooklib-backed)
EbookNormalizer/SimpleBook- Uses
ebooklibto read EPUBs and produce a serialized dict.
- Uses
serialize(book: SimpleBook) -> dict
Serialization schema (minimal)
{
"metadata": {
"title": "The Hobbit",
"language": "en",
"identifiers": ["urn:uuid:..."]
},
"manifest": [
{ "href": "text/part0001.html", "media_type": "application/xhtml+xml", "role": "content" }
],
"spine": ["text/part0001.html", "text/part0002.html"],
"chapters": [
{ "id": "ch-001", "label": "Chapter I", "href": "text/part0006.html", "text": "..." }
],
"toc": [
{ "label": "Chapter I", "href": "text/part0006.html", "chapter_id": "ch-001" }
],
"chunks": [
{ "id": "ch-001:0001", "chapter_id": "ch-001", "text": "...", "ordinal": 1 }
],
"assets": [
{ "href": "images/cover.jpeg", "media_type": "image/jpeg", "is_cover": true }
],
"artifacts": {
"container": "META-INF/container.xml",
"opf": "content.opf",
"toc_ncx": "toc.ncx",
"toc_nav": "toc.xhtml"
}
}
Simplifications vs current flow
- Drop ImportRun (tracking belongs to the host app).
- Drop Chapter as a first-class entity; represent it via TOC labels + section grouping.
- Drop validation status fields; leave validation to the host app.
- Keep only sections + chunks as canonical text units.
- No server/API layer — only Python classes + optional JSON export schema.
Deterministic derivation rules
- Sections are ordered by spine; label from TOC/nav or first heading.
- Chunks are generated from section text using deterministic sizing rules.
- No LLM usage; all rules are pure functions on EPUB contents.
Notes
- Must support OPF 2.0 + 3.0 and NCX/nav TOC.
- Never assume OPF at archive root.
Chapter detection (deterministic, no LLM)
We want real chapters, not every “section” in the EPUB.
Inputs used
- TOC/nav entries (preferred)
- Spine order (fallback + ordering)
- Heading text from content files (fallback)
Heuristics (apply in order)
- Accept TOC entries that look like chapters
- Labels that match:
Chapter,Ch.,Book,Part, Roman numerals, or numeric ordinals.
- Labels that match:
- Exclude front/back matter
- Labels like:
Titlepage,Cover,Copyright,Imprint,Dedication,Preface,Acknowledgments,Notes,Endnotes,Colophon,About the author,TOC,Contents.
- Labels like:
- Fallback to headings in spine
- If TOC is missing or unhelpful, treat first heading in each spine item as candidate.
- Normalize labels
- Standardize to
Chapter {N}when the label is numeric or roman.
- Standardize to
Outputs
chapters[]contains only accepted chapter candidates.toc[]mirrorschapters[]exactly.
Chunking
- Chunk paragraphs in chapter order.
- Each chunk has a chapter-local ordinal (
1..Nper chapter). - Target chunk size is deterministic and configurable (e.g., by char count).
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file simplebook-0.1.0.tar.gz.
File metadata
- Download URL: simplebook-0.1.0.tar.gz
- Upload date:
- Size: 17.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4831a1f7cfb805b601eafe7aac1cf2526c16b760340fdd724090c72f1fa2b3e0
|
|
| MD5 |
3e0b314a2eff9d340c6dbdb2db0eb3c3
|
|
| BLAKE2b-256 |
22c1e1e9c8269c67dbc5c2c3b170299e0eeea89bbf89d998866f3d1222091921
|
File details
Details for the file simplebook-0.1.0-py3-none-any.whl.
File metadata
- Download URL: simplebook-0.1.0-py3-none-any.whl
- Upload date:
- Size: 15.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1ee7fb6c6cdf1566b05b2fd36d363c4cfc3c5b90e80129e61872784736714c1e
|
|
| MD5 |
6e9030b24440f38e9420a7643ef7ac50
|
|
| BLAKE2b-256 |
94bdf725ca26ff0c254a1838d1a7845fc215d95b3e5b14334b77d957e740dd71
|