SimpleBook
Status: active (deterministic EPUB normalizer + tools).
Quick start
source ./activate
regtest --list
simplebook tests/epubs/the-hobbit.epub --preview -o /tmp/the-hobbit.json
simplebook tests/epubs/the-hobbit.epub --validate
gen-docs
Install (pip)
python -m pip install .
Editable install for development:
python -m pip install -e .[dev]
CLI
Run the normalizer with the installed command or the dev wrapper:
simplebook path/to/book.epub --preview --out /tmp/book.json
python -m simplebook path/to/book.epub --validate
scripts/simplebook.py path/to/book.epub
Flags:
--previewoutputs chapter structure/chunks only (omits paragraphs).--outwrites JSON to a file instead of stdout.--validatevalidates output against the JSON schema.
Aliases and completions are loaded by source ./activate (see devtools/aliases.sh).
Generated docs
Docs are generated into docs/generated. Regenerate anytime with:
gen-docs
Purpose
Provide a deterministic, no‑LLM ingest + structuring pipeline that turns an EPUB into canonical book structure (sections/chapters/chunks) plus normalized artifacts.
Owns
- Container resolution + OPF parsing (EPUB2/3)
- TOC extraction (NCX + nav)
- Spine resolution + reading order
- Canonical text extraction (HTML/XHTML → text)
- Section/Chapter inference (deterministic)
- Chunking (paragraph/segment splitting)
- Asset index + manifest normalization
- Raw artifact capture (OPF/NCX/nav, source file map)
Inputs
- EPUB file (zip)
Outputs
- Normalized manifest JSON (stable schema)
- Structured content: sections → chapters → chunks
- Raw artifact capture (OPF, NCX, nav)
- Asset map (images/css/fonts) + cover reference
Plan (deterministic, no LLMs)
-
Resolve package
- Read
META-INF/container.xml→ locate OPF. - Parse OPF 2.0/3.0 metadata, manifest, spine.
- Read
-
Build reading order
- Resolve spine items to hrefs; verify existence.
- Collect candidate content files (.html/.xhtml).
-
Extract navigation
- Prefer EPUB3 nav (
toc.xhtml), fallback to NCX. - Normalize to a single TOC list with labels + hrefs.
- Prefer EPUB3 nav (
-
Extract text deterministically
- Parse XHTML/HTML.
- Strip non-content elements (nav, scripts, styles, hidden).
- Normalize whitespace; preserve basic structure (headings, paragraphs).
-
Structure + chunk
- Deterministically segment into sections (toc entries + headings).
- Derive chapters from toc labels/heading levels when present.
- Chunk text into stable sizes with paragraph boundaries.
-
Normalize artifacts
- Canonical manifest with stable paths and metadata.
- Record raw artifacts for debugging.
- Build asset map (images/css/fonts), cover reference.
Non-goals
- No LLM usage (observations/extractions handled elsewhere).
- No styling or formatting preservation beyond basic structure.
- No “pretty” output for external ebook standards.
Data model plan (minimal, public repo)
Goal: make this a standalone, general-purpose Python module with a small, stable output surface. It is not a server or service — just classes and pure functions that can be imported by FantasyWorldBible (or any other app).
Minimal output model (serialization only)
NormalizedBook
metadata(title, language, identifiers)chapters[](ordered, real chapters only)toc[](mirrors chapters; chapter-only TOC)chunks[](paragraph chunks, numbered per chapter)artifacts(raw OPF/NCX/nav + container for debugging)
Public surface (ebooklib-backed)
EbookNormalizer/SimpleBook- Uses
ebooklibto read EPUBs and produce a serialized dict.
- Uses
serialize(book: SimpleBook) -> dict
Serialization schema (minimal)
{
"metadata": {
"title": "The Hobbit",
"language": "en",
"identifiers": ["urn:uuid:..."]
},
"manifest": [
{ "href": "text/part0001.html", "media_type": "application/xhtml+xml", "role": "content" }
],
"spine": ["text/part0001.html", "text/part0002.html"],
"chapters": [
{ "id": "ch-001", "label": "Chapter I", "href": "text/part0006.html", "text": "..." }
],
"toc": [
{ "label": "Chapter I", "href": "text/part0006.html", "chapter_id": "ch-001" }
],
"chunks": [
{ "id": "ch-001:0001", "chapter_id": "ch-001", "text": "...", "ordinal": 1 }
],
"assets": [
{ "href": "images/cover.jpeg", "media_type": "image/jpeg", "is_cover": true }
],
"artifacts": {
"container": "META-INF/container.xml",
"opf": "content.opf",
"toc_ncx": "toc.ncx",
"toc_nav": "toc.xhtml"
}
}
Simplifications vs current flow
- Drop ImportRun (tracking belongs to the host app).
- Drop Chapter as a first-class entity; represent it via TOC labels + section grouping.
- Drop validation status fields; leave validation to the host app.
- Keep only sections + chunks as canonical text units.
- No server/API layer — only Python classes + optional JSON export schema.
Deterministic derivation rules
- Sections are ordered by spine; label from TOC/nav or first heading.
- Chunks are generated from section text using deterministic sizing rules.
- No LLM usage; all rules are pure functions on EPUB contents.
Notes
- Must support OPF 2.0 + 3.0 and NCX/nav TOC.
- Never assume OPF at archive root.
Chapter detection (deterministic, no LLM)
We want real chapters, not every “section” in the EPUB.
Inputs used
- TOC/nav entries (preferred)
- Spine order (fallback + ordering)
- Heading text from content files (fallback)
Heuristics (apply in order)
- Accept TOC entries that look like chapters
- Labels that match:
Chapter,Ch.,Book,Part, Roman numerals, or numeric ordinals.
- Labels that match:
- Exclude front/back matter
- Labels like:
Titlepage,Cover,Copyright,Imprint,Dedication,Preface,Acknowledgments,Notes,Endnotes,Colophon,About the author,TOC,Contents.
- Labels like:
- Fallback to headings in spine
- If TOC is missing or unhelpful, treat first heading in each spine item as candidate.
- Normalize labels
- Standardize to
Chapter {N}when the label is numeric or roman.
- Standardize to
Outputs
chapters[]contains only accepted chapter candidates.toc[]mirrorschapters[]exactly.
Chunking
- Chunk paragraphs in chapter order.
- Each chunk has a chapter-local ordinal (
1..Nper chapter). - Target chunk size is deterministic and configurable (e.g., by char count).
Release files for simplebook 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| simplebook-0.1.0.tar.gz | 17.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| simplebook-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 33.0 kB
Release files / simplebook-0.1.0.tar.gz
| Download URL | simplebook-0.1.0.tar.gz |
|---|---|
| Size | 17.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4831a1f7cfb805b601eafe7aac1cf2526c16b760340fdd724090c72f1fa2b3e0
|
|
BLAKE2b-256 checksum How to use checksums |
22c1e1e9c8269c67dbc5c2c3b170299e0eeea89bbf89d998866f3d1222091921
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.5
|
Release files / simplebook-0.1.0-py3-none-any.whl
| Download URL | simplebook-0.1.0-py3-none-any.whl |
|---|---|
| Size | 15.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1ee7fb6c6cdf1566b05b2fd36d363c4cfc3c5b90e80129e61872784736714c1e
|
|
BLAKE2b-256 checksum How to use checksums |
94bdf725ca26ff0c254a1838d1a7845fc215d95b3e5b14334b77d957e740dd71
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.5
|