📚 DABOOK
Compile any book-like PDF into a traceable, validated, editable Book Graph, then export it as ML-ready datasets (
raw/pretrain/sft/rag/code) — resumable, multi-book, parallel, with a live localhost control room.
Version
Current release: v0.1.0 (Pre-Alpha)
This is early-stage software. The core pipeline (S00–S13) and worker/supervisor runtime are functional. Some features listed in the CLI help (e.g. cache gc, --watch, --autotune) are planned but not yet implemented.
See CHANGELOG.md for release history.
⚡ Quick Start
Prerequisites
- Python 3.11 or newer
- SQLite 3.35 or newer (ships with Python ≥ 3.12; check with
python -c "import sqlite3; print(sqlite3.sqlite_version)") - ~5 GB free disk space per book
Install
# Recommended: install with uv
uv tool install dabook
# Or with pip
pip install dabook
Run
# Process a single book — starts workers, launches dashboard
dabook run my_book.pdf
# Process a whole directory
dabook run ./books/
# Resume an interrupted run with zero re-work
dabook resume
# Check that your environment is ready
dabook doctor
The live control room dashboard opens automatically at http://127.0.0.1:8765.
🌟 Key Features
- Guaranteed Resumable —
kill -9or power loss at any instant;dabook resumepicks up with zero re-work thanks to two-phase atomic filesystem commits and content-addressed caching. - Zero-Redo Caching — Same PDF + same config = instant skip. The cache key is
sha256(pdf) + sha256(stage_params). - Multi-Book Parallelism — Process batches of books concurrently. Worker counts, book concurrency, and memory ceilings are all live-adjustable from the dashboard without restart.
- Live Localhost Control Room — Real-time SSE dashboard at
http://127.0.0.1:8765showing task throughput, worker pool status, per-book progress, error logs, and ETA. - Full Provenance Chain — Every token traces back: PDF → Page → BBox → Block → Node → Export Record.
- Multiple ML Export Formats — Produces five JSONL files per book (see Export Formats below).
- Human-Editable — The Book Graph is plain JSON; you can inspect and patch any node and re-compile.
🏗️ Pipeline Stages (S00–S13)
PDF → S00 → S01 → S02 (×shards) → S03 → S04 → S05 → S06
↓
datasets ← S13 ← S12 ← S11 ← S10 ← S09 ← S08 ← S07
| Stage | Name | Pool | What it does |
|---|---|---|---|
| S00 | s00_register |
io |
Hash the PDF (SHA-256), create workspace dirs, write source metadata |
| S01 | s01_inspect |
cpu |
Extract page count, detect born-digital vs scanned, parse PDF bookmarks (TOC) |
| S02 | s02_extract |
cpu |
Sharded parallel extraction — text blocks, bounding boxes, char counts per page |
| S03 | s03_merge |
cpu |
Merge all extraction shards into a single sorted page stream |
| S04 | s04_furniture |
cpu |
Detect and mark running headers, footers, page numbers, and watermarks |
| S05 | s05_reading_order |
cpu |
XY-cut column detection; reconstruct true top-to-bottom, left-to-right reading order |
| S06 | s06_structure |
cpu |
Reconcile PDF TOC bookmarks with font-weight heuristics to label headings and chapters |
| S07 | s07_continuity |
cpu |
Stitch paragraphs broken across page boundaries; flag cross-page continuations |
| S08 | s08_typed_content |
cpu |
Classify blocks as code, list_item, or paragraph; detect x86/x64 assembly and C |
| S09 | s09_graph |
cpu |
Build the canonical Book Graph — a typed node tree (book → chapter → section → block) |
| S10 | s10_clean |
cpu |
Unicode NFC normalization, ligature healing (fi nd → find), dehyphenation |
| S11 | s11_validate |
cpu |
Compute quality score; check word count, heading count, and code block count |
| S12 | s12_semantic |
cpu |
Add breadcrumb paths to every node (e.g. ["Chapter 3", "Memory Layout"]) |
| S13 | s13_compile |
io |
Export all five dataset formats (JSONL); consolidate workspace-level manifests |
All stages run through the same crash-safe commit protocol: write to .partial/ → atomic rename → write _SUCCESS. A stage output is valid if and only if _SUCCESS exists.
📦 Export Formats
S13 produces five JSONL files inside <workspace>/books/<sha12>/datasets/:
| File | Description |
|---|---|
raw.jsonl |
Every non-root node with its text, type, page index, and breadcrumbs |
pretrain.jsonl |
Document-level Markdown chapters with fenced code blocks; suitable for continued pre-training |
sft.jsonl |
Supervised fine-tuning pairs — concept Q&A, code walkthroughs, and exercise answers; Alpaca and ChatML format |
code.jsonl |
Multi-line code and assembly snippets with language tag and surrounding context |
rag.jsonl |
~400-word semantic chunks bounded by section boundaries, each with full breadcrumb path |
A manifest.json and <workspace>/datasets/manifest.json (consolidated across all books) are also written.
💻 CLI Reference
# ── Processing ──────────────────────────────────────────────────────
dabook run <path> [<path>...] # Process PDFs or directories
--workspace / -w <dir> # Workspace dir (default: ./dabook_workspace)
--profile fast|balanced|accurate|lowram
--workers-cpu N # Override CPU worker count
--workers-gpu N # Override GPU worker count
--book-concurrency N # Max books processed at once
--port 8765 # Dashboard port
--no-ui # Skip dashboard
--no-open # Don't auto-open browser
dabook resume [-w <dir>] # Resume all work in a workspace
# ── Monitoring ──────────────────────────────────────────────────────
dabook status [-w <dir>] [--json] # Print book/task counts by state
dabook ui [-w <dir>] [--port 8765] # Start dashboard for an existing workspace
dabook doctor [-w <dir>] # Check env, deps, SQLite version, disk space
# ── Queue Control ───────────────────────────────────────────────────
dabook pause [-w <dir>] # Pause the processing queue
dabook continue [-w <dir>] # Resume a paused queue
dabook stop [-w <dir>] # Soft-stop (finish current task then exit)
# ── Recovery ────────────────────────────────────────────────────────
dabook retry [-w <dir>] [--dead] [--failed] [--book <sha_prefix>]
# ── Cache ────────────────────────────────────────────────────────────
dabook cache stats [-w <dir>] # Count committed stage artifacts
dabook cache verify [-w <dir>] # Deep-verify SHA-256 of all artifacts
dabook cache gc # (planned, not yet implemented)
⚙️ Worker Pools
Workers are spread across four resource classes. The supervisor spawns and drains them automatically; all counts are live-adjustable in the dashboard.
| Pool | Default | Used for |
|---|---|---|
cpu |
4 | Most pipeline stages (S01–S12) |
io |
2 | Register (S00), merge (S03), compile (S13) |
gpu |
1 | Reserved for future GPU-accelerated backends |
llm |
0 | Reserved for future LLM-assisted stages |
The supervisor uses a crash loop breaker: if a worker lane crashes 3+ times within 5 minutes, spawning is paused for 30 seconds before retrying.
🎯 Design Principles
- The block is the unit of meaning, never the page. Page boundaries are printing artifacts; semantic blocks transcend them.
- Deterministic first, AI second. Regex, font heuristics, and bbox geometry run before any neural model.
- Strict source-faithfulness. Raw text is never discarded; all generated/inferred content is flagged separately.
- Three text representations.
raw(exact PDF glyphs) →normalized(ligatures, hyphens fixed) →clean(furniture stripped, Unicode NFC). - Infrastructure before intelligence. WAL-mode SQLite, process leases, atomic commits, and crash recovery come before ML pipelines.
🗄️ State Store
All state lives in <workspace>/state.db (SQLite, WAL mode). Key tables:
| Table | Purpose |
|---|---|
books |
One row per PDF; tracks state (queued → active → done/failed) |
tasks |
One row per stage+shard; owns the lease, heartbeat, and output path |
workers |
Live worker registry with PID, RSS, and current task |
settings |
Live-editable key-value config polled every 2 s by workers |
events |
Append-only event log (shown in dashboard log tail) |
samples |
Downsampled telemetry (CPU/RAM/disk/GPU) for history charts |
🤝 Contributing
git clone https://github.com/brovk2008/Dabook.git
cd Dabook
uv venv --python 3.11
source .venv/bin/activate # Windows: .\.venv\Scripts\Activate.ps1
uv pip install -e ".[dev]"
pytest # run the test suite
pytest -m "not slow" # skip crash-safety tests
See CONTRIBUTING.md, CODE_OF_CONDUCT.md, SECURITY.md, and SUPPORT.md.
⚖️ License
Distributed under the Apache License 2.0. See LICENSE for details.
Optional third-party plugins or proprietary model weights have separate licenses.
Metadata
Release files for dabook 0.2.26
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dabook-0.2.26.tar.gz | 101.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dabook-0.2.26-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 198.0 kB
Release files / dabook-0.2.26.tar.gz
| Download URL | dabook-0.2.26.tar.gz |
|---|---|
| Size | 101.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
49039f07cda697240192abbd48cfbc2810ffeb62f2ce62431525886f3d1ac879
|
|
BLAKE2b-256 checksum How to use checksums |
3ae2a9cad708045bbecc3b878ce42cd4d2691acc9a5c777b2d9485f46c828339
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / dabook-0.2.26-py3-none-any.whl
| Download URL | dabook-0.2.26-py3-none-any.whl |
|---|---|
| Size | 96.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
37a6643773eb96bc4399f43930f46b1ad66815e3899a1b3c0dc377b56b81b925
|
|
BLAKE2b-256 checksum How to use checksums |
15fb84bcf2bba058347792ae1d99bb0aef5bf596b98f5d512af4a4afeb8e6846
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log