Skip to main content

Document Outline Discovery.

Project description

DoD (Document Outline Discovery)

This project is intended to make local deployment and usage of PageIndex easier. It is not an official PageIndex service. Check their official PageIndex repo here.

DoD is a local-first document structure extraction toolkit built around PageIndex.

What it does:

  • ingests PDFs and builds normalized page artifacts
  • extracts page-level text (pymupdf or pytesseract)
  • generates hierarchical TOC/document outline trees
  • exposes both CLI and server+MCP interfaces for agent workflows

Use cases:

  • document-grounded Q&A assistants over private document libraries
  • TOC/outline extraction for scanned or OCR-heavy PDFs
  • page-targeted retrieval pipelines for downstream RAG/agent systems

Advantages:

  • local deployment for sensitive documents
  • structured outputs (TOC tree + page tables + page images)
  • agent-friendly retrieval via stable server and MCP tools

Why this can make more sense than traditional RAG for long manuals:

  • instead of flat chunk retrieval, PageIndex builds explicit document structure (sections/subsections + page mapping)
  • this improves navigation, targeted retrieval, and answer grounding for long technical documents
  • see PageIndex's technical-manual discussion: https://pageindex.ai/blog/technical-manuals

Example (Codex using DoD MCP for section-grounded Q&A):

Codex + DoD MCP example


Table of Contents

DoD turns a document into:

  1. page_table.jsonl (page text + metadata)
  2. toc_tree.json (hierarchical Table-of-Contents tree)
  3. image_page_table.jsonl (page-image index, including image paths and image payload fields)
  4. images/ (all pages rendered as image files)

Core flow:

  1. PDF normalization to page images
  2. Text extraction (pymupdf / pytesseract)
  3. TOC generation (PageIndex)
  4. Artifact writing (JSON/JSONL + manifest)

Project Structure

  • src/DoD/cli/ - CLI entrypoint
  • src/DoD/pipeline.py - end-to-end pipeline orchestration
  • src/DoD/normalize/ - PDF/image normalization to per-page images
  • src/DoD/text_extractor/ - text extraction backends
  • src/DoD/page_table.py - page table data model + writer
  • src/DoD/pageindex/ - PageIndex TOC builder
  • src/DoD/toc/ - TOC adapters
  • src/DoD/server/ - FastAPI server mode
  • src/scripts/ - executable entrypoints (dod, dod-server, dod-mcp)
  • .agents/skills/ - example agent skills
  • makefile - install/check/test developer workflow
  • src/DoD/conf/config.yaml - default configuration

0. LLM API Configuration Required

Before running either the CLI package mode or the server mode, set an OpenAI-compatible endpoint and API key:

export PAGEINDEX_API_KEY="<your_api_key>"
export PAGEINDEX_BASE_URL="<your_openai_compatible_base_url>"

Example (Snowflake Cortex):

export PAGEINDEX_API_KEY="<snowflake_pat>"
export PAGEINDEX_BASE_URL="https://<account-identifier>.snowflakecomputing.com/api/v2/cortex/v1"

Then choose any model available on your configured endpoint via toc.model.

1. Install and Dev

Use these commands for setup and day-to-day development:

make install  # bootstrap toolchain + Python deps + hooks
make check    # lint + format + type-check
make test     # run test suite

1.1 Use as a PyPI package

Install from PyPI:

pip install "dod-outline-discovery[pageindex,text_extractor,pdf,server,mcp]"

Then run:

dod --help
dod-server --help
dod-mcp --help

One-off execution without a persistent install:

uvx --from "dod-outline-discovery[pageindex,text_extractor,pdf,server,mcp]" dod --help
uvx --from "dod-outline-discovery[pageindex,text_extractor,pdf,server,mcp]" dod-server --help
uvx --from "dod-outline-discovery[pageindex,text_extractor,pdf,server,mcp]" dod-mcp --help

2. Run As A Package CLI

Run one document:

Choose extractor first:

  • use text_extractor.backend=pymupdf for PDFs with built-in text layers
  • use text_extractor.backend=pytesseract for image-based/scanned PDFs
dod \
  input_path=/path/to/document.pdf \
  text_extractor.backend=pymupdf \
  toc.backend=pageindex \
  toc.model=claude-sonnet-4-5

If you are running from source repo instead of a PyPI install, use uv run dod.

Where output is written:

  • Hydra run dir: outputs/<YYYY-MM-DD>/<HH-MM-SS>/
  • Artifacts folder: outputs/<YYYY-MM-DD>/<HH-MM-SS>/artifacts/

Main artifact files:

  • page_table.jsonl
  • image_page_table.jsonl
  • toc_tree.json
  • manifest.json

3. Run As A Server

3.1 Start server

export DOD_SERVER_HOST=0.0.0.0
export DOD_SERVER_PORT=8000
export DOD_SERVER_MAX_CONCURRENT_DOCS=4
export DOD_SERVER_JOB_TIMEOUT_SECONDS=300
export DOD_SERVER_WORK_DIR=outputs/server_jobs
dod-server

If you are running from source repo instead of a PyPI install, use uv run dod-server.

3.2 Health check

curl http://localhost:8000/healthz

3.3 Make requests

3.3.1 Single PDF wait for final result

curl -s -X POST "http://localhost:8000/v1/digest" \
  -F "file=@/path/to/document.pdf" \
  -F "text_extractor_backend=pymupdf" \
  -F "toc_backend=pageindex" \
  -F "toc_model=claude-sonnet-4-5" \
  -F "toc_concurrent_requests=4" \
  > result.json

This call blocks until the job is done and writes full result JSON to result.json.

3.3.2 Async job submit then poll

Submit:

SUBMIT_JSON=$(curl -s -X POST "http://localhost:8000/v1/digest?wait=false" \
  -F "file=@/path/to/document.pdf" \
  -F "text_extractor_backend=pymupdf" \
  -F "toc_backend=pageindex" \
  -F "toc_model=claude-sonnet-4-5")
JOB_ID=$(echo "$SUBMIT_JSON" | jq -r '.job_id')
JOB_REF=$(echo "$SUBMIT_JSON" | jq -r '.job_ref')

Check status:

curl -s "http://localhost:8000/v1/jobs/$JOB_REF"

Get final result:

curl -s "http://localhost:8000/v1/jobs/$JOB_REF/result" > result.json

3.3.3 Process More Than One PDF At The Same Time

Use async submit (wait=false) + parallel curl:

mkdir -p jobs
printf "%s\n" \
  "/path/to/a.pdf" \
  "/path/to/b.pdf" \
  "/path/to/c.pdf" \
| xargs -I{} -P 3 sh -c '
  name=$(basename "{}" .pdf)
  curl -s -X POST "http://localhost:8000/v1/digest?wait=false" \
    -F "file=@{}" \
    -F "text_extractor_backend=pymupdf" \
    -F "toc_backend=pageindex" \
    -F "toc_model=claude-sonnet-4-5" \
    -F "toc_concurrent_requests=4" \
  > "jobs/${name}.submit.json"
'

Poll and download all results:

for f in jobs/*.submit.json; do
  job_id=$(jq -r '.job_id' "$f")
  name=$(basename "$f" .submit.json)
  until curl -sf "http://localhost:8000/v1/jobs/$job_id/result" > "jobs/${name}.result.json"; do
    sleep 2
  done
done

Notes:

  • -P 3 controls how many submit requests run in parallel.
  • Server-side processing concurrency is capped by DOD_SERVER_MAX_CONCURRENT_DOCS.
  • TOC runs in strict mode: if PageIndex fails, the job status becomes failed (no fallback TOC).

4. Output What And Where

4.1 Output in API response JSON

/v1/digest (sync) or /v1/jobs/{job_id}/result returns:

  • result.toc_tree - TOC tree JSON
  • result.page_table - page records (JSON array parsed from JSONL)
  • result.image_page_table - image records (JSON array parsed from JSONL)
  • result.artifact_paths - filesystem paths for written artifacts
  • result.manifest - run metadata + config

Convert arrays back to JSONL if needed:

jq -c '.result.page_table[]' result.json > page_table.jsonl
jq -c '.result.image_page_table[]' result.json > image_page_table.jsonl
jq '.result.toc_tree' result.json > toc_tree.json

4.2 Output on disk Server mode

For each job:

  • Job folder: ${DOD_SERVER_WORK_DIR}/<job_id>/
  • Input copy: ${DOD_SERVER_WORK_DIR}/<job_id>/input.pdf
  • Artifact folder: ${DOD_SERVER_WORK_DIR}/<job_id>/artifacts/

Inside artifacts/:

  • page_table.jsonl
  • image_page_table.jsonl
  • toc_tree.json
  • manifest.json

Server-level job index:

  • ${DOD_SERVER_WORK_DIR}/jobs.json
    • persists job metadata across server restarts
    • used by GET /v1/jobs and MCP list_jobs()

Timeout:

  • default from src/DoD/conf/config.yamlserver.job_timeout_seconds
  • runtime override via DOD_SERVER_JOB_TIMEOUT_SECONDS

4.3 Configure output paths PyPI package

When installed from PyPI, outputs are still filesystem-based and default to the current working directory.

  • CLI (dod): defaults to Hydra output under outputs/<YYYY-MM-DD>/<HH-MM-SS>/artifacts
  • Server (dod-server): defaults to outputs/server_jobs

Use absolute paths in production/local deployments:

export DOD_SERVER_WORK_DIR="/absolute/path/to/server_jobs"
export DOD_LLM_CACHE_DIR="/absolute/path/to/llm_cache"

Optional CLI override for one run:

dod input_path=/path/to/document.pdf artifacts.output_dir=/absolute/path/to/artifacts

5. Request Fields Server /v1/digest

Multipart form fields:

  • file (required, .pdf)
  • text_extractor_backend (optional)
    • use pymupdf for PDFs with built-in text layers
    • use pytesseract for image-based/scanned PDFs
  • normalize_max_pages (optional int)
  • toc_backend (optional, typically pageindex)
  • toc_model (optional model name)
  • toc_concurrent_requests (optional int)
  • toc_check_page_num (optional int)
  • toc_api_key (optional per-request override)
  • toc_api_base_url (optional per-request override)

Query parameter:

  • wait (default true)
    • true: request returns when job finishes
    • false: request returns immediately with job metadata (job_id, job_ref, status/result URLs)

6. Simple MCP Setup

If you want agents to call DoD as tools, use the included MCP wrapper at src/scripts/dod_mcp.py. Agents do not upload PDFs. Human users submit jobs first via /v1/digest, then agents use job_ref to retrieve targeted outputs.

make install already installs all extras including MCP/server dependencies.

6.1 Install from PyPI for MCP use

uv tool install "dod-outline-discovery[pageindex,text_extractor,pdf,server,mcp]"

This installs dod, dod-server, and dod-mcp as system tools.

6.2 Start DoD HTTP server

dod-server

If you are running from source repo instead of PyPI install, use uv run dod-server.

6.3 Configure MCP client recommended

Most MCP hosts should launch dod-mcp themselves via command config, for example (~/.codex/config.toml):

[mcp_servers.dod_mcp]
command = "dod-mcp"

Alternative without tool install:

[mcp_servers.dod_mcp]
command = "uvx"
args = ["--from", "dod-outline-discovery[pageindex,text_extractor,pdf,server,mcp]", "dod-mcp"]

In this mode, do not start dod-mcp manually. Keep only dod-server running.

6.4 Optional manual MCP run debug only

dod-mcp

If you are running from source repo instead of a PyPI install, use uv run dod-mcp. Use this only for debugging MCP transport behavior.

6.5 Available MCP tools

  • list_jobs()
    • Returns: { jobs: [{ file_name, job_id, job_ref, status, created_at }] }.
  • get_toc(job_ref)
    • Returns: { job_id, job_ref, status, toc_tree }.
  • get_page_texts(job_ref, pages)
    • Returns: { job_id, job_ref, status, requested_pages, returned_pages, pages } where pages is a selected subset of { page_id, text }.
  • get_page_images(job_ref, pages, mode)
    • Returns: { job_id, job_ref, status, mode, requested_pages, returned_pages, pages } where pages is a selected subset of { page_id, image_path } (mode=path) or { page_id, image_b64 } (mode=base64).
  • pages accepts flexible specs like "110,111,89-100".
  • Retrieval guardrails come from src/DoD/conf/config.yaml under retrieval:
    • max_chars_per_page (null means full page text)
    • max_pages_per_call

6.6 Example agent skill

An example Codex skill for DoD library Q&A is included at:

  • .agents/skills/dod-library-qa/SKILL.md

7. Third-Party Licensing

This project includes third-party license attributions in:

  • THIRD_PARTY_NOTICES.md

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dod_outline_discovery-0.1.2.tar.gz (463.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dod_outline_discovery-0.1.2-py3-none-any.whl (53.9 kB view details)

Uploaded Python 3

File details

Details for the file dod_outline_discovery-0.1.2.tar.gz.

File metadata

  • Download URL: dod_outline_discovery-0.1.2.tar.gz
  • Upload date:
  • Size: 463.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.11

File hashes

Hashes for dod_outline_discovery-0.1.2.tar.gz
Algorithm Hash digest
SHA256 c1fd59eb32d2060e53826ccdda550994297a26ef604a78e41b3849564c80291e
MD5 fa7c130f254fe134b0c9e3cd820fe803
BLAKE2b-256 4624ca1cdff4b114973772f6e76421dc07a3d83588fae8fd80bf50c21b369684

See more details on using hashes here.

File details

Details for the file dod_outline_discovery-0.1.2-py3-none-any.whl.

File metadata

File hashes

Hashes for dod_outline_discovery-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 696c081633feea40277d12e44c9b9f0c686215492967f12faa57b653586b74b6
MD5 e17c8640e2c7b005da56515f469cbdeb
BLAKE2b-256 3f2b365bcbec3b151715cf0cd06a1a86a8f99a897c3d18e00bf292cb468acde3

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page