Skip to main content

headcleaner

Walk a folder, convert every document to Markdown (with frontmatter), OKF v0.2 (with frontmatter), or both — with an omp-style animated TUI.

headcleaner convert ~/Documents/inbox --format both --output ~/Documents/inbox.clean

headcleaner is a Python CLI that scans a directory you provide, identifies each document by extension, runs the appropriate extraction engine (OfficeCLI for Office formats, pdfplumber for PDFs, BeautifulSoup for HTML, etc.), and emits clean normalized output — either side-by-side Markdown and OKF, or just one.

  • Output formats: --format md (Markdown), --format okf (OKF v0.2 bundle), --format both (default)
  • Engine coverage: 7 formats out of the box (XLSX, DOCX, PPTX, PDF, HTML, HTM, TXT) — see docs/FORMAT_MATRIX.md for the 16-format v1.0 roadmap
  • TUI: omp-inspired animated terminal (box-drawing panels, neon palette, powerline separators)
  • Linter: headcleaner lint reviews the converted Markdown / OKF for formatting issues
  • Localization: gettext catalogs for English, Spanish (--lang es), and Simplified Chinese (--lang zh-CN) across CLI/TUI runtime status
  • Legacy Office: .doc, .xls, and .ppt convert through LibreOffice headless, then follow the normal Office extraction pipeline
  • PST: per-message concepts with full bodies and attachments through readpst, including MSYS2-aware Windows discovery
  • Heuristic cleanup: headcleaner convert --clean runs a 12-stage any2md-inspired cleanup pipeline
  • all2md fallback: Auto-handles 38 extra formats (Jupyter, LaTeX, reST, sourcecode, etc.) when all2md is installed
  • headcleaner mcp: Run headcleaner as an MCP server exposing 14 okf_* tools to any MCP agent host (Claude Code, Cursor, etc.) — install with uv pip install "headcleaner[mcp]"
  • Diagnostics: headcleaner doctor checks Python, PATH, OfficeCLI, output permissions, and the @slug registry, then prints a GO/NO-GO verdict
  • Adapter plugins: Third-party packages register formats through the headcleaner_plugin entry-point group
  • zsv CSV: World's-fastest SIMD CSV parser (~10-100x stdlib) when zsv is on PATH
  • Trust attestation: headcleaner attest builds a Merkle root + ed25519 signature; verify checks it
  • Local browse: headcleaner serve <bundle> exposes a FastAPI UI for browsing + search
  • Honest defaults: OKF trust fields filled with unverified / human:pending, never invented

Install

# 1. The Office engine — single binary, no Office install needed
npm install -g @officecli/officecli

# 2. The CLI itself (Python ≥3.12, uv-managed)
uv tool install headcleaner

# Or for development:
git clone <this repo>
cd headcleaner-cli
uv sync
uv run headcleaner --help

For other install methods (curl | bash, pip, brew, Windows PowerShell), see docs/INSTALL.md.

Quick start

headcleaner ~/Documents/inbox --format both --output ./clean

This produces:

clean/
├── manifest.json                  # run summary: per-file status, engine, sha256
├── REPORT.md                      # count, average time, and error rate by engine
├── _md/                           # plain Markdown (one file per source)
│   ├── notes.docx.md
│   ├── q3.pdf.md
│   └── ...
└── okf/                           # OKF v0.2 bundle (one concept per source)
    ├── index.md                   # auto-generated directory index
    ├── notes.md                   # OKF concept: type=Document
    ├── q3.pdf.md
    └── ...

CLI reference

headcleaner convert <INPUT_DIR> [OPTIONS]

Options:
  -f, --format {md,okf,both}   Output format(s) [default: both]
  -o, --output DIR             Output directory [default: ./out]
  --ocr                        Enable Tesseract OCR for scanned PDFs
  --officecli-timeout <secs>   Timeout per OfficeCLI subprocess call (default: 60)
  --include, -i GLOB           Include glob (may be repeated)
  --exclude, -e GLOB           Exclude glob (may be repeated)
  --jobs, -j N                Parallel worker processes (default: 1 = sequential)
  --no-cache                  Re-convert every file (skip the SHA-256 cache)
  --no-continue-on-error       Stop on the first failure
  --obsidian-compat            Add Obsidian-friendly flat fields to OKF frontmatter
  --clean                       Run the 12-stage heuristic cleanup pipeline (any2md-inspired) on each body
  --tui / --no-tui             Force / disable the animated TUI (default: auto-detect TTY)
  --no-okf-index               Skip OKF directory index.md generation

Other commands: headcleaner doctor [--output-dir DIR] Run install and permission diagnostics headcleaner templates List supported formats headcleaner agents Show engine install status headcleaner watch IN [--webhook-url URL] Re-convert on file changes (Ctrl+C to stop) headcleaner lint

Review converted Markdown / OKF for formatting issues headcleaner lint --fix Auto-repair safe issues to .fixed/ headcleaner serve Local HTTP browser for the OKF bundle headcleaner notion-import <EXPORT.zip> Reverse a Notion workspace export headcleaner attest Compute Merkle root + optional ed25519 signature headcleaner verify Verify an attestation against the bundle


## Languages

HeadCleaner uses standard-library gettext catalogs at runtime. English is the
fallback; Spanish and Simplified Chinese are available for conversion and TUI
runtime status:

```bash
headcleaner --lang es convert ./inbox --output ./out --no-tui
headcleaner --lang zh-CN convert ./inbox --output ./out

# Equivalent process default (the CLI option takes precedence)
HEADCLEANER_LANG=es headcleaner convert ./inbox --output ./out

Catalog sources live under src/headcleaner/locales/; they are compiled with Babel during development and shipped in the wheel. Add new source strings to the .po catalogs, compile them with uv run pybabel compile -d src/headcleaner/locales -D headcleaner, and test each supported locale.

Why OKF?

OKF (Open Knowledge Format, v0.2) is just markdown + YAML frontmatter in a directory hierarchy. That means:

  • Every concept is a single .md file you can cat, grep, edit in any text editor
  • Bundles live in git — pull requests, diffs, blame all work
  • Obsidian, Notion, MkDocs, Hugo, Jekyll all consume OKF natively
  • Required frontmatter key is just type — anything beyond that is producer freedom

See docs/OKF_NOTES.md for the OKF v0.2 specifics this CLI emits.

Trust stance (honest defaults)

We never auto-claim review. Every emitted OKF concept gets:

  • status: unverified
  • verified: human:pending
  • generated: human:<user>@<host> (OKF §7 actor convention)
  • stale_after: <today + 180d>
  • sources: [{uri: file://..., sha256: ...}]

A human can grep human:pending later to find concepts needing review. See docs/OKF_NOTES.md for the full contract.

Supported formats

See docs/FORMAT_MATRIX.md for the full engine × library table. At a glance:

Format Engine Library
.docx, .xlsx, .pptx OfficeCLI binary (native DOM)
.pdf pdfplumber (text-layer), pytesseract if --ocr pdfplumber / pytesseract
.html, .htm BeautifulSoup beautifulsoup4
.txt chardet + read chardet
.md, .markdown pass-through + frontmatter inject stdlib
.csv, .tsv Sniffer dialect + GFM table (zsv SIMD when installed) stdlib csv (or zsv binary)
.json pretty-print + fenced block stdlib json
.eml headers + text/html body + attachments stdlib email
.epub per-chapter HTML → MD ebooklib (+ bs4 fallback)
.rtf control-word stripping striprtf (+ regex fallback)
.odt, .ods, .odp paragraph/row extraction + GFM tables odfpy (+ raw-XML fallback)
.msg Outlook headers + body + attachments extract-msg
.pst per-message (one OKF concept per email) readpst (libpst) + libpff-python fallback
.docx, .xlsx, .pptx office_oxide (primary, ~100x faster), OfficeCLI binary (fallback) office_oxide 0.1.8 (PyO3)
.ipynb, .latex, .rst, sourcecode, .enex, .chm, etc. (38 formats) all2md (when installed) all2md 1.12
.doc, .xls, .ppt LibreOffice headless → office_oxide / OfficeCLI LibreOffice + modern Office engine

Live mode

headcleaner watch ~/inbox --output ~/out --webhook-url https://hooks.slack.com/...

Re-runs the conversion automatically when files change under ~/inbox. Each re-run POSTs the manifest to the webhook URL (optional). Press Ctrl+C to stop.

Obsidian vault sync

headcleaner convert ~/inbox --format okf \
    --output ~/Documents/MyVault/Concepts \
    --obsidian-compat

Adds Obsidian-friendly flat fields (source, sha256, generated_by, verified_by, stale_on) to the OKF frontmatter so the concept shows up correctly in Obsidian's property panel. Original OKF fields stay intact for round-tripping.

Review (human sign-off)

Auto-conversion sets verified: human:pending. The headcleaner review TUI walks every pending concept in a bundle and lets a human flip each to:

  • approved → verified: human:reviewed, status: verified, reviewed_at, reviewed_by, reviewed_via
  • rejected → verified: human:rejected, status: rejected, optional rejection_reasons[]
  • skipped → leaves the concept as pending
headcleaner review ./out/okf
# Textual TUI: a=approve, r=reject, s=skip, n=next, p=prev, q=quit

If Textual isn't available (e.g. headless CI), a plain-mode REPL falls back automatically.

Distribution

  • PyPI: pip install headcleaner (built via uv, published via OIDC trusted publishing on tag push)
  • Homebrew: brew install headcleaner (formula in packaging/homebrew/)
  • Docker: docker pull ghcr.io/jamesdsizemore/headcleaner-cli (multi-stage image with Tesseract and readpst)
  • Windows: winget install headcleaner, scoop install headcleaner, choco install headcleaner
  • Static binary: pip install pyinstaller && pyinstaller packaging/pyinstaller/headcleaner.spec

Full release checklist in RELEASE.md.

CLI surface

headcleaner view <bundle> (add --tui to browse in the terminal) renders an OKF bundle as a single self-contained HTML graph (no backend, opens in any browser). See docs/VIEWER.md for full options.

headcleaner convert         IN_DIR [flags]    # walk + convert
headcleaner watch           IN_DIR [flags]    # live mode + webhooks
headcleaner review          BUNDLE            # human sign-off TUI/REPL
headcleaner attest          BUNDLE [--private-key PEM]   # Merkle root + optional ed25519 sig
headcleaner verify          BUNDLE [--public-key PEM]    # verify an attestation
headcleaner serve           BUNDLE [--host] [--port]    # local HTTP browser for the bundle
headcleaner glob            DIR               # interactive include REPL (Textual)
headcleaner notion-import   EXPORT.zip OUT    # reverse a Notion workspace export
headcleaner lint            DIR [--fix]       # OKF + MD rule checks
headcleaner doctor          [--output-dir]    # dependency and permission preflight
headcleaner agents          [stdout]          # emit AGENTS.md
headcleaner templates                        # list supported formats

Documentation

Document Purpose
README.md this file — install, quick start, CLI reference
docs/INSTALL.md all install paths (curl, pip, brew, PowerShell, uv, Docker)
docs/USAGE.md detailed usage guide with worked examples
docs/ARCHITECTURE.md how the pipeline fits together, where to extend
docs/FORMAT_MATRIX.md every supported format × engine × library
docs/OKF_NOTES.md OKF v0.2 contract this CLI emits + trust policy
docs/SCHEMA.md OKF frontmatter JSON Schema and editor/CI integration
docs/PLUGINS.md third-party adapter entry-point protocol
docs/TROUBLESHOOTING.md common errors and fixes
docs/FAQ.md frequently asked questions
docs/CONTRIBUTING.md how to add a new format / engine / emitter
docs/CHANGELOG.md release history
docs/ENHANCEMENTS.md 44+ shipped enhancements + future ideas
vscode-extension/README.md HeadCleaner VS Code extension (Concept Explorer + Trust Inspector)

Troubleshooting

officecli not found — install with npm install -g @officecli/officecli. Run headcleaner agents to verify.

PDF with no extractable text — your PDF is image-only. Re-run with --ocr (requires pytesseract + Tesseract binary on PATH).

Hidden files skipped — intentional. Files starting with . are dropped by the walker.

OKF index.md missing for root — auto-generated when the bundle has ≥1 concept. Use --no-okf-index to opt out.

More — see docs/TROUBLESHOOTING.md.

Development

git clone <this repo>
cd headcleaner-cli
uv sync
uv run pytest                # 314 tests, ~14s
uv run headcleaner convert ./tests/fixtures --format both --output ./out

Architecture

src/headcleaner/
├── walk.py         # recursive folder walker
├── router.py       # extension → engine dispatch
├── normalize.py    # CanonicalDoc + OKF/MD frontmatter builders
├── lint.py         # post-conversion linter (OKF + Markdown)
├── run.py          # pipeline orchestrator
├── cli.py          # Click CLI (headcleaner command)
├── tui.py          # Textual TUI (omp-style)
├── engines/
│   ├── base.py     # Adapter ABC
│   ├── officecli.py
│   ├── pdf.py
│   ├── html.py
│   └── txt.py
└── emit/
    ├── markdown.py
    ├── okf.py
    ├── okf_index.py
    └── manifest.py

Adding a new format: drop a module in engines/, register the adapter in router.py, add a row to docs/FORMAT_MATRIX.md. See docs/CONTRIBUTING.md for the full extension guide.

License

Apache-2.0

Metadata

Release files for headcleaner 0.14.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for headcleaner 0.14.0
File Size Uploaded
headcleaner-0.14.0.tar.gz 368.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for headcleaner 0.14.0
File Interpreter ABI Platform
headcleaner-0.14.0-py3-none-any.whl Python 3 none any Details

Total release size: 515.4 kB

Release files / headcleaner-0.14.0.tar.gz

Download URL headcleaner-0.14.0.tar.gz
Size 368.5 kB
Tags Source
SHA-256 checksum
How to use checksums
525758aa05926909482c790f18b66b555b2e98b3e4412241ac356ade709f5f8f
BLAKE2b-256 checksum
How to use checksums
7fed587fc102f5b82f020ebb0cfa0ea7069a5345b9246894214d843fcf852273
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.

Transparency log

Release files / headcleaner-0.14.0-py3-none-any.whl

Download URL headcleaner-0.14.0-py3-none-any.whl
Size 146.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
fe58002a6afb03193a50850e2241f241cccfe06ffa0760d534baffa62c5b7eec
BLAKE2b-256 checksum
How to use checksums
2018bfd1fc58a273ea59617f61f1c53c02f2a17719f3ac1a8a2e9e0835ae0857
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.14.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page