headcleaner
Walk a folder, convert every document to Markdown (with frontmatter), OKF v0.2 (with frontmatter), or both — with an omp-style animated TUI.
headcleaner convert ~/Documents/inbox --format both --output ~/Documents/inbox.clean
headcleaner is a Python CLI that scans a directory you provide, identifies each document by extension, runs the appropriate extraction engine (OfficeCLI for Office formats, pdfplumber for PDFs, BeautifulSoup for HTML, etc.), and emits clean normalized output — either side-by-side Markdown and OKF, or just one.
- Output formats:
--format md(Markdown),--format okf(OKF v0.2 bundle),--format both(default) - Engine coverage: 7 formats out of the box (XLSX, DOCX, PPTX, PDF, HTML, HTM, TXT) — see docs/FORMAT_MATRIX.md for the 16-format v1.0 roadmap
- TUI: omp-inspired animated terminal (box-drawing panels, neon palette, powerline separators)
- Linter:
headcleaner lintreviews the converted Markdown / OKF for formatting issues - Localization: gettext catalogs for English, Spanish (
--lang es), and Simplified Chinese (--lang zh-CN) across CLI/TUI runtime status - Legacy Office:
.doc,.xls, and.pptconvert through LibreOffice headless, then follow the normal Office extraction pipeline - PST: per-message concepts with full bodies and attachments through
readpst, including MSYS2-aware Windows discovery - Heuristic cleanup:
headcleaner convert --cleanruns a 12-stage any2md-inspired cleanup pipeline - all2md fallback: Auto-handles 38 extra formats (Jupyter, LaTeX, reST, sourcecode, etc.) when all2md is installed
headcleaner mcp: Run headcleaner as an MCP server exposing 14okf_*tools to any MCP agent host (Claude Code, Cursor, etc.) — install withuv pip install "headcleaner[mcp]"- Diagnostics:
headcleaner doctorchecks Python, PATH, OfficeCLI, output permissions, and the@slugregistry, then prints a GO/NO-GO verdict - Adapter plugins: Third-party packages register formats through the
headcleaner_pluginentry-point group - zsv CSV: World's-fastest SIMD CSV parser (~10-100x stdlib) when
zsvis on PATH - Trust attestation:
headcleaner attestbuilds a Merkle root + ed25519 signature;verifychecks it - Local browse:
headcleaner serve <bundle>exposes a FastAPI UI for browsing + search - Honest defaults: OKF trust fields filled with
unverified/human:pending, never invented
Install
# 1. The Office engine — single binary, no Office install needed
npm install -g @officecli/officecli
# 2. The CLI itself (Python ≥3.12, uv-managed)
uv tool install headcleaner
# Or for development:
git clone <this repo>
cd headcleaner-cli
uv sync
uv run headcleaner --help
For other install methods (curl | bash, pip, brew, Windows PowerShell), see docs/INSTALL.md.
Quick start
headcleaner ~/Documents/inbox --format both --output ./clean
This produces:
clean/
├── manifest.json # run summary: per-file status, engine, sha256
├── REPORT.md # count, average time, and error rate by engine
├── _md/ # plain Markdown (one file per source)
│ ├── notes.docx.md
│ ├── q3.pdf.md
│ └── ...
└── okf/ # OKF v0.2 bundle (one concept per source)
├── index.md # auto-generated directory index
├── notes.md # OKF concept: type=Document
├── q3.pdf.md
└── ...
CLI reference
headcleaner convert <INPUT_DIR> [OPTIONS]
Options:
-f, --format {md,okf,both} Output format(s) [default: both]
-o, --output DIR Output directory [default: ./out]
--ocr Enable Tesseract OCR for scanned PDFs
--officecli-timeout <secs> Timeout per OfficeCLI subprocess call (default: 60)
--include, -i GLOB Include glob (may be repeated)
--exclude, -e GLOB Exclude glob (may be repeated)
--jobs, -j N Parallel worker processes (default: 1 = sequential)
--no-cache Re-convert every file (skip the SHA-256 cache)
--no-continue-on-error Stop on the first failure
--obsidian-compat Add Obsidian-friendly flat fields to OKF frontmatter
--clean Run the 12-stage heuristic cleanup pipeline (any2md-inspired) on each body
--tui / --no-tui Force / disable the animated TUI (default: auto-detect TTY)
--no-okf-index Skip OKF directory index.md generation
Other commands: headcleaner doctor [--output-dir DIR] Run install and permission diagnostics headcleaner templates List supported formats headcleaner agents Show engine install status headcleaner watch IN [--webhook-url URL] Re-convert on file changes (Ctrl+C to stop) headcleaner lint
Review converted Markdown / OKF for formatting issues headcleaner lint --fix Auto-repair safe issues to .fixed/ headcleaner serve Local HTTP browser for the OKF bundle headcleaner notion-import <EXPORT.zip> Reverse a Notion workspace export headcleaner attest Compute Merkle root + optional ed25519 signature headcleaner verify Verify an attestation against the bundle
## Languages
HeadCleaner uses standard-library gettext catalogs at runtime. English is the
fallback; Spanish and Simplified Chinese are available for conversion and TUI
runtime status:
```bash
headcleaner --lang es convert ./inbox --output ./out --no-tui
headcleaner --lang zh-CN convert ./inbox --output ./out
# Equivalent process default (the CLI option takes precedence)
HEADCLEANER_LANG=es headcleaner convert ./inbox --output ./out
Catalog sources live under src/headcleaner/locales/; they are compiled with
Babel during development and shipped in the wheel. Add new source strings to
the .po catalogs, compile them with uv run pybabel compile -d src/headcleaner/locales -D headcleaner, and test each supported locale.
Why OKF?
OKF (Open Knowledge Format, v0.2) is just markdown + YAML frontmatter in a directory hierarchy. That means:
- Every concept is a single
.mdfile you cancat,grep, edit in any text editor - Bundles live in git — pull requests, diffs, blame all work
- Obsidian, Notion, MkDocs, Hugo, Jekyll all consume OKF natively
- Required frontmatter key is just
type— anything beyond that is producer freedom
See docs/OKF_NOTES.md for the OKF v0.2 specifics this CLI emits.
Trust stance (honest defaults)
We never auto-claim review. Every emitted OKF concept gets:
status: unverifiedverified: human:pendinggenerated: human:<user>@<host>(OKF §7 actor convention)stale_after: <today + 180d>sources: [{uri: file://..., sha256: ...}]
A human can grep human:pending later to find concepts needing review. See docs/OKF_NOTES.md for the full contract.
Supported formats
See docs/FORMAT_MATRIX.md for the full engine × library table. At a glance:
| Format | Engine | Library |
|---|---|---|
.docx, .xlsx, .pptx |
OfficeCLI binary | (native DOM) |
.pdf |
pdfplumber (text-layer), pytesseract if --ocr |
pdfplumber / pytesseract |
.html, .htm |
BeautifulSoup | beautifulsoup4 |
.txt |
chardet + read | chardet |
.md, .markdown |
pass-through + frontmatter inject | stdlib |
.csv, .tsv |
Sniffer dialect + GFM table (zsv SIMD when installed) | stdlib csv (or zsv binary) |
.json |
pretty-print + fenced block | stdlib json |
.eml |
headers + text/html body + attachments | stdlib email |
.epub |
per-chapter HTML → MD | ebooklib (+ bs4 fallback) |
.rtf |
control-word stripping | striprtf (+ regex fallback) |
.odt, .ods, .odp |
paragraph/row extraction + GFM tables | odfpy (+ raw-XML fallback) |
.msg |
Outlook headers + body + attachments | extract-msg |
.pst |
per-message (one OKF concept per email) | readpst (libpst) + libpff-python fallback |
.docx, .xlsx, .pptx |
office_oxide (primary, ~100x faster), OfficeCLI binary (fallback) | office_oxide 0.1.8 (PyO3) |
.ipynb, .latex, .rst, sourcecode, .enex, .chm, etc. (38 formats) |
all2md (when installed) | all2md 1.12 |
.doc, .xls, .ppt |
LibreOffice headless → office_oxide / OfficeCLI | LibreOffice + modern Office engine |
Live mode
headcleaner watch ~/inbox --output ~/out --webhook-url https://hooks.slack.com/...
Re-runs the conversion automatically when files change under ~/inbox.
Each re-run POSTs the manifest to the webhook URL (optional). Press
Ctrl+C to stop.
Obsidian vault sync
headcleaner convert ~/inbox --format okf \
--output ~/Documents/MyVault/Concepts \
--obsidian-compat
Adds Obsidian-friendly flat fields (source, sha256, generated_by,
verified_by, stale_on) to the OKF frontmatter so the concept shows
up correctly in Obsidian's property panel. Original OKF fields stay
intact for round-tripping.
Review (human sign-off)
Auto-conversion sets verified: human:pending. The headcleaner review
TUI walks every pending concept in a bundle and lets a human flip each
to:
- approved →
verified: human:reviewed,status: verified,reviewed_at,reviewed_by,reviewed_via - rejected →
verified: human:rejected,status: rejected, optionalrejection_reasons[] - skipped → leaves the concept as
pending
headcleaner review ./out/okf
# Textual TUI: a=approve, r=reject, s=skip, n=next, p=prev, q=quit
If Textual isn't available (e.g. headless CI), a plain-mode REPL falls back automatically.
Distribution
- PyPI:
pip install headcleaner(built via uv, published via OIDC trusted publishing on tag push) - Homebrew:
brew install headcleaner(formula inpackaging/homebrew/) - Docker:
docker pull ghcr.io/jamesdsizemore/headcleaner-cli(multi-stage image with Tesseract andreadpst) - Windows:
winget install headcleaner,scoop install headcleaner,choco install headcleaner - Static binary:
pip install pyinstaller && pyinstaller packaging/pyinstaller/headcleaner.spec
Full release checklist in RELEASE.md.
CLI surface
headcleaner view <bundle> (add --tui to browse in the terminal) renders an OKF bundle as a single self-contained HTML graph (no backend, opens in any browser). See docs/VIEWER.md for full options.
headcleaner convert IN_DIR [flags] # walk + convert
headcleaner watch IN_DIR [flags] # live mode + webhooks
headcleaner review BUNDLE # human sign-off TUI/REPL
headcleaner attest BUNDLE [--private-key PEM] # Merkle root + optional ed25519 sig
headcleaner verify BUNDLE [--public-key PEM] # verify an attestation
headcleaner serve BUNDLE [--host] [--port] # local HTTP browser for the bundle
headcleaner glob DIR # interactive include REPL (Textual)
headcleaner notion-import EXPORT.zip OUT # reverse a Notion workspace export
headcleaner lint DIR [--fix] # OKF + MD rule checks
headcleaner doctor [--output-dir] # dependency and permission preflight
headcleaner agents [stdout] # emit AGENTS.md
headcleaner templates # list supported formats
Documentation
| Document | Purpose |
|---|---|
| README.md | this file — install, quick start, CLI reference |
| docs/INSTALL.md | all install paths (curl, pip, brew, PowerShell, uv, Docker) |
| docs/USAGE.md | detailed usage guide with worked examples |
| docs/ARCHITECTURE.md | how the pipeline fits together, where to extend |
| docs/FORMAT_MATRIX.md | every supported format × engine × library |
| docs/OKF_NOTES.md | OKF v0.2 contract this CLI emits + trust policy |
| docs/SCHEMA.md | OKF frontmatter JSON Schema and editor/CI integration |
| docs/PLUGINS.md | third-party adapter entry-point protocol |
| docs/TROUBLESHOOTING.md | common errors and fixes |
| docs/FAQ.md | frequently asked questions |
| docs/CONTRIBUTING.md | how to add a new format / engine / emitter |
| docs/CHANGELOG.md | release history |
| docs/ENHANCEMENTS.md | 44+ shipped enhancements + future ideas |
| vscode-extension/README.md | HeadCleaner VS Code extension (Concept Explorer + Trust Inspector) |
Troubleshooting
officecli not found — install with npm install -g @officecli/officecli. Run headcleaner agents to verify.
PDF with no extractable text — your PDF is image-only. Re-run with --ocr (requires pytesseract + Tesseract binary on PATH).
Hidden files skipped — intentional. Files starting with . are dropped by the walker.
OKF index.md missing for root — auto-generated when the bundle has ≥1 concept. Use --no-okf-index to opt out.
More — see docs/TROUBLESHOOTING.md.
Development
git clone <this repo>
cd headcleaner-cli
uv sync
uv run pytest # 314 tests, ~14s
uv run headcleaner convert ./tests/fixtures --format both --output ./out
Architecture
src/headcleaner/
├── walk.py # recursive folder walker
├── router.py # extension → engine dispatch
├── normalize.py # CanonicalDoc + OKF/MD frontmatter builders
├── lint.py # post-conversion linter (OKF + Markdown)
├── run.py # pipeline orchestrator
├── cli.py # Click CLI (headcleaner command)
├── tui.py # Textual TUI (omp-style)
├── engines/
│ ├── base.py # Adapter ABC
│ ├── officecli.py
│ ├── pdf.py
│ ├── html.py
│ └── txt.py
└── emit/
├── markdown.py
├── okf.py
├── okf_index.py
└── manifest.py
Adding a new format: drop a module in engines/, register the adapter in router.py, add a row to docs/FORMAT_MATRIX.md. See docs/CONTRIBUTING.md for the full extension guide.
License
Apache-2.0
Metadata
Release files for headcleaner 0.14.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| headcleaner-0.14.0.tar.gz | 368.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| headcleaner-0.14.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 515.4 kB
Release files / headcleaner-0.14.0.tar.gz
| Download URL | headcleaner-0.14.0.tar.gz |
|---|---|
| Size | 368.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
525758aa05926909482c790f18b66b555b2e98b3e4412241ac356ade709f5f8f
|
|
BLAKE2b-256 checksum How to use checksums |
7fed587fc102f5b82f020ebb0cfa0ea7069a5345b9246894214d843fcf852273
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.
Transparency logRelease files / headcleaner-0.14.0-py3-none-any.whl
| Download URL | headcleaner-0.14.0-py3-none-any.whl |
|---|---|
| Size | 146.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
fe58002a6afb03193a50850e2241f241cccfe06ffa0760d534baffa62c5b7eec
|
|
BLAKE2b-256 checksum How to use checksums |
2018bfd1fc58a273ea59617f61f1c53c02f2a17719f3ac1a8a2e9e0835ae0857
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.
Transparency log