grab2md
grab2md is a local-first HTML-to-Markdown tool. Its CLI converts individual
local HTML files and HTTP(S) URLs, processes link lists in batches, and archives
bounded site crawls. Its unpacked Chrome extension converts the active page,
main article, or current selection.
[!IMPORTANT] This is an alpha-stage, pre-1.0 release. The primary workflows are covered by end-to-end tests, but compatibility and real-world extraction quality remain under evaluation. The Python CLI release channel is PyPI; the Chrome extension remains an unpacked development release and is not published in the Chrome Web Store. Review the limitations and security boundaries before using either interface on sensitive or unattended workloads.
Status and support
- Release version:
0.4.2 - Release channel: public PyPI alpha
- Tested Python versions: 3.11, 3.12, and 3.13
- Production PyPI distribution:
grab2md - Installed command and Python import:
grab2md - Required gates: tests and production coverage, Ruff, Black, mypy, requirement export consistency, wheel smoke, extension runtime tests, Bandit, and dependency audit
- TestPyPI artifacts are staging-only; no Web Store release or stable Python API compatibility promise exists yet
The primary tested paths are local conversion, URL conversion, batch link processing, sequential crawling, interruption/resume, configuration recovery, and the extension's full-page/article/selection conversion modes.
Installation
Create an isolated environment and install the CLI from PyPI:
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install grab2md
grab2md --help
To install the current source tree instead:
git clone https://github.com/jkindrix/grab2md.git
cd grab2md
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install .
grab2md --help
Contributors running the complete development and release gates need Poetry 2.4.1 and Node.js; see CONTRIBUTING.md for that pinned toolchain and its canonical commands.
JavaScript rendering is an isolated optional installation:
python -m pip install "grab2md[render]"
python -m playwright install chromium
grab2md https://example.com/app --render-js
See docs/browser-rendering.md for its resource,
network, and authentication boundaries.
The production distribution is
grab2md on PyPI. TestPyPI artifacts are
release rehearsals and are not supported production releases.
Quick start
Convert a URL or local HTML file:
grab2md https://example.com --output example.md
grab2md page.html --output page.md
The hidden compatibility alias grab2md convert SOURCE remains accepted for
pre-release scripts. Direct sources are the primary interface and the alias is
deliberately omitted from the command table. Run grab2md help to inspect every
direct-conversion option without promoting that alias.
Process Markdown files or plain URL lists and rewrite links between successful local outputs:
grab2md batch links.md urls.txt --output-dir documentation
Crawl sequentially with robots.txt enabled by default:
grab2md crawl https://docs.example.com \
--output-dir documentation \
--max-depth 3 \
--max-pages 100 \
--rate-limit 30
Inspect and resume crawl state:
grab2md state list
grab2md state info CRAWL_ID
grab2md state resume CRAWL_ID
Run grab2md help for the complete, configuration-aware direct-conversion
option list. Run grab2md COMMAND --help for batch, crawl, configuration, and
state workflows.
Commands
| Command | Purpose |
|---|---|
grab2md SOURCE... |
Convert one or more HTTP(S) URLs or local HTML files. |
help |
Show every option accepted by the primary direct-conversion form. |
batch |
Extract links from input files, convert them, and rewrite successful local links. |
crawl |
Recursively fetch and convert pages using a sequential, robots-aware policy. |
config |
Inspect, validate, back up, restore, and change configuration. |
state |
List, inspect, export, import, clean, and resume crawl state. |
Python API support
The supported pre-1.0 interface is the grab2md command (and equivalent
python -m grab2md entry point). The installed Python modules are internal and
may change between alpha releases; importing conversion, crawler, cookie, or
transport implementation modules is not a supported compatibility contract.
grab2md.__version__ is exposed for metadata inspection only. A library API
will be considered separately if real use cases establish the required result,
exception, and compatibility contract.
Global options include --log-level (default WARNING), --debug-log,
--banner, and metadata-backed --version.
Conversion
Useful options include:
--content full|main|selectorfor explicit content selection (full is the whole-document default), with--selectorrequired by selector mode;--output/-oto write one source to a file instead of stdout (multiple sources cannot share one output path);--cookie-jsonfor an owner-only portable cookie export on every platform;--browser-cookiesfor compatible Firefox databases or the narrow legacy Windows Chrome DPAPI path, optionally with a one-shot--cookie-paththat does not modify global configuration;--headers-filefor an owner-only JSON object of target request headers;--storage-statewith--render-jsfor owner-only Playwright session state;--enhanced-headers/--basic-headersand--user-agent-contactfor an honest, versioned request identity;--download-imageswith a configurable--images-dir;--allow-private-networkonly for explicitly trusted intranet, loopback, or development destinations;--insecureonly for trusted hosts with invalid certificates; and--fancyfor decorated progress output.
Automatic Firefox database extraction is supported on Windows, macOS, and
Linux for recognized profile/database schemas. Current Chrome normally uses
app-bound (v20) encryption, so direct database extraction is generally
unavailable and fails closed with export guidance. Only legacy Windows
DPAPI-backed Chrome keys are supported; Chrome Keychain/keyring retrieval is
not implemented on macOS or Linux. Exported, owner-private cookie JSON is the
primary portable authentication path on every platform. Password submission is
not supported.
Batch output
Batch mode supports preserved paths, flattened domain output, a single flat directory, hierarchical domain folders, optional visualization, quiet output, and a Markdown report. Only successfully written files enter the local-link mapping; failed URLs remain remote links. Link rewriting is isolated per artifact, but any read or atomic-write failure makes the overall batch result partial and exits nonzero while preserving successfully converted files.
Crawl policy
Crawls are intentionally sequential. Available controls include:
--follow(domain-onlyfor the same host and port,host-onlyfor the exact hostname,subdomainfor that hostname and dot-delimited descendants, or an explicit regular expression);--max-depth, a cumulative per-start--max-pagespage-attempt budget that includes failures and explicit retries, and jittered--delay;--respect-robots/--ignore-robots;- hard-maximum requests-per-minute
--rate-limitfor each destination origin (host and port), with adaptive slowing and a circuit breaker; --politefor at least one second between sequential requests and twice any larger explicit delay;- content selection, progress, output layout, visualization, and quiet-mode switches.
Robots access rules follow RFC 9309 semantics: duplicate exact product-token
groups are combined case-insensitively, wildcard groups apply only when no exact
group exists, the longest matching rule wins, equivalent Allow rules win
ties, URI octets are normalized for comparison, and */$ patterns are
supported. A missing (4xx) robots file permits access; a network error or 5xx
response fails closed. Crawl-delay is honored as a non-standard extension.
Ctrl+C or termination checkpoints the active crawl and then preserves normal
signal behavior. Deferred URLs remain queued instead of being silently lost.
Final local-link rewriting is isolated per artifact, but any rewrite failure
makes the crawl or resume result partial and exits nonzero without deleting
successfully archived pages.
Configuration and state
Configuration is stored in the grab2md directory beneath the platform config
root ($XDG_CONFIG_HOME or ~/.config on Linux, Application Support on
macOS, and %APPDATA% on Windows). Writes are validated, atomic, backed up,
and recoverable. GRAB2MD_CONFIG_PATH overrides the complete file path. Run:
grab2md config show
grab2md config path
grab2md config show-options
grab2md config set-cli-default convert content_mode main
grab2md config set-cli-default crawl max_pages 250
grab2md config backup
grab2md config list-backups
convert in these configuration commands is the persisted namespace for
direct-conversion defaults and the hidden compatibility alias; users still run
direct conversions as grab2md SOURCE....
CLI defaults are typed and loaded at invocation time. Optional values accept
null; invalid updates fail without replacing the existing file. Concurrent
configuration changes from separate processes use last-write-wins semantics,
so serialize configuration commands in automation.
CSS selectors are caller-owned generic inputs; grab2md does not ship or
silently apply per-site extraction profiles. A selector can be supplied for one
run or configured as a CLI default together with content_mode=selector.
Crawl state supports list, resume, clean, export, import, and info.
State files default to ~/.grab2md/states, independently of the platform config
root, and use restrictive permissions on POSIX systems. state list prints
complete reusable IDs; other state commands also accept an unambiguous prefix
of at least eight characters. Identifiers and resolved state paths are confined
to that state directory.
Chrome extension
Load extension/ as an unpacked Manifest V3 extension in Chrome or Chromium.
The supported workflow operates on the active tab and provides:
- full-page, main-article, and current-selection conversion;
- preview, clipboard copy, and Markdown download;
- theme and conversion settings; and
- packaged Mozilla Readability extraction only when article mode is selected.
The extension uses activeTab, scripting, storage, downloads, and
clipboardWrite; it does not request persistent access to every site. URL-list,
batch, native CLI integration, background service-worker conversion, context
menus, and keyboard shortcuts are not supported.
See extension/README.md for installation and testing.
Security boundaries
- With a stable caller-owned output-root hierarchy, crawl and batch reject or sanitize page-derived traversal and symlink escapes. A concurrent local process that can rename or replace the selected root or one of its ancestor directories is outside this containment boundary; do not place output under a directory hierarchy writable by an untrusted local account.
- Browser cookie databases are copied into unpredictable owner-private temporary directories and removed after success, failure, or interruption.
- Configuration and crawl states are atomically replaced using
0600files in0700directories on POSIX systems. - Project-owned diagnostic logs are owner-only on POSIX systems and redact URL
userinfo, credential-bearing headers, complete cookie values, and recognized
credential query/fragment fields before emission. The rotating log defaults
to the platform's per-user application-log/state directory; set
GRAB2MD_LOG_PATHto choose an explicit file. Embedding applications remain responsible for handlers they attach to third-party or root loggers. - Target authentication accepts scoped browser cookies, an owner-only JSON cookie export, header file, or Playwright storage-state file for rendered conversion. Login flows remain outside grab2md, and authentication files are rejected when group- or world-readable or when the supplied final path is a symlink on POSIX systems. Resolution of caller-owned authentication-file ancestor directories remains an operating-system/filesystem trust boundary.
- Remote pages, crawl targets, robots files, and images allow only HTTP(S), resolve each origin once, connect only to validated numeric addresses, and manually revalidate redirects. Private, loopback, link-local, metadata, IPv4-mapped, 6to4, IPv4-compatible, and well-known NAT64 encodings of non-public IPv4 destinations are blocked by default. HTTPS redirects cannot downgrade to HTTP. Guarded traffic bypasses configured and environment proxies because proxy-side DNS would defeat address pinning.
- Static page/crawl responses are capped at 10 MiB and robots files at 1 MiB. Image acquisition additionally verifies MIME type and file signature, rejects every SVG image, and enforces 10 MiB per-image and 50 MiB per-conversion limits.
--allow-private-networkexplicitly relaxes destination classification for trusted internal or development targets; DNS pinning, redirect validation, URL validation, response limits, and TLS verification remain active.- Local image copying is restricted to regular files beneath the source HTML directory; parent traversal and symlink escapes are rejected.
--insecuredisables TLS verification and should be used only for hosts you control. It exposes the connection to interception but does not authorize private-network access.
See docs/network-security.md for the complete
outbound request contract and maintainer integration rule.
Windows relies on the current account's directory ACLs because POSIX mode bits are unavailable there.
Output contract
Remote relative links and image references are resolved against the final URL
and the document's valid <base> element. --metadata adds deterministic YAML
front matter for available title, author/date, canonical URL, description, and
language fields. Local references remain relative. See
docs/output-contract.md for the exact contract.
Known limitations
- Conversion uses
markdownify. Full-document mode deliberately preserves authored boilerplate; opt-in main-content mode uses substantial semantic regions and a confidence-gated readability fallback, whose quality varies by document. The measured decision is recorded inADR 0002. - JavaScript rendering is opt-in for direct URL conversion; batch and crawl remain static.
- Metadata extraction intentionally uses declared HTML/meta fields rather than text inference or executable structured data.
- Crawling is sequential; removed concurrency options are not advertised.
- Browser cookie support is deliberately bounded: Firefox database schemas can
change, while current Chrome's app-bound (
v20) cookies are not available to automatic extraction. The Chrome database path is limited to legacy Windows DPAPI-backed keys and excludes macOS Keychain and Linux keyring formats. Use an owner-private cookie JSON export as the primary portable path. - The extension must be installed unpacked; no Web Store release exists.
Development
Run the canonical local gates from a clean checkout:
poetry sync --with dev
poetry check
poetry run pre-commit run --all-files
poetry run pre-commit run --all-files --hook-stage pre-push
node --test extension/tests/*.test.js
node extension/tests/chromium-smoke.js
./deploy.sh --dry-run
Coverage uses the production package as its denominator and enforces the floor
documented in docs/coverage.md.
Project documentation
CHANGELOG.md: user-visible changes and changelog policyCONTRIBUTING.md: development setup and contribution gatesSECURITY.md: private vulnerability-reporting routeSUPPORT.md: alpha support and compatibility policydocs/releasing.md: reproducible release checklistdocs/deployment.md: local deployment detailsdocs/configuration-example.md: configuration exampledocs/adr/0001-defer-scale-out-crawling.md: scale-out architecture gatedocs/internal/: historical planning and review records
License
Copyright (c) 2025-2026 Justin Kindrix. Distributed under the MIT License.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file grab2md-0.4.2.tar.gz.
File metadata
- Download URL: grab2md-0.4.2.tar.gz
- Upload date:
- Size: 111.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2bd9f54c0d1779a023010f73811e45717cecf0e5edd900d2d7f43e6a67b97c12
|
|
| MD5 |
1c8a641b63acaf89bf3173c160526e13
|
|
| BLAKE2b-256 |
a1d98bf7d7b4a572a9f3111d45316019c9afff51faf38d5b13457b2e5aa556f3
|
Provenance
The following attestation bundles were made for grab2md-0.4.2.tar.gz:
Publisher:
publish.yml on jkindrix/grab2md
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
grab2md-0.4.2.tar.gz -
Subject digest:
2bd9f54c0d1779a023010f73811e45717cecf0e5edd900d2d7f43e6a67b97c12 - Sigstore transparency entry: 2241891780
- Sigstore integration time:
-
Permalink:
jkindrix/grab2md@958660f4c90bfb8d0aabff9f00be61ad1cea390a -
Branch / Tag:
refs/tags/v0.4.2 - Owner: https://github.com/jkindrix
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@958660f4c90bfb8d0aabff9f00be61ad1cea390a -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file grab2md-0.4.2-py3-none-any.whl.
File metadata
- Download URL: grab2md-0.4.2-py3-none-any.whl
- Upload date:
- Size: 135.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
40a71be77ec444b0b5d73edfbb6259ed4826a2d761803e7b744173e6d976538c
|
|
| MD5 |
4d542f7b4105dbe6f3232d5631dc0f6f
|
|
| BLAKE2b-256 |
be9f1ae59bb8ad4a354249e5a55e5b2b84515f9f40115fc288233d4dac5b1e7b
|
Provenance
The following attestation bundles were made for grab2md-0.4.2-py3-none-any.whl:
Publisher:
publish.yml on jkindrix/grab2md
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
grab2md-0.4.2-py3-none-any.whl -
Subject digest:
40a71be77ec444b0b5d73edfbb6259ed4826a2d761803e7b744173e6d976538c - Sigstore transparency entry: 2241892043
- Sigstore integration time:
-
Permalink:
jkindrix/grab2md@958660f4c90bfb8d0aabff9f00be61ad1cea390a -
Branch / Tag:
refs/tags/v0.4.2 - Owner: https://github.com/jkindrix
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@958660f4c90bfb8d0aabff9f00be61ad1cea390a -
Trigger Event:
workflow_dispatch
-
Statement type: