flint-slating
MCP server that reads PDFs and exposes them to LLM consumers as structured Markdown, plus the usual ancillaries: metadata, outline, images, tables.
Designed to pair with a separate "wiki" MCP server that handles the
writing side — an agent calls flint-slating to read PDFs and another
MCP to persist notes about them into a frontmattered-markdown knowledge
base.
What it does
Built on a permissive-license PDF stack:
| Library | License | Role |
|---|---|---|
| Docling | MIT | PDF → Markdown with heading hierarchy, multi-column reading order, and Markdown tables |
| pypdf | BSD-3 | metadata, TOC, page count, encryption checks, image enumeration |
| pdfplumber | MIT | per-page table extraction |
There is no PyMuPDF, no MuPDF, no AGPL or GPL anywhere in the dependency tree. A CI license-check job rejects PRs that pull in copyleft transitive deps.
Transports
Two transports off the same MCP server, selected via --transport:
| Transport | Run via | Use case |
|---|---|---|
| Streamable-HTTP (default) | uvx flint-slating or --transport http |
Long-lived local daemon, container, or shared service. |
| stdio | uvx flint-slating --transport stdio |
The standard MCP integration shape — drop into claude_desktop_config.json or any mcp.json. |
Run
As an HTTP daemon (default)
uvx flint-slating # listens on PORT (default 35833)
curl http://127.0.0.1:35833/health
Or pin it:
uv tool install flint-slating
flint-slating
As a stdio MCP server
uvx flint-slating --transport stdio
Wire into Claude Code's MCP config:
{
"mcpServers": {
"flint-slating": {
"command": "uvx",
"args": ["flint-slating", "--transport", "stdio"]
}
}
}
Docker
docker run --rm \
-p 35833:35833 \
-v $(pwd)/pdfs:/pdfs:ro \
-v flint-slating-data:/data \
ghcr.io/parkviewlab/flint-slating:latest
Or use docker-compose.yml for a persistent stack.
MCP tools
All PDF tools take a source argument with one of:
{"path": "/abs/path/to/file.pdf"}— local file{"url": "https://..."}— streamed to a content-addressed cache{"bytes_b64": "...", "filename": "x.pdf"}— base64 upload (size-capped)
| Tool | What it does |
|---|---|
pdf_info |
{page_count, metadata, is_encrypted, sha256} |
pdf_toc |
flat outline [{level, title, page}] |
pdf_read_text |
plain text by page range (fast — pypdf, no ML) |
pdf_read_markdown |
high-quality Markdown via Docling (hybrid sync/async — see below) |
pdf_read_chunks |
per-page Markdown chunks with tables/images/toc_items (hybrid sync/async) |
pdf_list_images |
enumerate images: [{page, index, name, width, height, ext}] |
pdf_extract_image |
base64 bytes of one image |
pdf_find_tables |
per-page Markdown tables via pdfplumber |
get_job_status |
poll a background job |
get_job_result |
fetch a finished job's artifact |
cancel_job |
cancel a running job |
Hybrid sync/async
pdf_read_markdown and pdf_read_chunks run inline when
page_count <= SYNC_PAGE_THRESHOLD (default 20). For larger PDFs they
queue a background job and return a job_id — poll get_job_status
until state=="done", then call get_job_result (or, in HTTP mode,
fetch output_url directly).
stdio mode transparently waits for the job inline — there's no HTTP server to download from, so the originating tool call blocks until the result is ready and returns it directly.
HTTP endpoints (HTTP mode only)
GET /health—{ok, version, uptime_seconds}GET /admin/version— package and dependency versions, Docling model statusGET /admin/jobs— recent job listGET /outputs/{job_id}/result.md— finished MarkdownGET /outputs/{job_id}/result.json— finished chunked outputGET /outputs/{job_id}/log.jsonl— append-only job logPOST /sse— MCP Streamable-HTTP transport
Configuration
| Env var | Default (daemon) | Default (container) | Purpose |
|---|---|---|---|
PORT |
35833 |
35833 |
HTTP bind port |
HOST |
0.0.0.0 |
0.0.0.0 |
HTTP bind address |
OUTPUT_ROOT |
./output |
/data/output |
Per-job output dirs |
CACHE_ROOT |
./cache |
/data/cache |
Materialized URL / base64 PDFs |
OUTPUT_EXPIRY_DAYS |
7 |
7 |
Sweep finished jobs older than N days |
MAX_INLINE_PDF_BYTES |
25 MB |
25 MB |
Cap on base64 upload size |
MAX_URL_PDF_BYTES |
200 MB |
200 MB |
Cap on URL download size |
SYNC_PAGE_THRESHOLD |
20 |
20 |
Inline-vs-job cutoff for Markdown conversion |
DOCLING_ARTIFACTS_PATH |
~/.cache/docling |
/opt/docling-models |
Docling layout-model cache |
ENABLE_OCR |
false |
false |
Enable Docling OCR (Tesseract required) |
PUBLIC_BASE_URL |
http://localhost:35833 |
http://localhost:35833 |
Used to build output_url |
Resource notes
- Docling downloads a ~200–500 MB layout model on first use. The
container image does not pre-fetch it (pre-fetching dominated
multi-arch build time under QEMU); the daemon warms it on startup,
and the first user-facing call pays the download. Operators can
populate
DOCLING_ARTIFACTS_PATH(default/opt/docling-modelsin the container) via volume mount for a hot start. - pypdf, pdfplumber, and the URL / base64 paths are fast and have no ML
overhead — use
pdf_info,pdf_toc,pdf_read_text, andpdf_find_tableswhenever Markdown isn't strictly needed.
Releasing
Tag-driven CI publishes to both PyPI (flint-slating) and GHCR (ghcr.io/parkviewlab/flint-slating):
# Bump version in pyproject.toml first, then:
git tag v0.1.0
git push origin v0.1.0
The release workflow refuses a tag that does not match pyproject.toml's version, that still carries a dev marker (.devN), that is not on origin/main, or that is not greater than the previous release tag.
Commit message convention
After the publish jobs, a changelog job generates the new CHANGELOG.md section — an LLM-written "Highlights" paragraph plus a categorized list written by dev-tools' generate-changelog — commits it back to main, and creates the GitHub Release with the same content as its body. Categorization uses Conventional Commits prefixes (the full list is in the ParkviewLab handbook's commits-and-changelogs.md):
| Title | Group in the notes | Notes |
|---|---|---|
any type with ! after it (feat!:), or a breaking-change footer |
Breaking changes | listed there once, whatever its type |
feat: |
Features | user-visible |
fix: |
Bug fixes | user-visible |
perf: |
Performance | user-visible |
refactor: |
Refactor | |
docs: |
Docs | |
test: |
Tests | |
revert: |
Reverts | GitHub's Revert button titles a PR Revert "…", which has no type |
build: / chore: / ci: / style: |
Maintenance | |
| any other title | Other changes | the whole title |
| a commit with no pull request | Direct commits | its subject and short hash |
A title without a recognised type is not dropped: it is listed whole under Other changes. So prefix your PR titles, and correct a title before the merge, since retitling afterwards does not change the commit. The groups appear in the order above, and an empty group is left out.
License
Licensed under either of
- Apache License, Version 2.0 (LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0), or
- MIT license (LICENSE-MIT or http://opensource.org/licenses/MIT)
at your option. In SPDX terms: MIT OR Apache-2.0.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in this work by you shall be dual-licensed as above, without any additional terms or conditions. See LICENSING.md.
flint-slating only depends on permissive-licensed libraries; the CI
license-check job enforces this on every PR. torch and torchvision are pinned
to the CPU-only PyTorch wheel index so
the distribution does not bundle NVIDIA's proprietary CUDA libraries. Inference
runs on CPU on Linux/Windows and on MPS (Metal) on Apple Silicon. See
THIRD_PARTY_LICENSES.md for the per-dependency
license breakdown.
© 2026 Gary Frattarola · Licensed under MIT OR Apache-2.0 · part of ParkviewLab
Metadata
Release files for flint-slating 0.1.8
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| flint_slating-0.1.8.tar.gz | 176.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| flint_slating-0.1.8-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 211.2 kB
Release files / flint_slating-0.1.8.tar.gz
| Download URL | flint_slating-0.1.8.tar.gz |
|---|---|
| Size | 176.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ef7afefde0b167b82682862565d604902bb21d24cbf10ab71fed9591754d7541
|
|
BLAKE2b-256 checksum How to use checksums |
c4de0074a578f89956598d738789e5e041bb1d0a358730cb845f0ef551e8797b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency logRelease files / flint_slating-0.1.8-py3-none-any.whl
| Download URL | flint_slating-0.1.8-py3-none-any.whl |
|---|---|
| Size | 35.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0f4cffc1397970ca514a822b96622083faad733ec4af168b3531b9302195ac6a
|
|
BLAKE2b-256 checksum How to use checksums |
d1d5f32ad54fac294ba6fdf6277ef3c462a9a0913e09761ecb76b81efd46a903
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency log