llm-markdown-sanitizer (Python)
Fix broken markdown that LLMs generate. Zero dependencies, one function.
A Java binding with the same behavior is also available — see the repository root for both.
Install
Requires Python 3.9+. No other dependencies get pulled in.
pip install llm-markdown-sanitizer
Using a virtual environment (recommended for any real project):
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install llm-markdown-sanitizer
Pin a specific version if you want reproducible builds:
pip install "llm-markdown-sanitizer==0.1.0"
Verify it installed correctly:
python -c "from llm_markdown_sanitizer import clean_markdown; print(clean_markdown('**hi**there'))"
# **hi** there
Use
The whole API is one function:
from llm_markdown_sanitizer import clean_markdown
clean_markdown("**Note**this needs a space")
# "**Note** this needs a space"
clean_markdown("| A | B | | --- | --- | | 1 | 2 |")
# "| A | B |\n| --- | --- |\n| 1 | 2 |"
Default settings handle the common failure modes without additional configuration.
In a FastAPI endpoint
A typical place to call this is right before a stored or freshly-generated LLM response goes out to a client:
from fastapi import FastAPI
from llm_markdown_sanitizer import clean_markdown
app = FastAPI()
@app.get("/lectures/{lecture_id}/summary")
def get_summary(lecture_id: int):
raw = db.get_ai_summary(lecture_id) # however you fetch/generate it
return {"summary": clean_markdown(raw)}
Streaming/multi-part LLM responses
Some SDKs return responses as a list of {"text": ...}-shaped chunks instead of one string. clean_markdown accepts that directly:
chunks = [{"text": "# Hello"}, {"text": "\n\nWorld"}]
clean_markdown(chunks)
# "# Hello\n\nWorld"
Why this exists
Ask an LLM to answer in markdown and eventually you'll get: the whole answer wrapped in a stray ```markdown fence, **bold**text glued directly onto the next word, list indentation that's inconsistent within the same response, and tables that are either collapsed onto one line or missing a separator row. Rendering that output as-is breaks the UI.
clean_markdown() fixes all of the above in a single left-to-right pass over the text — no whole-string regex backtracking, so it stays fast on long documents.
What it fixes
| Problem | Before | After |
|---|---|---|
| Wrapping code fence | ```markdown\n# Title\n``` |
# Title |
<br> outside tables |
Line one<br>Line two |
Line one\nLine two (left untouched inside table cells, where it's usually intentional) |
| Bold glued to text | **Note**this breaks |
**Note** this breaks |
| Inconsistent list indent | mixed 2/3/tab indents | normalized to 4 spaces per nesting level |
| Collapsed table | | A | B | | --- | --- | | 1 | 2 | |
proper one-row-per-line table |
| Broken table (no separator / mismatched columns) | renders as a wall of | |
dropped instead of rendering broken |
Protecting your own syntax
If your prompts produce custom tokens — a [[wiki]]-style syntax, template placeholders, etc. — that the cleanup passes above might mangle, they can be excluded explicitly:
import re
clean_markdown(text, protect_patterns=[re.compile(r"\[\[.*?\]\]")])
Origin
Extracted from the markdown-cleanup layer of a production RAG service, after months of hardening against real LLM output. The domain-specific parts — a custom wiki syntax, a Korean-language note pattern — were removed in favor of the general protect_patterns mechanism above, so callers can supply their own domain syntax instead.
Contributing
Bug fixes and small improvements are welcome. No CLA/DCO required — see CONTRIBUTING.md for guidelines and how to run the test suite locally. AI coding agents should pick up AGENTS.md / CLAUDE.md automatically.
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file llm_markdown_sanitizer-0.1.0.tar.gz.
File metadata
- Download URL: llm_markdown_sanitizer-0.1.0.tar.gz
- Upload date:
- Size: 10.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
29828c4ba3e34d051af75853b2e0f09110eb85f80b2a72f2eb6710d0b42d318d
|
|
| MD5 |
c507c2a6beddcfabb564ddafd88d3b7b
|
|
| BLAKE2b-256 |
c0c6a95c2ce472bb71f3d4553c49d236d1d21bef44700129f15ac28419c1b4c4
|
Provenance
The following attestation bundles were made for llm_markdown_sanitizer-0.1.0.tar.gz:
Publisher:
python-publish.yml on stlahxm/llm-markdown-sanitizer
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
llm_markdown_sanitizer-0.1.0.tar.gz -
Subject digest:
29828c4ba3e34d051af75853b2e0f09110eb85f80b2a72f2eb6710d0b42d318d - Sigstore transparency entry: 2498300932
- Sigstore integration time:
-
Permalink:
stlahxm/llm-markdown-sanitizer@6eaf30e1383f76b438c9e42fe9d199221b082dc4 -
Branch / Tag:
refs/tags/python-v0.1.0 - Owner: https://github.com/stlahxm
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@6eaf30e1383f76b438c9e42fe9d199221b082dc4 -
Trigger Event:
push
-
Statement type:
File details
Details for the file llm_markdown_sanitizer-0.1.0-py3-none-any.whl.
File metadata
- Download URL: llm_markdown_sanitizer-0.1.0-py3-none-any.whl
- Upload date:
- Size: 11.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0d1df81ddd7bc2e0afbbef31bd02e782031a5140db75cfaaa3ef488da49d0256
|
|
| MD5 |
8242c475870b9d89528ef91e01ee8304
|
|
| BLAKE2b-256 |
c9d4d1851b5e56dd3d82a67cedf0a1a04d3569e84ab980bd6739463df2d58969
|
Provenance
The following attestation bundles were made for llm_markdown_sanitizer-0.1.0-py3-none-any.whl:
Publisher:
python-publish.yml on stlahxm/llm-markdown-sanitizer
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
llm_markdown_sanitizer-0.1.0-py3-none-any.whl -
Subject digest:
0d1df81ddd7bc2e0afbbef31bd02e782031a5140db75cfaaa3ef488da49d0256 - Sigstore transparency entry: 2498300955
- Sigstore integration time:
-
Permalink:
stlahxm/llm-markdown-sanitizer@6eaf30e1383f76b438c9e42fe9d199221b082dc4 -
Branch / Tag:
refs/tags/python-v0.1.0 - Owner: https://github.com/stlahxm
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@6eaf30e1383f76b438c9e42fe9d199221b082dc4 -
Trigger Event:
push
-
Statement type: