llm-html
CMDOP Skill — install and use via CMDOP agent:
cmdop-skill install llm-html
LLM-optimized HTML cleaning: hydration extraction, token budgets, multiple output formats.
Install
pip install llm-html
Quick Start
from llm_html import HTMLCleaner, CleanerConfig, OutputFormat
# Basic cleaning
cleaner = HTMLCleaner()
result = cleaner.clean(html)
print(f"Reduction: {result.stats.reduction_percent}%")
# Hydration-first (extracts SSR data from Next.js, Nuxt, etc.)
if result.hydration_data:
data = result.hydration_data
else:
cleaned = result.html
Convenience Functions
from llm_html import clean, clean_to_json, clean_html, clean_for_llm
# Quick clean
result = clean(html)
# Get JSON if SSR data available, otherwise cleaned HTML
data = clean_to_json(html)
# Pipeline with full control
result = clean_html(html, max_tokens=5000)
result = clean_for_llm(html, output_format="markdown")
Output Formats
from llm_html import to_markdown, to_aom_yaml, to_xtree
md = to_markdown(html)
aom = to_aom_yaml(html)
xtree = to_xtree(html)
Downsampling
Token-budget targeting with D2Snap algorithm:
from llm_html import downsample_html, estimate_tokens
tokens = estimate_tokens(html)
if tokens > 10000:
html = downsample_html(html, target_tokens=8000)
Semantic Chunking
Split large pages into LLM-sized chunks:
from llm_html import SemanticChunker, ChunkConfig
config = ChunkConfig(max_tokens=8000, max_items=20)
chunker = SemanticChunker(config)
result = chunker.chunk(soup)
for chunk in result.chunks:
process(chunk.html)
Shadow DOM
Flatten Web Components for LLM visibility:
from llm_html import flatten_shadow_dom
flat = flatten_shadow_dom(html)
Helpers
from llm_html import html_to_text, extract_links, extract_images, json_to_toon
text = html_to_text(html)
links = extract_links(html, base_url="https://example.com")
images = extract_images(html)
toon = json_to_toon({"key": "value"})
License
MIT
Metadata
Release files for llm-html 0.1.14
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_html-0.1.14.tar.gz | 53.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_html-0.1.14-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 124.5 kB
Release files / llm_html-0.1.14.tar.gz
| Download URL | llm_html-0.1.14.tar.gz |
|---|---|
| Size | 53.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
50dce8f524b9b2d532c87cb09e664f5786d8d1981ae3069fa70279dc6930a79e
|
|
BLAKE2b-256 checksum How to use checksums |
275440733a4aa8dbd78d2429193570879e988969985da414bdb63251e4cede7a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.10.18
|
Release files / llm_html-0.1.14-py3-none-any.whl
| Download URL | llm_html-0.1.14-py3-none-any.whl |
|---|---|
| Size | 70.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d6c5ebddaa8fc5d224437cce89084179da1080cd9159faccea101028a4cd72ab
|
|
BLAKE2b-256 checksum How to use checksums |
da8de6183905dd70ab0b3d3e11a215be506754c3c959a3f80aa248d98bc19c9b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.10.18
|