Skip to main content

clean_html_for_llm

This library provides a method to clean HTML content by removing specified tags and attributes while keeping specified attributes. It is particularly useful for preprocessing HTML data to remove noisy tags, making it easier for language models (LLMs) to understand the HTML and generate accurate responses. This is helpful if you are querying LLMs with your HTML data. PyPI Version License

Installation

You can install the clean_html_for_llm library using pip:

pip install clean-html-for-llm

Usage

from clean_html_for_llm import clean_html

html_content = '<div id="main" style="color:red">Hello <script>alert("World")</script></div>'
cleaned_html = clean_html(html_content, tags_to_remove=['script'], attributes_to_keep=['id'])
print(cleaned_html)
# Output: '<div id="main">Hello </div>'

The clean_html function takes the following arguments:

  • html_to_clean (str): The HTML content to clean.
  • tags_to_remove (List[str]): List of tags to remove from the HTML content. Default is ['style', 'svg', 'script'].
  • attributes_to_keep (List[str]): List of attributes to keep in the HTML tags. Default is ['id', 'href'].

You can customize the tags and attributes to remove or keep based on your requirements.

Examples

Example 1:

html_content = '<div id="content" class="main">This is a <span style="font-size: 18px;">paragraph</span>.</div>'
cleaned_html = clean_html(html_content)
print(cleaned_html)
# Output: '<div id="content">This is a <span>paragraph</span>.</div>'

Example 2:

html_content = '<p class="content">Click <a href="https://example.com">here</a> for more information.</p>'
cleaned_html = clean_html(html_content, tags_to_remove=['a'], attributes_to_keep=['class'])
print(cleaned_html)
# Output: '<p class="content">Click </p>'

License

This library is released under the MIT License. See LICENSE for details.

Metadata

Release files for clean-html-for-llm 1.3.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for clean-html-for-llm 1.3.2
File Size Uploaded
clean_html_for_llm-1.3.2.tar.gz 3.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for clean-html-for-llm 1.3.2
File Interpreter ABI Platform
clean_html_for_llm-1.3.2-py3-none-any.whl Python 3 none any Details

Total release size: 7.4 kB

Release files / clean_html_for_llm-1.3.2.tar.gz

Download URL clean_html_for_llm-1.3.2.tar.gz
Size 3.3 kB
Tags Source
SHA-256 checksum
How to use checksums
d39f1fe6761ff84f1917ced1c04985aa2411100f8ea3169475d55c78024971ae
BLAKE2b-256 checksum
How to use checksums
f0efbdaccf53409f54dae633ca584ce420d4f8868745d17e2d534192c0615fe2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.0.0 CPython/3.11.3

Release files / clean_html_for_llm-1.3.2-py3-none-any.whl

Download URL clean_html_for_llm-1.3.2-py3-none-any.whl
Size 4.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0d33c726627be03b63bcc9880ef9f7cab9863969443de31fe4d26d84e7733545
BLAKE2b-256 checksum
How to use checksums
f90f32dab024ae329c19da55a69defecfbd6e658ef00ed6ea440f8aa1e8991eb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.0.0 CPython/3.11.3

Release history Release notifications | RSS feed

This release

1.3.2 This release

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page